This Australian Research Council Discovery Project is a cross-institutional collaboration between ANU (Glenn Roe and Robert Wellington), The University of Melbourne (Erin Helyard), The University of Sydney (Mark Ledbury), and Oxford University (Nicholas Cronk). Through the study of a unique and ambitious eighteenth-century songbook – Jean-Benjamin de Laborde’s Choix de Chansons (1773) – our project provides a workable solution to these questions by way of the notion of ‘transdisciplinarity’. First described by the developmental psychologist and philosopher Jean Piaget as a superior stage of interdisciplinary relationships, the transdisciplinary approach implies a total system of interrelated knowledge without established disciplinary boundaries; a system that has much in common with that imagined by Diderot and d’Alembert in the Encyclopédie. We propose that the Choix de Chansons is also in many ways a quintessential transdisciplinary object. As such, it requires a new methodological approach that operates at the interface of interdisciplinary collaboration, rich historical contextualization, and new media dissemination.

This is the first project of its kind to address the complex transdisciplinary and transmedial nature of both Laborde’s Choix de Chansons, and of eighteenth-century print culture more generally. The complementary disciplines of musicology, art history, and French literature will create a unique transdisciplinary matrix in which our team will ‘perform’ Laborde’s text in order to recreate its original modes of reception, evoking the ways eighteenth-century participants appreciated, decoded, and debated the intersections of music, visual art, and literature.

Engraving from Laborde's Choix de chansons

Engraving from Laborde’s Choix de chansons

Most often, this kind of cultural consumption was enacted publicly and sensationally at the opera. But, unlike such multimedia events, and based on the quasi-democratic principles of Masonic culture, Laborde’s text is meant for a small community of like-minded individuals who commingle their performative experiences in the intimacy of a salon around a single instrument: harp or harpsichord. In many respects, Laborde’s project aims to reproduce – albeit, in miniature – the operatic experience by simulating the close connections between image, music and text. The novel aspect in Laborde’s scenarios is that these connections take place not in the public arena of the opera box, where spectators perform only as audience members, but rather in the chamber, where the participants are no longer merely spectators but themselves performers.

This emphasis on individual expression as a meaningful component of a close engagement with others in a culture of sociability and sensibilité echoes the contemporaneous musical, philosophical and social trends. By implication, any attempt at recapturing the creative, receptive, and performative complexity of Laborde’s songbook – or any other complex cultural artefact for that matter – today requires new models of cross-disciplinary collaboration and multimedia dissemination. Our project will provide one such model, aimed at reproducing digitally the cultural context of the Chansons, both as an object of transdisciplinary communication – one that actively speaks from the nexus of image music, and text – and as a product of the cultural and intellectual networks of the time, from courtly and salon culture to the more progressive sociability of the Masonic societies.

By moving the Chansons from print to digital media we can not only incorporate multiple layers of remediation (image, music, text), but also shed greater light on the various strata – social, intellectual, political, philosophical – that informed the work’s production and its relationship to the cultural networks mentioned above. We will develop a new digital edition of the Chansons that will present high-resolution scans of its pages and engravings alongside transcriptions of the poetry and recordings of the songs. This juxtaposition will expose the image-music-text relationship inherent to the illustrated songbook.

text by Glenn Roe, Erin Helyard, Mark Ledbury, and Robert Wellington.

Digitizing Raynal

A collaborative digital research project

On the heels of Cecil Courtney and Jenny Mander’s recent publication, Raynal’s ‘Histoire des deux Indes’ colonialism, networks and global exchange (OSE, 2015), I am pleased to announce a new international research project aimed at further exploring Raynal’s monumental work and its impact on Enlightenment thought. Thanks to the generous support of the Consortium for the Study of the Premodern World at the University of Minnesota, the Centre for Digital Humanities Research at the Australian National University, Stanford University Libraries, and The ARTFL Project at the University of Chicago, we have recently completed the digitization and text encoding (in TEI-XML) of the three primary editions of the Histoire philosophique et politique des établissements et du commerce des Européens dans les deux Indes. These editions – the first edition of 1770, the second of 1774, and the 1780 third edition – were those that Raynal himself oversaw during his lifetime.

Our digital editions are based on high quality PDFs provided by the BNF’s Gallica online library (1770 and 1780 editions) and the Bodleian’s Oxford Google Books Project (1774 edition). A preliminary search interface has been built using the ARTFL Project’s PhiloLogic software and can be accessed here: Raynal search form. Users can query one or all of the above editions, which represent the first publicly available full-text digital edition(s) of the Histoire des deux Indes. In the coming months we will release a new version of the database running on ARTFL’s state-of-the-art PhiloLogic4 system, along with a preliminary ‘intertextual interface’ that will aim to incorporate the text of the three separate editions into one reading interface.


Title page and frontispiece of the 1780 edition of Raynal’s Histoire des deux Indes (Gallica).

Diderot, Hornoy, and the 1780 edition

What is perhaps most exciting about these new digital resources is the inclusion of a unique 1780 edition of the Histoire des deux Indes recently made available by the BNF. Acquired at public auction in March 2015, this particular edition had been conserved since the late 18th century in the private library of Alexandre Marie Dompierre d’Hornoy (1742-1828). A lawyer at the Parlement de Paris and great-nephew of Voltaire – he in fact inherited Jean-Baptiste Pigalle’s infamous nude statue of Voltaire upon his great-uncle’s death – Hornoy corresponded with many of the philosophes, Diderot included. His copy of the Histoire contains pencil marks in the margins of some passages, an unremarkable fact, perhaps, were it not for a note written by Hornoy just above a three-page insert at the beginning of the first tome. The handwritten tables included in the insert list all the sections marked in pencil over the four volumes of text: ‘mourceaux qui sont de M. Diderot’, Hornoy writes, ‘marqués en crayon par Mme de Vandeul’. Madame de Vandeul was, of course, Diderot’s daughter.


Handwritten insert of the 1780 edition (Gallica)

The existence of such an annotated volume of the Histoire was posited in the 19th century, notably by Joseph Marie Quérard in his Supercheries littéraires dévoilées (5 vols., 1845-1856). Quérard claimed that there supposedly existed a copy of the 1780 edition on which Diderot himself had marked in pencil all the passages that belonged to him [1]. According to Quérard, this copy became the property of Madame de Vandeul shortly after Diderot’s death. Whether or not the copy acquired by the BNF is the same as that owned by Vandeul we cannot say for sure, but Herbert Dieckmann, in his inventory of the ‘fonds Vandeul’, also mentions the hypothetical existence of a copy of the in-4o edition (e.g. 1780) that was purportedly annotated by hand, but that had since been lost [2].

Some preliminary experiments

While consensus as to the validity of Hornoy’s assertion that the marked sections are in fact those authored by Diderot will most likely take years to accrue, we can begin, using the new digital edition, to ask some basic questions as to the authorship claims indicated in the text. Thanks to extensive markup in TEI-XML notation, sections purportedly belonging to Diderot are clearly indicated, and perhaps more importantly, can be extracted as one test corpus. Using some basic statistical measures drawn from authorship attribution studies, or Stylometry, we can begin to think about how the ‘Diderot’ sections may, or may not, differ stylistically – i.e. in terms of comparative word usage over the most common words, an established metric of ‘authorship’ in stylometry and forensic linguistics – from the rest of the text.


Page from 1780 edition with ‘Diderot’ section marked in pencil (Gallica)

Working with the Centre for Literary and Linguistic Computing at the University of Newcastle (Australia), and in particular with their Intelligent Archive software for stylistic and statistical text analysis, we extracted the top 200 words for each ‘author’ (e.g. those drawn from sections putatively by Diderot, and the remaining ‘Raynal’ sections). As a result, we were left with 4 ‘Diderot’ tomes (containing all of the text marked in pencil) and 4 ‘Raynal’ tomes (containing the remainder), representing their unique word lists over the entire edition. For a first preliminary test, we ran a cluster analysis on the 8 tomes to see if they would cluster together or separately:


Cluster analysis of ‘Diderot’ tomes vs. ‘Raynal’ tomes, based on top 200 word lists

Cluster analysis works by separating (or clustering) the most similar texts first and the most distinct last, in this case into 2 branches. A division like the one above, clearly separated into two distinct ‘trees’ is a very clear indication that the texts in each of the two branches are highly likely to be those of two different authors.

Principal component analysis (PCA) provides another method of examining our corpora. PCA is a procedure for identifying a smaller number of uncorrelated variables, called ‘principal components’, from a large set of data. The goal of PCA is to explain the maximum amount of variance with the fewest number of principal components. In our case, it is a technique that allows for the first two principal components of our two sets of texts, i.e. their word variance, to be plotted on a bi-axial or two-dimensional graph. One of these plots (using the 100 most frequent words of the full text) with both text corpora divided into 10,000 word blocks, is shown below.


Principal component analysis using 10,000 word blocks and 100 most frequent words

The disparity in size of our two test corpora meant that while there were 68 text sections for Raynal (in green), there were only 14 for Diderot (in blue). Nonetheless, the separation between the two authorial sets is almost complete, with just two of the Diderot sections located in the outer fringes of the Raynal set. Since the word variables underlying this plot were the 100 most frequent words of the whole text, this is a convincing stylistic division, one that suggests a strong distinction in terms of authorship signal between the two sets.

In order to account for the size discrepancy between the two corpora, we ran another PCA test but this time we increased the number of Diderot sections by segmenting his text into 5,000 word blocks and running these against the previous Raynal 10,000-word sections. This plot is shown below:


Principal component analysis on 5,000 word blocks (Diderot) and Raynal, using 100 most frequent words

Here we see the same sort of authorial/stylistic separation as we saw above, but this time (with the Diderot sections halved in size) the distinction is even stronger, as there is only one section located within the Raynal set of entries, indicating an even greater likelihood that the sections marked in pencil were written by a different author than the rest of the 1780 edition.

These are obviously very rudimentary experiments, but they nonetheless indicate several promising future avenues of exploration. Moving forward, we intend to apply a full suite of computational and stylistic approaches to the 1780 edition and its predecessors, including sequence alignment tools developed by ARTFL, text collation software, and the MEDITE system developed by the labex OBVIL at the Sorbonne for computational genetic criticism. All of these approaches will allow us to explore the textual evolution of the Histoire from 1770 to 1780 in an unprecedented manner, as well as its relationship to other Enlightenment texts and text collections such as Electronic Enlightenment, TOUT Voltaire, and the Encyclopédie.

*I would especially like to thank Alexis Antonia and the Centre for Literary and Linguistic Computing at Newcastle for their generous help with the above stylistic analyses.

[1] See Michèle Duchet, Diderot et l’Histoire des deux Indes ou l’écriture fragmentaire, Paris, Nizet, 1978, p. 22.

[2] Herbert Dieckmann, Inventaire du fonds Vandeul et inédits de Diderot, Genève, Droz, 1951.

ViTA: Visualization for Text Alignment

A preliminary outcome of our Commonplace Cultures Digging into Data project, we have developed a web-based visual analytics system called ViTA: Visualization for Text Alignment. Hosted by the Oxford e-Research Centre at the University of Oxford, ViTA is a web-based visual analytics interface that enables domain experts to construct a text alignment pipeline, visualize the components and connections for any given method (i.e., an alignment model) using image processing techniques, and then test assumptions about the corresponding inputs and outputs. Rather than visualizing the alignment results in a post hoc manner – as is often the case with many available alignment packages – ViTA’s interactive pipeline editing facility essentially becomes a visual programming interface from which users can iteratively build and export more efficient text alignment methods.

ViTA Editor panel

ViTA Editor panel

Screen shot of a ViTA text alignment

Screen shot of a ViTA text alignment

We are hoping to use the ViTA interface to refine our existing PhiloLine-PAIR alignment algorithms, with the goal of identifying ‘commonplaces’ and other forms of large-scale text reuse in the Gale-Cengage Eighteenth Century Collections Online (ECCO) database. A classic ‘big data’ humanities collection, ECCO currently contains more than 32 million digitized pages from 182,898 titles in 205,639 volumes.

Digging into Data

I am very pleased to be one of the co-investigators for a winning project in the third round of the Digging into Data Challenge, an international grant scheme that brings together teams working in computer science and the humanities in the US, Canada, UK, and Netherlands. Our project, “Commonplace Cultures: Mining Shared Passages in the 18th Century using Sequence Alignment and Visual Analytics”, aims to explore 18th-century literary culture through the lens of the early modern practice of commonplacing. Leveraging previous work on data mining and automatic classification of Enlightenment texts (link), machine learning approaches to textual borrowings and source criticism in the 18th century (link), sequence alignment techniques for identifying intertextuality (link) and citation practices in the Encyclopédie (link), we plan to use these same approaches to examine commonplaces and to visualise their deployment over the largest collection of 18th-century works ever assembled.

This project is a partnership between the ARTFL Project and Computation Institute (CI) at the University of Chicago and the University of Oxford’s e-Research Centre (OeRC) and Voltaire Foundation (VF). Bringing together world-class centres for Enlightenment studies (ARTFL, VF) and multi-disciplinary computing applications (CI, OeRC), the team consists of 18th-century scholars: Robert Morrissey (PI, Chicago) and Nicholas Cronk (Co-I, Oxford); computer scientists: Min Chen (PI, Oxford) and Ian Foster (Co-I, Chicago); and digital humanists: Mark Olsen (Chicago), and me (ANU), among other participants.

See the new Project Website for more updates.

TOUT Voltaire…

09The Voltaire Foundation, in collaboration with the ARTFL Project, is pleased to announce the public release of the TOUT VOLTAIRE online database. This database brings you in fully searchable form all of Voltaire’s works apart from his correspondence (which can be searched separately, in Electronic Enlightenment).

Currently publishing the Complete works of Voltaire in print, the Voltaire Foundation plans to unveil an online version of this definitive critical edition sometime after 2018. In the meantime, this plain text version of Voltaire’s writings (without critical apparatus or notes) is the most reliable version available anywhere on the web.

The various editions used to establish this database are clearly marked: from the Voltaire Foundation’s own Complete works of Voltaire to nineteenth-century editions by Beuchot and Moland, among others. When possible we have included Voltaire’s notes, as well as some textual variants depending on the edition. Pagination, however, is often not representative of the print editions, so if you wish to cite Voltaire for scholarly purposes, you should always consult the list of the best critical editions currently available.

The TOUT VOLTAIRE database is built using ARTFL’s full-text search and retrieval engine PhiloLogic, one of the oldest and most successful text analysis systems in the digital humanities. With a wide variety of search and reporting functions, users can look for words, groups of words, or phrases over Voltaire’s entire corpus, or in individual works (and even parts of works). Results can be displayed in context, as frequency reports (by title, by decade, etc.), or as a collocation table and word cloud.

Example searches could include:

For more search tips, please visit the PhiloLogic user manual.

This research tool is made available free of charge by the Voltaire Foundation (University of Oxford) and the ARTFL Project (University of Chicago). If you wish to make a contribution to our work, please contact the Voltaire Foundation.