From Philology to Semantic Triples: Displaying Ancient Inscriptions with Digital Tools

Translate this post

This summer, I had the pleasure of serving as an intern for the IDEA Project, an initiative that documents information related to the ancient Syrian city of Dura-Europos. My job included transcribing and digitally representing a series of inscriptions from the Palace of the Dux Ripae, a military building created during the city’s Roman occupation in the third century. The inscriptions cover a wide range of functions: some were painted requests for the master of the house to be remembered, while others were mundane lists or messages scribbled on the walls by ancient passers-by. The inscriptions I worked on also vary in language, from Greek to Latin to Persian, reflecting the multilingual culture of the frontier city. 

The remains of the Palace of the Dux Ripae, set against a clear blue sky (Marsyas, CC BY-SA 3.0 https://creativecommons.org/licenses/by-sa/3.0, via Wikimedia Commons)

But despite the abundance of riches uncovered in this palace full of inscriptions, the contents of the Palace of the Dux Ripae were not fully accessible when I started my internship. The individual inscriptions were discussed in 20th-century field reports, but most had left very few traces elsewhere. Therefore, working on the inscriptions required, at minimum, a strong understanding of English, knowledge of ancient languages and epigraphy, the training to parse archaeological records that are almost 100 years old, the ability to see and interpret images that lack alternative text for a screen reader, and access to library credentials. It is these barriers to entry that motivate interventions by the IDEA Project. By extracting the inscriptions and their associated contextual details from an exclusive, monolingual format and weaving them into the Wikidata knowledge graph, I was able to represent them in a medium that makes them easy to find, highly accessible, and available to researchers in a more inclusive range of languages.

Immersing myself in Wikidata, I found that one of my most gratifying summer experiences was expanding my understanding of the uses of language. Although I’ve spent my academic career studying language, working with the IDEA Project required me to read, communicate, and process information in an entirely new way. Learning the language of Wikidata is a task different from the traditional activities of Classicists, but using computer-oriented languages creates potential for accessibility, collaboration, and innovation in the representation of information about the Roman Empire.

I came to the IDEA project with the traditional training of a literary Classicist, a training that is language-centered (or, to use the academic word, philological). I’ve racked up hours studying vocabulary, committing verb forms to memory, and puzzling over the intricate syntax of Greek and Latin. When I need to communicate the information I’ve received and processed, I often use the scholarly language of academic essays and presentations.

It was certainly true that this philological training benefitted me as I transcribed the inscriptions. For instance, my knowledge of ancient languages helped me to understand the linguistic ambiguities presented by fragmentary inscriptions. Even the erosion of one letter can seriously impact an inscription’s message, so researchers who aim to reconstruct the inscriptions’ text must sometimes keep more than one option in mind; for instance, an eroded line in one inscription I worked on could refer to either tyros (cheese) or Tyria (a type of currency). When I read discussions of textual reconstructions, my knowledge of Greek and Latin helped me appreciate the different linguistic possibilities. Since my philological work with the ancient languages is often grounded in works of polished, poetic literature, reading colloquial messages both expanded my understanding of the languages’ possibilities and brought a smile to my face: scratching the walls of the palace, people left behind references to members of the army, messages to Zeus, and sometimes just random scribbles or symbols.

But as much as my philological training assisted me, working with the IDEA project also required me to acquire entirely new language skills. To turn my notes on the inscriptions into information that Wikidata could understand and display, I had to work with a concept called a semantic triple. Semantic triples contain three parts that make information legible to a computer: a subject, a predicate, and an object. (They can also contain qualifiers, which further describe one of the elements.) Those who have studied grammar may recognize those terms, but their application in the context of the Semantic Web requires a bit of rethinking literary conventions.

Conventionally, when one sets out to write, goals might include communication that sounds natural, a feeling of clarity and precision, and variation in sentence length and word choice to make the prose interesting. If I were describing one of the Dux Ripae inscriptions in essayistic language, I might say something like “Inscription 960 is fragmentary graffiti that is an archaeological artefact.”

But communicating this information to Wikidata required me to think of communication differently. The sentence I wrote in natural language contains too much information for a semantic triple, which is composed of three simple parts. I have to break it up into a few semantic triples:

Inscription 960: instance of: inscription (qualifier: object of statement has role: fragment).

Inscription 960: instance of: graffiti (qualifier: object of statement has role: fragment).

Inscription 960: instance of: archaeological artefact.

A screenshot from Wikidata displaying “instance of” statements for Inscription 959 from Dux Ripae, Dura-Europos.

And I render information in a similar way for each piece of information about the inscription: its date, location, language, material, dimensions, text, and visual representaions, among other elements.

Although displaying information in such a way doesn’t involve smooth, flowing prose, using semantic triples has many advantages. One of the most important is language accessibility. Each part of a semantic triple derives from a label that exists independently on Wikidata. Volunteers on Wikidata translate labels into different languages. Since, for instance, the labels “instance of,” “inscription,” “graffiti,” and “archaeological artefact” have all independently been translated by human contributors into Spanish, the semantic triples involving these elements can automatically be displayed in Spanish. In sum, each bit of information previously confined to the English archaeological reports can now be rendered in any of the languages supported in the Wikidata environment. In addition, Wikidata’s text and images can be parsed with a screen reader, making information accessible to visually impaired users.

A screenshot from Wikidata displaying “instance of” statements for Inscription 959 from Dux Ripae, Dura-Europos in Spanish.

Another advantage of uploading an item to Wikidata is the context and connection that representation in the Knowledge Graph provides. Part of this data comes from the IDEA Project’s focus on consolidating the fragmented landscape of information related to the 20th-century excavations. When I uploaded an inscription, I included data about the physical item and its ancient culture as well as its country of origin (Syria) and time of discovery (1935-1936); I also noted that it was described by a specific field report. When viewed in museums, ancient objects are often stripped of their modern context. But a collection object’s web of Wikidata relationships can give visibility to the contexts that have colored the interpretation of objects on view in modern institutions.

What’s more, uploading items to Wikidata and fleshing them out with data connects the items in a way that no physical museum could. Using Wikidata’s coding language, SPARQL, it’s possible to search Wikidata for all items that match precisely defined criteria: for instance, every item that contains the statement “instance of: inscription” and “location of discovery: Palace of the Dux Ripae.” Wikidata can provide a list of every item that fits these categories, enabling users to see that these inscriptions are not isolated finds but connected to other items (which are in turn connected to archival documentation, publications, ruinous building remains, and so on).

A screenshot of a screen bubble visualization of entities related to the Palace of the Dux Ripae and its inscriptions.

Some Classicists in this day and age are defensive of philology, arguing that non-traditional activities undermine the discipline of Classics. These Classicists tend to be wary of modern methodology, especially that which concerns race, identity, or marginal areas of study. As a philologist who adores studying language and literature, I understand when people express a strong appreciation for textual study. But in my opinion, a narrow-minded focus on philology to the detriment of other ways of participation in Classics obscures the exciting possibilities of ancient studies in the modern day. Rather than threaten each other, the study of ancient and modern ways of communication – of natural and computer-oriented languages – complement each other, allowing for simultaneous flourishing. One does not gain while the other loses.

Many digital initiatives are enriching the field of Classics, from the Perseus Digital Library to Logeion, and, perhaps most exciting of all, the Vesuvius Challenge, which aims to digitally unwrap carbonized papyri to reveal lost works of literature and philosophy. These projects, which increase Classics’ visibility and accessibility, are not only meaningful but thrilling: we get to see our discipline evolve with the digital age, leading to new discoveries and insights as a result. If we wish to contribute to such projects, it’s crucial that, in addition to gaining a thorough training in ancient studies, we also expand our understanding of what Classicists are capable of learning and doing. To be clear, I’m not suggesting a sort of Classics utilitarianism, where the study of the ancient world must integrate digital tools to prove itself to be practical or worthwhile. Rather than drive home existential anxiety or the pressure for reform, I simply wish to express a spirit of optimism: even as certain technological developments such as generative AI may unsettle us, I believe that appreciators of the ancient world have much to look forward to in the digital age. At the very least, having learned to communicate with Wikidata, I believe there is much pleasure and fulfillment to be had in allowing oneself to learn new languages, in whatever format they manifest, again and again.

Can you help us translate this article?

In order for this article to reach as many people as possible we would like your help. Can you translate this article to get the message out?