Research
Research focus areas, current projects, open datasets, open tools, open corpora released by the lab
Research focus areas
The lab organises its research under six thematic strands:
Contact through biblical translation
Investigation of language-contact effects in the Greek Septuagint and biblical retranslations. Analysis of Hebrew-Greek syntactic interference, argument-structure borrowing, and the evolution of translation norms across successive retranslations from antiquity through medieval periods.
Medieval translation contact
Study of contact phenomena in medieval literary and religious translations. Examination of Latin-Greek, Arabic-Greek, and other medieval translation traditions, focusing on how translation practice created sustained contact situations and enabled structural borrowing.
Written-language borrowing
Analysis of syntactic and lexical borrowing specifically in written registers. Investigation of how written contact differs from spoken contact in terms of borrowing constraints, the role of literacy and prestige in enabling structural transfer, and the persistence of written borrowings across generations.
Diachronic retranslation studies
Comparative analysis of successive retranslations of the same source texts across centuries. Study of how retranslation norms evolve, the accumulation of contact effects through multiple translation cycles, and the role of influential earlier translations in shaping later versions.
Corpus-based contact analysis
Development of computational methods for detecting and quantifying contact phenomena in historical treebanks. Application of statistical techniques to distinguish translation interference from contact-induced change and to identify genuine structural borrowings in translated texts.
Syntactic transfer mechanisms
Theoretical investigation of how syntactic structures transfer between languages in written-contact situations. Analysis of constraints on borrowability in written versus spoken contact, the role of argument structure in enabling transfer, and grammaticalisation through contact in literary texts.
Current research projects
Five research projects are active in the lab at v0.4:
CVL-CDSAML, A Corpus-based Valency Lexicon for a Contrastive and Diachronic Study of Ancient and Medieval Languages
The flagship project. Develops a corpus-based valency (the number and type of arguments a verb takes) lexicon for systematic comparison of syntactic patterns across ancient and medieval language varieties. Funded by HFRI Project No. 20577 + Greece 2.0 NRRP. Compute on GRNET ARIS, allocation pa260305. Deliverables include the AthDGC platform, the diachronic-Greek treebank with IE parallels, and the valency-frame database (v0.5).
Written Language Contact Corpus
Development of annotated corpora of biblical and medieval translations for systematic analysis of contact effects. Extends the PROIEL Treebank with post-classical and medieval Greek texts, providing resources for diachronic contact-linguistics research. The corpus includes successive retranslations, enabling comparison of contact phenomena across translation traditions.
Septuagint Contact Phenomena
Comprehensive study of Hebrew-Greek contact effects in the Septuagint, analysing argument-structure patterns, word-order interference, and the development of syntactic constructions under contact influence. Examines how biblical translation created a sustained contact situation that influenced later Greek syntax.
Renaissance Retranslation Patterns
Analysis of Early Modern English and Greek retranslations. Investigates how Renaissance translation practices differed from medieval approaches and examines the accumulation of contact effects through successive translation cycles. Focus on Tyndale's New Testament and contemporary Greek biblical translations.
Contrastive Written Contact
Cross-linguistic comparison of written-contact phenomena in multiple language pairs, examining whether contact effects in translation show universal patterns or language-specific constraints. Integration with the CVL-CDSAML valency lexicon enables systematic contrastive analysis.
AthDGC platform
The lab's flagship computational platform. PROIEL (the dependency-treebank (a collection of sentences whose grammatical structure has been analysed and stored) standard for early Indo-European languages, developed at Oslo)-XML 2.0 dependency-parsed treebank of the entire Greek language (Homeric through Modern), with verse-level cross-lingual alignment to four IE witnesses at v0.4 and five more queued at v0.7.
| Item | Status | URL |
|---|---|---|
| Public showcase | live | https://athdgc.github.io |
| Source repository | live | https://github.com/AthDGC/Diachronic-Linguistics-Platform |
| Hugging Face (a public hosting platform for machine-learning models and datasets) mirror | live (3 model repos) | https://huggingface.co/AthDGC |
| PyPI (the Python Package Index, the central archive that pip downloads packages from) package | live (stub) | https://pypi.org/project/athdgc-tools/ |
| Concept DOI (a Digital Object Identifier that always points to the latest version of a deposited record) | live (v0.4.0) | 10.5281/zenodo (an open research-data repository hosted at CERN that mints permanent DOIs).20439182 |
Open-source toolkit
Fourteen modules under OSI-approved licences. Highlights:
- LightSIDE (an open-source text-mining workbench developed at Carnegie Mellon)-AthDGC (the Lavidas-extension fork of LightSIDE that operates on syntactic features rather than only text features), LightSIDE fork for PROIEL syntactic features (dependency arcs (the head-to-dependent links between words in a parsed sentence), argument-structure frames, morphology bundles). BSD-3-Clause + Apache-2.0 (a permissive open-source licence widely used for software).
- Fine-tuned (further trained on new data to adapt it to a particular text type) Stanza (Stanford's open-source Python workflow that automatically tags, lemmatises, and parses sentences) checkpoints (saved snapshots of a trained model),
grc_byz_proiel,grc_lbem_proiel,grc_mod_proielfor diachronic Greek. Apache-2.0. Hosted at https://huggingface.co/AthDGC. - PROIEL XML 2.0 (the file format that stores each sentence as a tree of word-by-word grammatical relations, developed at Oslo) validator (v0.5), schema + relation-inventory linter. Apache-2.0.
- House-style check (v0.5), automated style enforcement for repository PRs. Apache-2.0.
- Quarto (an open-source publishing system that builds websites, slides, papers, and posters from a single source) template pack, the multi-output Quarto pack that builds athdgc.github.io and this lab site. MIT (a short permissive open-source licence).
Full module list: https://athdgc.github.io/tools.html.
Open-access corpus inputs
Every primary source text used by the lab is open-access (public domain, CC-BY, CC-BY-SA, or equivalent). Greek sources draw on Perseus Digital Library, Open Greek and Latin / First1K (Leipzig), SBL Greek NT, Tischendorf and Westcott-Hort, Rahlfs LXX via openscriptures.org, Papyri.info, Patrologia Graeca via Documenta Catholica Omnia, Bibliotheca Augustana, Anemi (UoC), and Wikisource el. IE parallels draw on Vulsearch + Latin Library, the Wulfila Project (University of Antwerp), TITUS (Frankfurt), Digilib Armenian, GRETIL (Goettingen), SARIT, TEAMS, the DOE corpus, and the National Library of Ukraine. Full per-period source map: https://athdgc.github.io/samples.html.
Working Papers
Open-access pre-prints + platform launch reports self-published as the GlossaContactLab Working Papers, digital edition.
Funding
Funded by HFRI Project No. 20577 + Greece 2.0 National Recovery and Resilience Plan. Compute on GRNET ARIS (the Greek national high-performance computing cluster, run by GRNET). Project: CVL-CDSAML.