About the position
Context This position is part of the ANR project ROMAM², led by Thibault Clérice at Inria Paris within the ALMAnaCH project-team (Automatic Language Modelling and Analysis & Computational Humanities, led by Benoît Sagot). ROMAM² treats pre-editorial normalisation (PEN) of graphemic automatic text recognition (ATR) output as a dedicated NLP task. PEN is a traceable process in which every editorial inference (abbreviation expansion, post-correction of recognition errors, spelling normalisation) stays anchored to the manuscript it comes from. The project works on two medieval languages: Latin, which is heavily abbreviated, and Old French, whose spelling varies in linguistically meaningful ways. A central claim of the project is that editorial normalisation is not neutral. Printed critical editions silently expand abbreviations and regularise spelling. In doing so, they erase variation that is evidence for the history of the language. Dees' quantitative geography of Old French and its successors rest largely on such editions, and Morin has shown how this can distort dialectal conclusions. However, no study has yet measured this distortion on a controlled parallel corpus. This postdoctoral position is designed to produce that study and the gold data it requires. The postdoctoral researcher will be supervised by Thibault Clérice. They will work closely with: the project's PhD candidate in NLP, co-supervised by Thibault Clérice, Benoît Sagot and Rachel Bawden. The PhD candidate will use the gold data and evaluation framework produced by the postdoc to train and evaluate normalisation models. David Smith (Northeastern University), a specialist in aligning noisy historical data. the ANR JCJC project Phil•IA, coordinated by Ariane Pinche at CIHAM (UMR 5648, ENS de Lyon). A regular collaboration is expected on Old French graphemic transcription, digital editing and TEI encoding. The position is based at Inria Paris, within ALMAnaCH. The team brings together researchers in NLP, language modelling and computational humanities, and offers a rare environment for philologists and linguists who want to work directly with NLP researcher. The postdoc will also benefit from existing community resources developed by the team: the CATMuS dataset (the largest ATR dataset for medieval manuscripts), the CoMMA corpus (3.3 billion tokens of Latin and Old French from over 32,000 manuscripts) and the upcoming work on Biblissima-Textes. The position is for 18 months, starting March 2027. Bibliography Clérice, T., Bawden, R., Glaise, A., Pinche, A., & Smith, D. (2026). Pre-Editorial Normalization for Automatically Transcribed Medieval Manuscripts in Old French and Latin. In Proceedings of the Fourth Workshop on Language Technologies for Historical and Ancient Languages (LT4HALA) @ LREC 2026 . https://arxiv.org/abs/2602.13905 Clérice, T., Pinche, A., Vlachou-Efstathiou, M., Chagué, A., Camps, J.-B., et al. (2024). CATMuS Medieval: A multilingual large-scale cross-century dataset in Latin script for handwritten text recognition and beyond. In Proceedings of ICDAR 2024 (LNCS 14806, pp. 174–194). Springer. https://doi.org/10.1007/978-3-031-70543-4_11 Clérice, T., Gabay, S., Vlachou-Efstathiou, M., Pinche, A., & Sagot, B. (2026). CoMMA, a Large-scale Corpus of Multilingual Medieval Archives. In Proceedings of the Fifteenth Language Resources and Evaluation Conference . ELRA. https://inria.hal.science/hal-05299220 Dees, A. (1985). Dialectes et scriptae à l'époque de l'ancien français. Revue de Linguistique Romane , 49(193–194), 87–117. Morin, Y. C. (2006). Histoire du corpus d'Amsterdam : le traitement des données dialectales. In Le Nouveau Corpus d'Amsterdam. Actes de l'atelier de Lauterbad . Scheer, T., & Brun-Trigaud, G. (2022). L'atlas Dees électronique. Concordial , Grenoble. https://hal.science/hal-03912660 Kuparinen, O., & Scherrer, Y. (2024). Corpus-based dialectometry with topic models. Journal of Linguistic Geography , 12(1), 1–12. Pinche, A. (2022). Guide de transcription pour les manuscrits du Xe au XVe siècle . https://hal.archives-ouvertes.fr/hal-03697382 Duval, F. (2012). Transcrire le français médiéval : de l'« Instruction » de Paul Meyer à la description linguistique contemporaine. Bibliothèque de l'École des chartes , 170(2), 321–342. Assignment The postdoctoral researcher, trained in philology or historical linguistics with skills in digital humanities, will build the philological foundations and evaluation resources of ROMAM² and carry out a controlled study of how editorial practices affect the dialectometry of Old French. The work involves: Building a multi-layer gold corpus from the Nouveau Corpus d'Amsterdam (NCA, formerly Dees' corpus). Using Tobias Scheer's Atlas Dees Électronique mapping between NCA editions and their source manuscripts, the postdoc will sample each text (~500 words per document, ~100,000 words in total) and transcribe them in eScriptorium, align the edited text with a graphemic transcription of the manuscript, and annotate tokens as abbreviated or not in XML-TEI. Abbreviation-aware dialectometry. The postdoc will quantify abbreviation practices across the corpus and compare the resulting feature maps and dialectal distances with those of Dees and his successors (e.g. Scherrer's Dialektkarten ). They will assess how far ambiguous abbreviation resolution, of the kind Morin identified in Floovant , changes the conclusions, using several dialectometric methods. Auditing and augmenting the training data. In collaboration with David Smith, the postdoc will audit the automatically aligned corpus from the prototype PEN work. The goal is to separate valid alignments (identity, ATR post-correction, abbreviation expansion) from invalid ones (literary variants, spelling variants), and to explore inter- and intra-manuscript alignment between witnesses. This work also yields a corpus for studying abbreviation practices across textual traditions. Contributing to the evaluation framework , jointly with the PhD candidate. This includes lossless conversion between ALTO, plain text and TEI that preserves uncertainty markup ( <choice> , <abbr> , <expan> ), stage-specific metrics, and a fine-grained error taxonomy (overnormalisation, variant insertion, hallucination, morphosyntactic errors, ambiguity collapse, etc.). This builds on the expertise of Ariane Pinche and the Phil•IA project. Supervising annotation work carried out by hourly-paid annotators, with the PI, to ensure the linguistic and editorial quality of the gold data (Old French and Latin if the applicant knows Latin). The research component lies at the intersection of historical linguistics, material philology and NLP evaluation. The postdoc is expected to publish results in both communities: a paper on the impact of normalisation on dialectal attribution in Old French, open datasets (the TEI NCA sample and the new gold PEN dataset), and a proposed panel at the International Medieval Congress (Leeds). Main activities The main activities of the applicant will include: carrying out research on the topic outlined in the job descriptions, including developing new ideas, positioning the work with respect to related work in historical linguistics and NLP, and validating the methodology through corpus construction, experiments and analysis producing and releasing open, reusable datasets in XML-TEI using the ParamHTRs interface (TEI NCA sample with abbreviation annotation, gold PEN data for Old French and Latin) working closely with the project's PhD candidate (starting September 2027), so that the gold data and evaluation framework directly support model development collaborating with the ANR Phil•IA project (CIHAM, Lyon) and with the project's external experts (David Smith, Ariane Pinche) supervising and checking annotation work carried out by annotators presenting work both internally and externally in conference, journal and workshop papers, in NLP and humanities venues (e.g. Revue de linguistique romane , IMC Leeds, CHR, LREC) Organizing and taking part in the project's workshops and exchanging with colleagues on NLP and philological topics Skills Required: PhD in philology, historical linguistics, medieval studies or a related field Strong knowledge of Old French Training in palaeography, with experience reading medieval manuscripts Working knowledge of XML-TEI Ability to do statistics and to process or parse structured data with at least one programming language (Python or R) Highly appreciated: Experience in dialectometry, or a demonstrated interest in the dialects and scriptae of medieval French Knowledge of medieval Latin Soft skills: Ability to work in an interdisciplinary team with NLP researchers Good written and oral communication in English; French is an asset Good organization skills
This listing was collected from a public source and is reproduced here for
information only. Always confirm the details on the original posting before applying.
View the original posting