My research focuses on deep learning for natural language processing (NLP). I am interested in systems that use data in multiple languages and in how their quality can be evaluated.
Highlights from my research:
- PhD thesis on Model-Based Evaluation of Multilinguality: Summary, Published thesis
- SwissBERT, the multilingual language model for Switzerland: Explainer, Model weights
Questions that intrigue me:
- How can maximum-quality text be generated with language models?
- How can large language models be adapted to local data?
- Since language models estimate word probabilities – what are creative ways of using those probabilities? Ideas we developed previously include: Contrastive Conditioning, Translation Cross-Likelihood, Omission Error Detection and Translation Direction Detection
Thesis Supervision
The next semester in which I can supervise new Bachelor's or Master's theses is Spring 2027. Please reach out via email, and please provide a list of 2–3 topic ideas. This will help me understand your research interests and give you optimal advice.
Short CV
- Since January 2024: Academic associate at the Department of Computational Linguistics
- Collaborator InvestigaDiff project
- Collaborator Language Technology for Romansh Idioms
- April 2023 − December 2023: Postdoctoral researcher, MUTAMUR project
- 2019 − March 2023: PhD student at the Department of Computational Linguistics, supervised by Rico Sennrich, Lena A. Jäger and Martin Volk.
- Summer 2022: Applied Science Internship with Amazon AI Translate, Berlin
- 2018−2019: Research internship at Munich Re (NLP for Reinsurance Development)
- 2018−2019: Graduate teaching assistant for Prof. Dr. Hinrich Schütze, CIS Munich
- 2017−2019: M.Sc. in Computational Linguistics (major) and Computer Science (minor) at LMU Munich
- 2015−2017: Full-Stack Web Developer at Arteria GmbH, Basel
- 2011−2015: B.A. in Computer Science and Philosophy from the University of Basel
Recent Publications
Jannis Vamvas, Ignacio Pérez Prat, Angela Heldstab, Dominic P. Fischer, Sina Ahmadi and Rico Sennrich. 2026. Translation Asymmetry in LLMs as a Data Augmentation Factor: A Case Study for 6 Romansh Language Varieties. Accepted to Findings of EMNLP 2026. [cite] [data] [model] [code]
Michelle Wastl, Jannis Vamvas and Rico Sennrich. 2026. Scaling Unsupervised Word Alignment to Documents via Structural Constraints. Accepted to EMNLP 2026. [cite] [data] [code, code]
Hanxu Hu, Zdeněk Šnajdr, Pinzhen Chen, Jannis Vamvas and Rico Sennrich. 2026. Reinforcement Learning Elicits Contextual Learning of Unseen Language Translation. Accepted to Findings of EMNLP 2026. [cite] [code]
Michelle Wastl, Jannis Vamvas and Rico Sennrich. 2026. SwissGov-RSD: A Human-annotated, Cross-lingual Benchmark for Token-level Recognition of Semantic Differences Between Related Documents. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 31134–31163, San Diego, California, United States. Association for Computational Linguistics. [cite] [data] [code] ★ Selected as SAC Highlight
Apertus Team. 2026. Apertus: Democratizing Open and Compliant LLMs for Global Language Environments. In Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 46877–46955, San Diego, California, United States. Association for Computational Linguistics. [cite] [model]
Recent Teaching
| Fall 2026 | Lecturer Fundamentals of Large Language Models |
| Fall 2026 | Lecturer Mathematical Foundations for Language Technology 1 |
| Fall 2026 | Co-instructor Ethics of AI for Language and Speech |
| Spring 2026 | Lecturer CAS Generative AI |
| Spring 2026 | Lecturer Text Generation with Language Models |