PhD Position F/M Robust and generalizable sign-to-text translation for French and German sign languages
Contract type : Fixed-term contract
Level of qualifications required : Graduate degree or equivalent
Fonction : PhD Position
Context
The PhD project will be carried out in the MULTISPEECH team at Loria. The PhD candidate will be supervised by Slim Ouni (professor, University of Lorraine) and Mostafa Sadeghi (researcher, Inria), and will benefit from the research environment, expertise, and powerful computational resources (GPUs & CPUs) of the team.
This PhD position is fully funded within the framework of RoGSiLT: Robust and Generalizable Sign Language Translation, a collaborative research project between Inria in France and DFKI in Germany. The project aims to advance AI-based translation technologies for French Sign Language (LSF) and German Sign Language (DGS), covering both translation from sign-language videos into written language and translation from written language into sign language. The PhD candidate will work with the Inria team and collaborate closely with DFKI, with a specific focus on the data, alignment, and representation-learning foundations required for robust and generalizable sign-to-text translation.
Assignment
Motivation and context
Sign languages are natural languages with their own grammar, structure, and rich visual-expressive features. Unlike spoken languages, they convey meaning through a combination of hand movements, body posture, facial expressions, mouthing, spatial organization, and timing. Automatic translation from sign language videos into written text therefore requires modelling complex visual and temporal information, rather than simply recognizing isolated gestures.
A common approach in sign language translation relies on glosses, i.e., written labels that roughly represent signs using words from a spoken language. Glosses are useful as an intermediate representation, but they are costly to annotate, scarce for many sign languages, and unable to capture the full linguistic richness of signed communication, such as facial grammar, spatial references, or prosody. As a result, existing systems often depend on limited gloss-annotated data, generalize poorly across signers and recording conditions, and produce translations that lack fluency, completeness, or semantic accuracy [1].
Main activities
Project description
The PhD project will contribute to the development of robust and generalizable sign-to-text translation for low-resource sign languages, with a focus on LSF and DGS. Rather than concentrating only on the final translation model, the PhD will address the upstream components that are essential for reliable sign-to-text translation: data collection and curation, privacy-aware data processing, automatic video-text alignment, and sign-language representation learning.
A first objective will be to identify, collect, structure, and curate sign-language video resources for LSF and DGS, for which available resources remain more limited than for some other sign languages and domains [2,3]. The work will exploit existing and newly identified video-text resources, including open repositories and potential data sources from associations or broadcasters. Since sign-language data typically consists of videos of identifiable signers, the PhD will also address privacy and ethical challenges, including consent, anonymization, secure data handling, metadata quality, and linguistic diversity.
A second objective will be to develop automatic methods for aligning sign-language videos with corresponding written transcriptions, subtitles, or other textual material. Such alignment is essential for transforming raw or weakly structured video-text data into resources that can be used for machine learning. The PhD will therefore investigate multimodal approaches that connect visual sign-language information with textual representations under low-resource conditions, using techniques from computer vision, natural language processing, and representation learning [4,5].
A third objective will be to study how sign-language videos can be represented in a way that captures the visual and temporal complexity of signed communication, including manual signs, body posture, facial expression, mouthing, spatial structure, and timing. The PhD will explore how sign-language-specific pre-trained and multimodal models can be adapted to LSF and DGS data, in order to support video-text alignment, representation learning, and downstream sign-to-text translation [9,10].
Implementation plan
The first stage of the PhD will focus on data collection and curation. The candidate will identify and gather relevant LSF and DGS video-text resources from open repositories and other potential sources. This will include the analysis of available metadata, text-video correspondence, recording conditions, signer variation, and data quality. Particular attention will be paid to privacy-aware processing, including anonymization strategies, secure data handling, and the possible use of pose-based or derived representations to reduce identifiability.
The second stage will focus on automatic alignment between sign-language videos and written text. The candidate will investigate methods for aligning videos with subtitles, transcriptions, or other textual material, to transform weakly structured data into usable training and evaluation resources. This work may involve temporal segmentation, visual feature extraction, sentence-level semantic representations, multimodal similarity learning, and Transformer-based or CLIP-like architectures adapted to sign-language data [4,5].
The third stage will explore the use of pre-trained and foundation models for sign-language representation learning. The candidate will study how computer vision models, multimodal models, and large language models can be adapted or fine-tuned for sign-language data. This may include keypoint-based, video-based, or hybrid representations, as well as weakly supervised or contrastive learning strategies for connecting visual and textual information.
In the final stage, the developed data-processing, alignment, and representation-learning components will be assessed in relation to downstream sign-to-text translation. The work will include quantitative and qualitative analysis of alignment quality, robustness across data sources and signers, and the usefulness of the learned representations for translation-oriented tasks. While the PhD is not primarily centred on evaluation methodology, the candidate will use appropriate automatic and human-informed analyses to validate the proposed methods [6].
Expected outcomes
The PhD candidate is expected to contribute to both scientific knowledge and practical resources for the sign language technology community. Expected outcomes may include:
- Curated or enriched LSF and DGS video-text resources;
- Improved automatic alignment methods for sign videos and written text;
- Multimodal representations of sign-language videos suitable for downstream translation;
- Adaptation of pre-trained computer vision, multimodal, and language models to sign-language data;
- Open-source implementations of data-processing, alignment, and modelling components;
- Experimental analyses demonstrating the usefulness of the proposed methods for sign-to-text translation;
- Publications in leading conferences and workshops in NLP, computer vision, or sign language technology.
References
[1] N. C. Camgöz, S. Hadfield, O. Koller, H. Ney, and R. Bowden, “Neural Sign Language Translation,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
[2] J. Forster, C. Schmidt, O. Koller, M. Bellgardt, and H. Ney, “Extensions of the Sign Language Recognition and Translation Corpus RWTH-PHOENIX-Weather,” in Proceedings of the International Conference on Language Resources and Evaluation (LREC), 2014.
[3] J. Halbout, D. Fabre, Y. Ouakrim, J. Lascar, A. Braffort, et al., “Matignon-LSF: A Large Corpus of Interpreted French Sign Language,” in Proceedings of the LREC-COLING Workshop on the Representation and Processing of Sign Languages, 2024.
[4] H. Bull, T. Afouras, G. Varol, S. Albanie, L. Momeni, and A. Zisserman, “Aligning Subtitles in Sign Language Videos,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021.
[5] Y. Hamidullah, J. van Genabith, and C. España-Bonet, “Sign Language Translation with Sentence Embedding Supervision,” in Proceedings of the Annual Meeting of the Association for Computational Linguistics (ACL), Short Papers, 2024.
[6] M. Müller et al., “Findings of the Second WMT Shared Task on Sign Language Translation,” in Proceedings of the Conference on Machine Translation (WMT), 2023.
[7] G. Fauré, M. Sadeghi, S. Bigeard, and S. Ouni, “Towards Skeletal and Signer Noise Reduction in Sign Language Production via Quaternion-Based Pose Encoding and Contrastive Learning,” in Proceedings of SLTAT 2025: 9th Workshop on Sign Language Translation and Avatar Technologies, 2025.
[8] G. Fauré, M. Sadeghi, S. Bigeard, and S. Ouni, “The Impact of VAE Design on Latent Pose Representations for Diffusion-based Sign Language Production,” in Proceedings of GenSign: Generative AI for Sign Language, CVPR 2026 Workshop, 2026.
[9] Z. Jiang, G. Sant, A. Moryossef, M. Müller, R. Sennrich, and S. Ebling, “SignCLIP: Connecting Text and Sign Language by Contrastive Learning,” in Proceedings of the Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 9171–9193, 2024.
[10] P. Jiao, Y. Min, and X. Chen, “Visual Alignment Pre-training for Sign Language Translation,” in Proceedings of the European Conference on Computer Vision (ECCV), pp. 349–367, 2024.
Skills
Applicants should have, or be close to obtaining, a Master’s degree or equivalent in computer science, artificial intelligence, computational linguistics, natural language processing, computer vision, machine learning, or a related field.
A strong candidate will have experience or interest in several of the following areas:
- Machine learning and deep learning;
- Natural language processing or neural machine translation;
- Computer vision or video understanding;
- Multimodal learning;
- Large language models.
Programming experience in Python and familiarity with deep learning frameworks such as PyTorch are expected. Knowledge of sign languages, LSF, DGS, or Deaf studies is welcome but not mandatory. The candidate should be willing to work in an interdisciplinary and collaborative setting.
Benefits package
- Subsidized meals
- Partial reimbursement of public transport costs
- Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
- Possibility of teleworking (after 6 months of employment) and flexible organization of working hours
- Professional equipment available (videoconferencing, loan of computer equipment, etc.)
- Social, cultural and sports events and activities
- Access to vocational training
- Social security coverage
Remuneration
2300€ gross/month
General Information
- Theme/Domain :
Language, Speech and Audio
Statistics (Big data) (BAP E) - Town/city : Villers lès Nancy
- Inria Center : Centre Inria de l'Université de Lorraine
- Starting date : 2026-11-01
- Duration of contract : 3 years
- Deadline to apply : 2026-08-15
Warning : you must enter your e-mail address in order to save your application to Inria. Applications must be submitted online on the Inria website. Processing of applications sent from other channels is not guaranteed.
Instruction to apply
Defence Security :
This position is likely to be situated in a restricted area (ZRR), as defined in Decree No. 2011-1425 relating to the protection of national scientific and technical potential (PPST).Authorisation to enter an area is granted by the director of the unit, following a favourable Ministerial decision, as defined in the decree of 3 July 2012 relating to the PPST. An unfavourable Ministerial decision in respect of a position situated in a ZRR would result in the cancellation of the appointment.
Recruitment Policy :
As part of its diversity policy, all Inria positions are accessible to people with disabilities.
Contacts
- Inria Team : MULTISPEECH
-
PhD Supervisor :
Sadeghi Mostafa / mostafa.sadeghi@inria.fr
The keys to success
Required documents
- Curriculum vitae (CV);
- Cover letter describing research interests and motivation for the topic;
- Academic transcripts;
- Names and contact details of referees;
- (Optional) Publications, code repositories, Master’s thesis, or other relevant research outputs.
About Inria
Inria, the French national institute for research in digital science and technology, supports the French government in national research and innovation strategies in the digital field, acting as Digital Programs Agency. Inria leads over 300 research and innovation projects with its 3,500 scientists, engineers, and support staff, in partnership with universities and the digital ecosystem (businesses, entrepreneurs, and public stakeholders). Together, we explore strategic fields such as artificial intelligence, cybersecurity, quantum computing, cloud technologies, digital transformation in healthcare, digital twins, and digital technologies for defence. We develop practical solutions such as software, tech startups, partnerships with national companies, and cutting-edge training programmes. Our goal is to drive scientific, technological, and industrial excellence to ensure France’s digital sovereignty.