PhD Position F/M Privacy-Preserving Speech Understanding with Multimodal Signals for Clinical Applications

Contract type : Fixed-term contract

Level of qualifications required : Graduate degree or equivalent

Other valued qualifications : Engineering degree or Master’s degree in computer science, artificial intelligence, signal or speech processing, applied mathematics, data science, or a related field.

Fonction : PhD Position

About the research centre or Inria department

L'employeur sera l'Université de Lorraine. 

Context

The PhD student will join the MULTISPEECH team at LORIA, a joint research unit of Université de Lorraine, Inria and CNRS, located in Villers-lès-Nancy, France. The three-year PhD is funded by the AI Grand Est ENACT research chair and is embedded in a recognised research environment in speech processing, deep learning, voice privacy and artificial intelligence for healthcare.

The project offers an interdisciplinary environment through collaboration with CHRU Nancy, bringing together speech AI, emergency medicine, public health and clinical research, data governance, health law and ethics, and biomedical imaging and signals. The student will have access to dedicated computing resources, speech-processing software, deep-learning frameworks and AI engineering support from the ENACT chair.

Clinical data will be processed within an appropriate ethical and regulatory framework, using secure HDS-certified health-data infrastructure. The position offers opportunities for international publications, clinical collaborations and contributions to open evaluation frameworks for privacy-preserving clinical speech.

Assignment

The PhD student will develop and evaluate privacy-preserving methods for clinical spoken language understanding. The work will investigate the trade-off between protecting speaker identity and spoken identifying content, on the one hand, and preserving clinically relevant information, including paralinguistic biomarkers and medical semantics, on the other.

The PhD will build on de-identified clinical speech-and-text corpora, speech and audio large language models (speech-LLMs), and, where feasible, an additional modality such as physiological signals, medical imaging or health-record metadata. The precise clinical use case and available data will be defined in collaboration with clinical partners. The objective is to deliver methods and evaluation protocols that rigorously reconcile voice privacy with clinical utility.

Keywords

Speech anonymization; voice privacy; spoken language understanding; speech processing; speech and audio language models; health data; artificial intelligence for healthcare; multimodality; deep learning; privacy–utility trade-off; privacy evaluation.

References

  • European Parliament and Council. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union, 2024.

  • European Parliament and Council. Regulation (EU) 2016/679 (General Data Protection Regulation, GDPR). Official Journal of the European Union, 2016.

  • Tomashenko, N., et al. “Introducing the VoicePrivacy Initiative.” Proceedings of Interspeech, 2020.

  • Tomashenko, N., et al. “The Third VoicePrivacy Challenge: Preserving Emotional Expressiveness and Linguistic Content in Voice Anonymization.” Computer Speech & Language, 2026.

  • Srivastava, B. M. L., et al. “Privacy and Utility of x-Vector Based Speaker Anonymization.” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2022.

  • Arasteh, S. T., et al. “Addressing Challenges in Speaker Anonymization to Maintain Utility While Ensuring Privacy of Pathological Speech.” Communications Medicine, 2024.

  • Diaz-Asper, C., et al. “Navigating the Tradeoff Between Personal Privacy and Data Utility in Speech Anonymization for Clinical Research.” npj Digital Medicine, 2025.

  • Cohn, I., et al. “Audio De-identification: A New Entity Recognition Task.” Proceedings of NAACL-HLT, 2019.

  • Radford, A., et al. “Robust Speech Recognition via Large-Scale Weak Supervision.” 2022.

  • Zhong, X., et al. “Considerations for Patient Privacy of Large Language Models in Health Care: Scoping Review.” Journal of Medical Internet Research, 2025.

Main activities

  • Select, curate and, where needed, annotate a de-identified clinical speech-and-text corpus in collaboration with clinical partners.

  • Develop baseline methods for voice anonymization and spoken-content de-identification.

  • Design clinical spoken language understanding models operating on anonymized inputs and based on speech and audio large language models.

  • Investigate disentangled representations to remove speaker identity and identifying information while retaining clinically relevant information.

  • Develop evaluation methods that combine privacy and clinical utility, including speaker verification, membership and attribute-inference attacks, task-performance measures and expert assessment.

  • Extend the approaches to an additional modality and develop privacy-preserving defenses, according to the data and clinical use case defined with partners.

  • Test, compare and validate the developed methods; write technical documentation, progress reports and scientific publications.

  • Present and disseminate results to scientific and clinical partners and through international conferences and journals.

 

Skills

 

  • Strong Python skills.

  • Good knowledge of machine-learning and deep-learning methods.

  • Proficiency in PyTorch, PyTorch Lightning and the Transformers library (Hugging Face), or another deep-learning framework.

  • Interest in speech processing and multimodal data, including audio, text and signals.

  • Ability to design, train, evaluate and compare models; knowledge of evaluation statistics.

  • Scientific rigor, autonomy, critical thinking and interest in interdisciplinary projects at the interface between AI and healthcare.

  • Ability to work with clinical teams and translate medical questions into modelling problems.

  • Valued skills: speech and natural-language processing; privacy and anonymization; federated learning or differential privacy; multimodal analysis; explainable AI; AI for healthcare; experience with sensitive or clinical data.

  • Previous scientific publications are appreciated but not mandatory.

  • Ability to communicate effectively in French or English, both orally and in writing.

 

Remuneration

The gross monthly salary is set at €2,300.