Post-Doctoral Research Visit F/M Continual Audiovisual Perception for Human-Robot Interaction
Type de contrat : CDD
Niveau de diplôme exigé : Thèse ou équivalent
Fonction : Post-Doctorant
A propos du centre ou de la direction fonctionnelle
The Centre Inria de l’Université de Grenoble groups together almost 450 people in 26 research teams and 9 research support departments.
Staff is present on three campuses in Grenoble, in close collaboration with other research and higher education institutions (Université Grenoble Alpes, CNRS, CEA, INRAE, …), but also with key economic players in the area.
The Centre Inria de l’Université Grenoble Alpes is active in the fields of high-performance computing, verification and embedded systems, modeling of the environment at multiple levels, and data science and artificial intelligence. The center is a top-level scientific institute with an extensive network of international collaborations in Europe and the rest of the world.
Contexte et atouts du poste
Recent approaches in conversational AI for social assistive robots primarily rely on a human-robot substitution model, where the
robot replaces a human agent. The ĀnandaBot project (PEPR eNSEMBLE, France 2030, 2026–2030) instead investigates a human
(human/robot) approach, in which a robot sidekick accompanies and supports a human agent in a triadic collaborative task — much
as Ananda supported Buddha. The project is coordinated by Fabrice Lefèvre (LIA, Avignon Université) and brings together LIA
(Avignon Université), Inria (RobotLearn team), LISN (CNRS/Université Paris-Saclay), and AP-HP (Broca Hospital), with an application scenario in a gerontology day-hospital setting.
ĀnandaBot develops an audiovisual processing chain enabling triadic verbal and non-verbal interactions between a human user, a
human agent, and their robot sidekick. In this context, the offer is about continual audiovisual robot perception, and aims to develop
methods and algorithms to continuously extract cues about human behaviour from audio and visual data in real-world, socially
situated interactions, with two central requirements: (i) robustness to real-world perturbations, with quantitative estimation of the
reliability of extracted cues, and (ii) continuous adaptation to variations of these perturbations over time. We will cover three tasks:
behaviour understanding from visual inputs, continual audiovisual adaptation, and adaptation to simulated environments.
What do we offer?
- A 24-month postdoctoral contract at Inria Grenoble Rhône-Alpes, within the RobotLearn team.
- Remuneration according to Inria’s postdoctoral salary scale, including standard Inria employee benefits (health insurance, paid leave, restaurant subsidy, etc.).
- Access to Inria’s computing infrastructure and to the ĀnandaBot project’s dedicated hardware (GPU servers, robot platforms).
- Integration in a well-funded, multi-site national consortium (LIA, Inria, LISN, AP-HP) with a clinical deployment site at Broca
Hospital. - Support for travel to project meetings, conferences, and the yearly ĀnandaBot workshop.
References:
1. M. Marge, C. Espy-Wilson, and N. Ward, "Spoken Language Interaction with Robots: Research Issues and Recommendations," Report from the NSF Future Directions Workshop, 2019.
2. Y. Liu et al., "Continual learning for VLMs: A survey and taxonomy beyond forgetting," 2025.
3. Y. Xu, Y. Ban, G. Delorme, C. Gan, D. Rus, and X. Alameda-Pineda, "Transcenter: Transformers with dense representations for multiple-object tracking," IEEE TPAMI, 2022.
4. L. Vaquero, Y. Xu, X. Alameda-Pineda, V. M. Brea, and M. Mucientes, "Lost and found: Overcoming detector failures in online multi-object tracking," ECCV, 2024.
5. Y. Ban, X. Alameda-Pineda, L. Girin, and R. Horaud, "Variational Bayesian inference for audio-visual tracking of multiple speakers," IEEE TPAMI, 2019.
6. X. Alameda-Pineda et al., "Socially pertinent robots in gerontological healthcare," International Journal of Social Robotics, 2025.
7. A. Golmakani, M. Sadeghi, X. Alameda-Pineda, and R. Serizel, "A weighted-variance variational autoencoder model for speech enhancement," ICASSP, 2024.
8. J.-E. Ayilo, M. Sadeghi, R. Serizel, and X. Alameda-Pineda, "Diffusion-based unsupervised audio-visual speech enhancement," ICASSP, 2025.
9. S. Sadok, S. Leglaive, L. Girin, X. Alameda-Pineda, and R. Séguier, "A multimodal dynamical variational autoencoder for audiovisual speech representation learning," Neural Networks, 2024.
10. A. Ballou, X. Alameda-Pineda, and C. Reinke, "Variational meta reinforcement learning for social robotics," Applied Intelligence, 2023.
11. R. Aljundi, K. Kelchtermans, and T. Tuytelaars, "Task-free continual learning," CVPR, 2019.
12. S. Mo, W. Pian, and Y. Tian, "Class-incremental grouping network for continual audio-visual learning," ICCV, 2023.
Mission confiée
The postdoctoral researcher will contribute to the design, implementation, and evaluation of the audiovisual perception pipeline of
ĀnandaBot, under the supervision of Xavier Alameda-Pineda and Karteek Alahari. The position covers the following topics:
- Behaviour understanding from visual input: multi-person localisation and tracking, extraction of audiovisual cues (speech,
gaze, posture) relevant to socially situated triadic interactions, with generalisation to previously unseen people, objects, and
acoustic/visual conditions. - Continual audiovisual adaptation: development of methods enabling the perception modules to adapt on-the-fly to distribution shifts (new speakers, new environments, sensor perturbations) without catastrophic forgetting, building on the team’s prior work in continual and self-supervised learning.
- Adaptation to simulated environments: transfer of perception modules between the real-world experimental platform (deployed at Broca Hospital, AP-HP) and a simulation platform used for lower-cost, privacy-preserving training and evaluation, in coordination with the dialogue/engagement modules.
Principales activités
- Design and implement audiovisual perception algorithms (multi-person detection/tracking, speaker diarization, cue fusion)
robust to real-world social conditions. - Develop continual-learning strategies for on-site, low-resource adaptation of audiovisual models, respecting the project’s
privacy-by-design and data-efficiency requirements. - Contribute to the simulation platform enabling transfer and evaluation of perception modules without requiring human
participants. - Collaborate with AP-HP (Broca Hospital) on the definition of experimental protocols and the evaluation of perception modules
in laboratory and real-world (gerontology day-hospital) settings. - Contribute to scientific publications, open-source releases, and the project’s data/privacy compliance (informed consent,
secure on-site storage, CNIL and Ethics Committee submissions). - Participate in project meetings (LIA, Inria, LISN, AP-HP) and the yearly ĀnandaBot workshop.
Compétences
- PhD in Computer Science, Machine Learning, Signal Processing, or a closely related field, completed or nearly completed at the
start date. - Strong background in computer vision and/or audio(-visual) machine learning (e.g., multi-object tracking, speaker/source
localisation, multimodal fusion). - Experience with deep learning frameworks (PyTorch or equivalent) and proficiency in Python.
- Familiarity with continual/lifelong learning and/or self-supervised learning is a strong plus.
- Experience with robotic platforms or human-robot interaction is a plus but not required.
- Good written and spoken English; French is not required but is a plus.
- Ability to work in a multidisciplinary, multi-site consortium.
Avantages
- Subsidized meals
- Partial reimbursement of public transport costs
- Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
- Possibility of teleworking and flexible organization of working hours
- Professional equipment available (videoconferencing, loan of computer equipment, etc.)
- Social, cultural and sports events and activities
- Access to vocational training
- Social security coverage under conditions
Rémunération
2788 € gross salary / month
Informations générales
- Thème/Domaine :
Vision, perception et interprétation multimedia
Ingénierie technique et de production (TIC) (BAP E) - Ville : Montbonnot
- Centre Inria : Centre Inria de l'Université Grenoble Alpes
- Date de prise de fonction souhaitée : 2026-10-01
- Durée de contrat : 2 ans
- Date limite pour postuler : 2026-08-30
Attention: Les candidatures doivent être déposées en ligne sur le site Inria. Le traitement des candidatures adressées par d'autres canaux n'est pas garanti.
Consignes pour postuler
Applications must be submitted online via the Inria website. Processing of applications submitted via other channels is not guaranteed.
Sécurité défense :
Ce poste est susceptible d’être affecté dans une zone à régime restrictif (ZRR), telle que définie dans le décret n°2011-1425 relatif à la protection du potentiel scientifique et technique de la nation (PPST). L’autorisation d’accès à une zone est délivrée par le chef d’établissement, après avis ministériel favorable, tel que défini dans l’arrêté du 03 juillet 2012, relatif à la PPST. Un avis ministériel défavorable pour un poste affecté dans une ZRR aurait pour conséquence l’annulation du recrutement.
Politique de recrutement :
Dans le cadre de sa politique diversité, tous les postes Inria sont accessibles aux personnes en situation de handicap.
Contacts
- Équipe Inria : ROBOTLEARN
-
Recruteur :
Alameda Pineda Xavier / xavier.alameda-pineda@inria.fr
A propos d'Inria
Inria, l'institut national de recherche dans les sciences et technologies du numérique, est en appui de l’État pour les stratégies nationales de recherche et d’innovation du numérique en tant qu'Agence de programmes. Inria mène plus de 300 projets de recherche et d’innovation avec ses 3500 scientifiques, ingénieurs et personnels d’appui, en partenariat avec les universités et l’écosystème numérique (entreprises, entrepreneurs, acteurs publics). Ensemble, nous explorons des domaines clés comme l'intelligence artificielle, la cybersécurité, l’informatique quantique, le Cloud, la transformation numérique de la santé, les jumeaux numériques ou encore les technologies numériques pour la défense. Nous construisons des solutions concrètes telles que des logiciels, des startups technologiques, des partenariats avec les entreprises du tissu national et des formations de pointe. Notre objectif : l’impact scientifique, technologique et industriel au service de la souveraineté numérique de la France.