Post-Doctoral Research Visit F/M Continual Audiovisual Perception for Human-Robot Interaction
Contract type : Fixed-term contract
Level of qualifications required : PhD or equivalent
Fonction : Post-Doctoral Research Visit
About the research centre or Inria department
The Centre Inria de l’Université de Grenoble groups together almost 450 people in 26 research teams and 9 research support departments.
Staff is present on three campuses in Grenoble, in close collaboration with other research and higher education institutions (Université Grenoble Alpes, CNRS, CEA, INRAE, …), but also with key economic players in the area.
The Centre Inria de l’Université Grenoble Alpes is active in the fields of high-performance computing, verification and embedded systems, modeling of the environment at multiple levels, and data science and artificial intelligence. The center is a top-level scientific institute with an extensive network of international collaborations in Europe and the rest of the world.
Context
Recent approaches in conversational AI for social assistive robots primarily rely on a human-robot substitution model, where the
robot replaces a human agent. The ĀnandaBot project (PEPR eNSEMBLE, France 2030, 2026–2030) instead investigates a human
(human/robot) approach, in which a robot sidekick accompanies and supports a human agent in a triadic collaborative task — much
as Ananda supported Buddha. The project is coordinated by Fabrice Lefèvre (LIA, Avignon Université) and brings together LIA
(Avignon Université), Inria (RobotLearn team), LISN (CNRS/Université Paris-Saclay), and AP-HP (Broca Hospital), with an application scenario in a gerontology day-hospital setting.
ĀnandaBot develops an audiovisual processing chain enabling triadic verbal and non-verbal interactions between a human user, a
human agent, and their robot sidekick. In this context, the offer is about continual audiovisual robot perception, and aims to develop
methods and algorithms to continuously extract cues about human behaviour from audio and visual data in real-world, socially
situated interactions, with two central requirements: (i) robustness to real-world perturbations, with quantitative estimation of the
reliability of extracted cues, and (ii) continuous adaptation to variations of these perturbations over time. We will cover three tasks:
behaviour understanding from visual inputs, continual audiovisual adaptation, and adaptation to simulated environments.
What do we offer?
- A 24-month postdoctoral contract at Inria Grenoble Rhône-Alpes, within the RobotLearn team.
- Remuneration according to Inria’s postdoctoral salary scale, including standard Inria employee benefits (health insurance, paid leave, restaurant subsidy, etc.).
- Access to Inria’s computing infrastructure and to the ĀnandaBot project’s dedicated hardware (GPU servers, robot platforms).
- Integration in a well-funded, multi-site national consortium (LIA, Inria, LISN, AP-HP) with a clinical deployment site at Broca
Hospital. - Support for travel to project meetings, conferences, and the yearly ĀnandaBot workshop.
References:
1. M. Marge, C. Espy-Wilson, and N. Ward, "Spoken Language Interaction with Robots: Research Issues and Recommendations," Report from the NSF Future Directions Workshop, 2019.
2. Y. Liu et al., "Continual learning for VLMs: A survey and taxonomy beyond forgetting," 2025.
3. Y. Xu, Y. Ban, G. Delorme, C. Gan, D. Rus, and X. Alameda-Pineda, "Transcenter: Transformers with dense representations for multiple-object tracking," IEEE TPAMI, 2022.
4. L. Vaquero, Y. Xu, X. Alameda-Pineda, V. M. Brea, and M. Mucientes, "Lost and found: Overcoming detector failures in online multi-object tracking," ECCV, 2024.
5. Y. Ban, X. Alameda-Pineda, L. Girin, and R. Horaud, "Variational Bayesian inference for audio-visual tracking of multiple speakers," IEEE TPAMI, 2019.
6. X. Alameda-Pineda et al., "Socially pertinent robots in gerontological healthcare," International Journal of Social Robotics, 2025.
7. A. Golmakani, M. Sadeghi, X. Alameda-Pineda, and R. Serizel, "A weighted-variance variational autoencoder model for speech enhancement," ICASSP, 2024.
8. J.-E. Ayilo, M. Sadeghi, R. Serizel, and X. Alameda-Pineda, "Diffusion-based unsupervised audio-visual speech enhancement," ICASSP, 2025.
9. S. Sadok, S. Leglaive, L. Girin, X. Alameda-Pineda, and R. Séguier, "A multimodal dynamical variational autoencoder for audiovisual speech representation learning," Neural Networks, 2024.
10. A. Ballou, X. Alameda-Pineda, and C. Reinke, "Variational meta reinforcement learning for social robotics," Applied Intelligence, 2023.
11. R. Aljundi, K. Kelchtermans, and T. Tuytelaars, "Task-free continual learning," CVPR, 2019.
12. S. Mo, W. Pian, and Y. Tian, "Class-incremental grouping network for continual audio-visual learning," ICCV, 2023.
Assignment
The postdoctoral researcher will contribute to the design, implementation, and evaluation of the audiovisual perception pipeline of
ĀnandaBot, under the supervision of Xavier Alameda-Pineda and Karteek Alahari. The position covers the following topics:
- Behaviour understanding from visual input: multi-person localisation and tracking, extraction of audiovisual cues (speech,
gaze, posture) relevant to socially situated triadic interactions, with generalisation to previously unseen people, objects, and
acoustic/visual conditions. - Continual audiovisual adaptation: development of methods enabling the perception modules to adapt on-the-fly to distribution shifts (new speakers, new environments, sensor perturbations) without catastrophic forgetting, building on the team’s prior work in continual and self-supervised learning.
- Adaptation to simulated environments: transfer of perception modules between the real-world experimental platform (deployed at Broca Hospital, AP-HP) and a simulation platform used for lower-cost, privacy-preserving training and evaluation, in coordination with the dialogue/engagement modules.
Main activities
- Design and implement audiovisual perception algorithms (multi-person detection/tracking, speaker diarization, cue fusion)
robust to real-world social conditions. - Develop continual-learning strategies for on-site, low-resource adaptation of audiovisual models, respecting the project’s
privacy-by-design and data-efficiency requirements. - Contribute to the simulation platform enabling transfer and evaluation of perception modules without requiring human
participants. - Collaborate with AP-HP (Broca Hospital) on the definition of experimental protocols and the evaluation of perception modules
in laboratory and real-world (gerontology day-hospital) settings. - Contribute to scientific publications, open-source releases, and the project’s data/privacy compliance (informed consent,
secure on-site storage, CNIL and Ethics Committee submissions). - Participate in project meetings (LIA, Inria, LISN, AP-HP) and the yearly ĀnandaBot workshop.
Skills
- PhD in Computer Science, Machine Learning, Signal Processing, or a closely related field, completed or nearly completed at the
start date. - Strong background in computer vision and/or audio(-visual) machine learning (e.g., multi-object tracking, speaker/source
localisation, multimodal fusion). - Experience with deep learning frameworks (PyTorch or equivalent) and proficiency in Python.
- Familiarity with continual/lifelong learning and/or self-supervised learning is a strong plus.
- Experience with robotic platforms or human-robot interaction is a plus but not required.
- Good written and spoken English; French is not required but is a plus.
- Ability to work in a multidisciplinary, multi-site consortium.
Benefits package
- Subsidized meals
- Partial reimbursement of public transport costs
- Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
- Possibility of teleworking and flexible organization of working hours
- Professional equipment available (videoconferencing, loan of computer equipment, etc.)
- Social, cultural and sports events and activities
- Access to vocational training
- Social security coverage under conditions
Remuneration
2788 € gross salary / month
General Information
- Theme/Domain :
Vision, perception and multimedia interpretation
IT Technical and production engineering (BAP E) - Town/city : Montbonnot
- Inria Center : Centre Inria de l'Université Grenoble Alpes
- Starting date : 2026-10-01
- Duration of contract : 2 years
- Deadline to apply : 2026-08-30
Warning : you must enter your e-mail address in order to save your application to Inria. Applications must be submitted online on the Inria website. Processing of applications sent from other channels is not guaranteed.
Instruction to apply
Applications must be submitted online via the Inria website. Processing of applications submitted via other channels is not guaranteed.
Defence Security :
This position is likely to be situated in a restricted area (ZRR), as defined in Decree No. 2011-1425 relating to the protection of national scientific and technical potential (PPST).Authorisation to enter an area is granted by the director of the unit, following a favourable Ministerial decision, as defined in the decree of 3 July 2012 relating to the PPST. An unfavourable Ministerial decision in respect of a position situated in a ZRR would result in the cancellation of the appointment.
Recruitment Policy :
As part of its diversity policy, all Inria positions are accessible to people with disabilities.
Contacts
- Inria Team : ROBOTLEARN
-
Recruiter :
Alameda Pineda Xavier / xavier.alameda-pineda@inria.fr
About Inria
Inria, the French national institute for research in digital science and technology, supports the French government in national research and innovation strategies in the digital field, acting as Digital Programs Agency. Inria leads over 300 research and innovation projects with its 3,500 scientists, engineers, and support staff, in partnership with universities and the digital ecosystem (businesses, entrepreneurs, and public stakeholders). Together, we explore strategic fields such as artificial intelligence, cybersecurity, quantum computing, cloud technologies, digital transformation in healthcare, digital twins, and digital technologies for defence. We develop practical solutions such as software, tech startups, partnerships with national companies, and cutting-edge training programmes. Our goal is to drive scientific, technological, and industrial excellence to ensure France’s digital sovereignty.