Post-Doctoral Research Visit F/M Editing and Conditional Generation with Text-to-Video Generation Models
Contract type : Fixed-term contract
Level of qualifications required : PhD or equivalent
Fonction : Post-Doctoral Research Visit
About the research centre or Inria department
The Inria Grenoble research center groups together almost 600 people in 27 research teams and 8 research support departments.
Staff is present on three campuses in Grenoble, in close collaboration with other research and higher education institutions (University Grenoble Alpes, CNRS, CEA, INRAE, …), but also with key economic players in the area.
Inria Grenoble is active in the fields of high-performance computing, verification and embedded systems, modeling of the environment at multiple levels, and data science and artificial intelligence. The center is a top-level scientific institute with an extensive network of international collaborations in Europe and the rest of the world.
Context
Titre : Editing and Conditional Generation with Text-to-Video Generation Models
Supervision : Dr Stéphane Lathuilière (INRIA-UGA)
Funding : BPI contract
Contexte :
Recent advancements in generative AI, and in particular diffusion models [1,2], have significantly enhanced the capabilities of text-to-video (T2V) models [3,4], allowing users to produce richly varied and imaginative scenes from natural language descriptions. These systems demonstrate strong scene diversity and flexibility, making them attractive for applications in entertainment, simulation, and human–computer interaction.
However, a persistent limitation lies in their inability to enforce fine-grained conditioning and maintain strict consistency for specific visual elements. For example, while a T2V model can generate a “person walking in a park,” it struggles to ensure the persistent appearance of a specific object, character identity, or detailed attribute (such as a specific garment [5]) across complex poses and dynamic environmental interactions.
In contrast, highly specialized image and video editing systems—such as those designed for virtual try-on [5], face swapping, or precise object insertion—excel at fine-grained conditioning on target individuals or objects. They can adapt elements to morphology, pose, and texture details with remarkable realism. Yet, these specialized approaches generally operate in isolation, lacking the scene diversity and broader contextual awareness that foundational T2V models offer.
Bridging these two paradigms offers a powerful opportunity: to synthesize realistic, precisely controllable subjects and objects embedded within richly described, dynamic environments. To achieve this, novel alignment and editing techniques are required. Specifically, post-training with Reinforcement Learning (RL) presents a highly promising methodology to overcome these limitations. By leveraging RL during the post-training phase, foundation T2V models can be explicitly optimized to follow complex conditioning signals, enforce temporal consistency, and align with specific human-defined objectives for fine-grained editing tasks without sacrificing their generative diversity.
Assignment
Research Objectives :
The primary mission of the Postdoctoral Research Fellow will be to advance the state-of-the-art in controllable and editable Text-to-Video (T2V) generation. The successful candidate will design, implement, and evaluate novel deep generative models and methodologies that address the current limitations of existing T2V systems. A core focus will be on achieving fine-grained conditional generation via post-training with Reinforcement Learning (RL), allowing users to specify complex temporal, spatial, and stylistic constraints, as well as enabling intuitive and high-fidelity post-generation editing of the video content. The research will aim to produce models that are not only photorealistic but also exhibit high semantic fidelity, temporal coherence, and practical usability in creative and industrial applications.
Main activities
2. Main Tasks
The Postdoctoral Research Fellow will be responsible for the following main tasks. They will engage in Model Design and Development by designing and implementing novel architectures (e.g., Diffusion Models, Transformers, VAEs) specifically tailored for high-resolution, temporally consistent, and controllable video generation. A key focus is to develop conditional generation techniques to guide the Text-to-Video process using various complex inputs beyond a simple text prompt, such as image references, motion skeletons, semantic masks, or detailed scene descriptions. They will extensively research Video Editing and Manipulation, developing methods for high-fidelity post-generation video editing, allowing for non-destructive modification of generated videos (e.g., object replacement, style transfer, background alteration) while maintaining strong temporal consistency. Furthermore, they will investigate in-context editing mechanisms that enable precise changes to specific segments or objects within a generated video based on new text or image prompts. A core part of the role is Addressing Key T2V Challenges. This includes tackling the fundamental challenge of temporal coherence and consistency, ensuring that generated videos do not suffer from "flickering" or object identity changes across frames, and developing strategies to improve semantic fidelity, resolving issues where models misinterpret complex text prompts. They will also explore methods for efficient training and inference to manage the significant computational cost associated with high-resolution, long-duration video generation, and address the difficulties of data scarcity and bias through techniques like data augmentation or cross-modal transfer learning. Finally, they will perform Evaluation and Benchmarking, establishing rigorous quantitative and qualitative metrics to assess the quality, editability, and controllability of the developed models. The fellow is expected to prioritize Dissemination and Collaboration, which involves documenting research findings and publishing high-quality papers in top-tier machine learning and computer vision venues, actively participating in departmental seminars, and contributing to collaborative projects.
Skills
Compétences techniques et niveau requis :We are seeking a motivated PhD candidate with a strong background in one or more the following areas :
- speech processing, computer vision, machine learning,
- solid programmming skills
- interest in connecting AI with human cognition Prior experience with LLM, SpeechLMs, RL algorithms, or robotic platforms is a plus, but not mandatory
Langues : Anglais
Benefits package
- Subsidized meals
- Partial reimbursement of public transport costs
- Leave: 7 weeks of annual leave + 10 extra days off due to RTT (statutory reduction in working hours) + possibility of exceptional leave (sick children, moving home, etc.)
- Possibility of teleworking (90 days / year) and flexible organization of working hours
- Professional equipment available (videoconferencing, loan of computer equipment, etc.)
- Social, cultural and sports events and activities
- Access to vocational training
- Complementary health insurance under conditions
Remuneration
2788€ brut / mois
General Information
- Theme/Domain :
Vision, perception and multimedia interpretation
Statistics (Big data) (BAP E) - Town/city : Montbonnot
- Inria Center : Centre Inria de l'Université Grenoble Alpes
- Starting date : 2026-10-01
- Duration of contract : 2 years
- Deadline to apply : 2026-08-20
Warning : you must enter your e-mail address in order to save your application to Inria. Applications must be submitted online on the Inria website. Processing of applications sent from other channels is not guaranteed.
Instruction to apply
Applications must be submitted online on the Inria website.
Processing of applications sent by other channels is not guaranteed.
Defence Security :
This position is likely to be situated in a restricted area (ZRR), as defined in Decree No. 2011-1425 relating to the protection of national scientific and technical potential (PPST).Authorisation to enter an area is granted by the director of the unit, following a favourable Ministerial decision, as defined in the decree of 3 July 2012 relating to the PPST. An unfavourable Ministerial decision in respect of a position situated in a ZRR would result in the cancellation of the appointment.
Recruitment Policy :
As part of its diversity policy, all Inria positions are accessible to people with disabilities.
Contacts
- Inria Team : ROBOTLEARN
-
Recruiter :
Lathuiliere Stephane / stephane.lathuiliere@inria.fr
About Inria
Inria, the French national institute for research in digital science and technology, supports the French government in national research and innovation strategies in the digital field, acting as Digital Programs Agency. Inria leads over 300 research and innovation projects with its 3,500 scientists, engineers, and support staff, in partnership with universities and the digital ecosystem (businesses, entrepreneurs, and public stakeholders). Together, we explore strategic fields such as artificial intelligence, cybersecurity, quantum computing, cloud technologies, digital transformation in healthcare, digital twins, and digital technologies for defence. We develop practical solutions such as software, tech startups, partnerships with national companies, and cutting-edge training programmes. Our goal is to drive scientific, technological, and industrial excellence to ensure France’s digital sovereignty.