Head of Research · Moments Lab
I lead research on multimodal retrieval and reasoning at Moments Lab, after a PhD at Institut Polytechnique de Paris on multimodal speaker diarization, advised by Jérôme Boudy and Gérard Chollet. One question governs my work: how can a model efficiently understand long-context data with complex temporal, spatial and narrative dependencies?
I also contribute to research on fairness, accessibility and disability representation while openly writing, speaking about my experience and advocating for a more inclusive science.
Finally, I built UpToCure, which assembles research on rare diseases into readable, sourced reports. I co-founded VocaCoach, a speech training platform (€1M valuation, VivaTech award, covered by Le Parisien), and serve on the board of Universal Wings Mobility, the French startup building AVI, a smart travel assistant for people with disabilities.
Research focus
Multimodal learning
A useful answer usually needs more than one signal at once: vision, language, audio, and time. I build models and pipelines that combine those modalities so retrieval and reasoning stay grounded in what actually happened in the input.
Long context efficiency
Long inputs are expensive to process and easy to summarise poorly. I work on selecting and compressing evidence so a model can handle hours of video or speech without losing the moments that matter.
Fairness and biases
Models inherit the gaps in their training data, and the people furthest from the distribution pay for it first. I measure how disability and demographic biases show up in vision and speech systems, and I build open data and benchmarks for those settings.
Selected work
Three papers that carry the agenda. The full record is under Publications; the code, datasets, and demos are under Artifacts.
-
Distils caption-conditioned relevance from a teacher model into a small frame selector, so a captioning model can work from far fewer frames.
-
HSI 2024 Best Presentation
Towards Retrieval Augmented Generation over Large Video Libraries
Answers questions about a large video library by retrieving the relevant segments first and generating only from what came back, so the answer can cite the moments it rests on.
-
Measures how image generation models depict disability. Prompted only for “a disabled man” and “a disabled woman”, Midjourney and SDXL Turbo returned a wheelchair every time.
See also
Research
The research agenda, publications, artifacts, and teaching and mentoring.
Blog & Advocacy
Technical notes, press, and talks on Blog & Press, and disability and accessibility work under advocacy.