Research
I study how multimodal models use evidence when the input is long, incomplete, or unevenly represented. An hour of video has almost no explicit structure, and the query mixes language, metadata, and time. I build systems that select which frames or segments to keep, retrieve the evidence that matters, and keep answers anchored to specific moments. Evaluation is part of the same problem: reported gains move when you change the frame sampling or how the benchmark was built, so the protocols have to make those dependencies visible instead of averaging them away.
The same choices about which evidence to keep decide whose signal a model can read at all. I study that when speech or movement falls outside the training distribution: fairness in speaker diarization, disability representation in image generation, and open data and benchmarks for capability-conditioned settings. Papers are under Publications; datasets, models, and demos under Artifacts.
Publications
Recent work is on multimodal retrieval and reasoning over video. The earlier papers are on multimodal speech processing, robustness, and fairness, and those questions carried over. Citation counts are on Google Scholar; code, datasets, and demos are under Artifacts.
2026
-
Video understanding Last authorPEEK: Picking Essential frames via Efficient Knowledge distillation
2025
-
Video understanding Co-authorFrame Sampling Strategies Matter: A Benchmark for small vision language models
-
Video retrieval Co-inventorComputer-based platforms and methods for efficient AI-based digital video shot indexing
2024
-
Video understanding Co-inventorSystems and methods for AI generation of image captions enriched with multiple AI modalities
- Video retrieval First author
Towards Retrieval Augmented Generation over Large Video Libraries
-
Fairness Sole authorDisability Representations: Finding Biases in Automatic Image Generation
-
Video understanding Co-authorMultimodal Chaptering for Long-Form TV Newscast Video
-
Video understanding First authorInserting Faces inside Captions: Image Captioning with Attention Guided Merging
- Speech processing Co-author
Privacy Preserving Personal Assistant with On-Device Diarization and Spoken Dialogue System for Home and Beyond
2023
- Speech processing Sole author
Diarisation multimodale: vers des modèles robustes et justes en contexte réel
- Speech processing First author
Détection d'activité vocale Multi-flux pour la Diarisation du locuteur
- Speech processing First author
Home monitoring for frailty detection through sound and speaker diarization analysis
-
Fairness First authorTowards measuring and scoring speaker diarization fairness
2022
- Speech processing First author
Multi-stream voice activity detection for robust speaker diarization
- Speech processing First author
The Newsbridge-Telecom SudParis VoxCeleb Speaker Recognition Challenge 2022 System Description
Artifacts
Datasets, models, code, and demos you can download and run, grouped by the research line they came from.
Video understanding research
Code and data behind the frame selection, captioning, retrieval, and diarization work, across video, image, and speech. The recurring question is how much of the available signal a model actually needs, and how to tell when a reported gain is real rather than an artefact of how the input was sampled.
- Code
- Code
- Dataset
- Model
- Dataset
AI for science
Retrieval and evidence grounding applied to scientific literature and clinical data, and the data work that makes capability-conditioned settings measurable at all.
- Code
- Collection
- Collection
Teaching and mentoring
Teaching
-
Speaker Diarization — Guest lecture
Diarization pipelines, multi-stream voice activity detection, evaluation, and fairness in real-world recording conditions.
Mentoring
- Killian Steunou PhD
- Anas Filali Razzouki Postdoc
- Ruben Leon Intern
- Marija Brkic Intern
- Estelle Zheng Intern
- Prateek Gulati Intern
Working with me
I am interested in supervising work on evidence selection under a budget, evaluation that survives contact with real data, and assistive systems where the signal depends on the person using them. If that overlaps with what you want to work on, get in touch.
Service
Reviewer for ICASSP 2026, ICPRAI 2026, and ICME 2025, and a member of the scientific committee of JETSAN 2025. Granted US patents on shot-level video indexing and on multimodal enrichment of generated captions are listed under Publications.