Research

I study how multimodal models use evidence when the input is long, incomplete, or unevenly represented. An hour of video has almost no explicit structure, and the query mixes language, metadata, and time. I build systems that select which frames or segments to keep, retrieve the evidence that matters, and keep answers anchored to specific moments. Evaluation is part of the same problem: reported gains move when you change the frame sampling or how the benchmark was built, so the protocols have to make those dependencies visible instead of averaging them away.

The same choices about which evidence to keep decide whose signal a model can read at all. I study that when speech or movement falls outside the training distribution: fairness in speaker diarization, disability representation in image generation, and open data and benchmarks for capability-conditioned settings. Papers are under Publications; datasets, models, and demos under Artifacts.

Publications

Recent work is on multimodal retrieval and reasoning over video. The earlier papers are on multimodal speech processing, robustness, and fairness, and those questions carried over. Citation counts are on Google Scholar; code, datasets, and demos are under Artifacts.

2026

  • Video understanding Last author

    PEEK: Picking Essential frames via Efficient Knowledge distillation

    Killian Steunou, Anas Filali Razzouki, Khalil Guetari, Mounim A. El Yacoubi, Yannis Tevissen · BMVC 2026 · arXiv:2605.31029 · arXiv

2025

  • Video understanding Co-author

    Frame Sampling Strategies Matter: A Benchmark for small vision language models

    Marija Brkic, Anas Filali Razzouki, Yannis Tevissen, Khalil Guetari, Mounim A. El Yacoubi · arXiv:2509.14769 · arXiv

  • Video retrieval Co-inventor

    Computer-based platforms and methods for efficient AI-based digital video shot indexing

    Frédéric Petitpont, Philippe Petitpont, Yannis Tevissen, Khalil Guetari · US Patent 12,288,377 · Apr 2025

2024

  • Video understanding Co-inventor

    Systems and methods for AI generation of image captions enriched with multiple AI modalities

    Frédéric Petitpont, Yannis Tevissen, Khalil Guetari · US Patent 12,148,233 · Nov 2024

  • Video retrieval First author

    Towards Retrieval Augmented Generation over Large Video Libraries

    Yannis Tevissen, Khalil Guetari, Frédéric Petitpont · HSI 2024 · Best Presentation Paper

  • Fairness Sole author

    Disability Representations: Finding Biases in Automatic Image Generation

    Yannis Tevissen · CVPR 2024 Workshop AVA · arXiv

  • Video understanding Co-author

    Multimodal Chaptering for Long-Form TV Newscast Video

    Khalil Guetari, Yannis Tevissen, Frédéric Petitpont · 2024 · arXiv

  • Video understanding First author

    Inserting Faces inside Captions: Image Captioning with Attention Guided Merging

    Yannis Tevissen, Khalil Guetari, Marine Tassel, Erwan Kerleroux, Frédéric Petitpont · arXiv:2405.02305 · arXiv

  • Speech processing Co-author

    Privacy Preserving Personal Assistant with On-Device Diarization and Spoken Dialogue System for Home and Beyond

    Gérard Chollet et al. · IHIET 2024 · arXiv

2023

  • Speech processing Sole author

    Diarisation multimodale: vers des modèles robustes et justes en contexte réel

    Yannis Tevissen · Institut Polytechnique de Paris · HAL

  • Speech processing First author

    Détection d'activité vocale Multi-flux pour la Diarisation du locuteur

    Yannis Tevissen, Jérôme Boudy, Gérard Chollet, Frédéric Petitpont · GRETSI 2023 · HAL

  • Speech processing First author

    Home monitoring for frailty detection through sound and speaker diarization analysis

    Yannis Tevissen et al. · JETSAN 2023 · arXiv

  • Fairness First author

    Towards measuring and scoring speaker diarization fairness

    Yannis Tevissen, Jérôme Boudy, Gérard Chollet, Frédéric Petitpont · arXiv:2302.09991 · arXiv

2022

  • Speech processing First author

    Multi-stream voice activity detection for robust speaker diarization

    Yannis Tevissen, Jérôme Boudy, Gérard Chollet · GDR ISIS 2022

  • Speech processing First author

    The Newsbridge-Telecom SudParis VoxCeleb Speaker Recognition Challenge 2022 System Description

    Yannis Tevissen, Jérôme Boudy, Frédéric Petitpont · VoxCeleb SRC 2022 Task 4 · arXiv

Artifacts

Datasets, models, code, and demos you can download and run, grouped by the research line they came from.

Video understanding research

Code and data behind the frame selection, captioning, retrieval, and diarization work, across video, image, and speech. The recurring question is how much of the available signal a model actually needs, and how to tell when a reported gain is real rather than an artefact of how the input was sampled.

  • Code

    PEEK

    A lightweight dynamic frame selector that distils caption-conditioned relevance from a teacher model, for video captioning under a small frame budget. BMVC 2026.

  • Code

    Frame sampling benchmark

    A controlled benchmark showing that measured video reasoning performance for small vision-language models shifts substantially with the frame sampling strategy alone.

  • Dataset

    AstroCaptions

    Captioning dataset built from archival space imagery, used for multimodal caption enrichment work.

  • Model

    Falcon-TSCPD

    A PEFT adapter for text-based speaker change point detection.

  • Dataset

    VoxConverse TSCPD

    Speaker change point annotations derived from VoxConverse, for evaluating text-side change detection in diarization pipelines.

    BSD 3-Clause

AI for science

Retrieval and evidence grounding applied to scientific literature and clinical data, and the data work that makes capability-conditioned settings measurable at all.

  • Code

    UpToCure

    Assembles published research on rare diseases into readable, sourced reports, applying retrieval and evidence grounding to scientific literature.

  • Collection

    Open SMA Research Data

    A provenance-first index of the spinal muscular atrophy data that is genuinely open, mirrored from the original studies with attribution and checksums, plus a derived reach-intent benchmark and a browser explorer. Licences, access conditions, and stated limitations sit with each repository.

  • Collection

    AI for Disability

    A curated collection of datasets, models, Spaces, and papers that apply AI to disability-related tasks.

Hugging Face profile · GitHub · Publications ↑

Teaching and mentoring

Teaching

  • Speaker Diarization — Guest lecture

    2024 · Télécom SudParis, graduate course on multimodal speech and speaker recognition

    Diarization pipelines, multi-stream voice activity detection, evaluation, and fairness in real-world recording conditions.

Mentoring

Working with me

I am interested in supervising work on evidence selection under a budget, evaluation that survives contact with real data, and assistive systems where the signal depends on the person using them. If that overlaps with what you want to work on, get in touch.

Service

Reviewer for ICASSP 2026, ICPRAI 2026, and ICME 2025, and a member of the scientific committee of JETSAN 2025. Granted US patents on shot-level video indexing and on multimodal enrichment of generated captions are listed under Publications.

Publications ↑ · Artifacts ↑ · Teaching & mentoring ↑