MINTIME: Multi-Identity Size-Invariant Video Deepfake Detection
About this Topic:
Despite the strong performance achieved by deepfake detection methods on benchmark datasets, their application to real-world videos remains challenging due to several conditions that are often overlooked. For example, manipulated videos may show several people, and faces can appear at a great variety of scales due to variations in distance and framing. Most of the literature approaches either deal with one person at a time or combine the face-level predictions by using simple aggregation methods, which may result in important information being lost.
This webinar introduces Multi-Identity Size-Invariant Video Deepfake Detection (MINTIME), a video deepfake detection approach which is designed to simultaneously model multiple identities and variations in face size. It combines convolutional feature extraction with a spatio-temporal transformer and includes an identity-aware attention mechanism to effectively model face sequences that belong to different people. The model also explicitly takes into account temporal information and the relative face size.
The presenter will examine the reasons that led to these design decisions, look at the architecture of MINTIME, and cover the way in which it was evaluated. The experimental results show the relevance of performing spatio-temporal analysis and the advantages of using identity-aware and size-aware modelling for robust video deepfake detection.
About the Presenter:
Davide Alessandro Coccomini received the B.Sc. degree in computer engineering, the M.Sc. degree in artificial intelligence and data engineering, and the Ph.D. degree in information engineering all from the University of Pisa, Pisa, Italy in 2019, 2021 and 2025 respectively.
He is currently a Machine Learning Engineer at Aryel, where he works on the research and development of machine learning solutions, with a particular focus on computer vision for immersive advertising experiences. During his doctoral studies, he was a Research Associate at the Institute of Information Science and Technologies of the Italian National Research Council (CNR) in Pisa, Italy, where his research focused on deepfake detection in images and videos. His research interests include deepfake detection, multimedia forensics, computer vision, deep learning, vision transformers, and the robustness and generalization of artificial intelligence models.
Want to learn more about upcoming events & webinars?
Visit the events section of the Signal Processing website to see all upcoming lectures, workshops, webinars, and more.