Massimiliano Todisco

Associate Professor, EURECOM Digital Security Department - France

Massimiliano Todisco is a professor of audio and speech technologies at the Digital Security Department of EURECOM in France. His research focuses on voice biometrics, deepfake detection, and privacy preservation, with the aim of securing personal data and enhancing the trustworthiness of AI-driven systems. He is widely recognised for proposing the Constant Q Cepstral Coefficients (CQCC)—a handcrafted audio feature that became a standard benchmark in spoofing detection prior to the rise of deep learning. This work earned him the ISCA Best Paper Award for the quinquennium 2015–2019.

His contributions have helped shape modern approaches to synthetic speech detection, advancing both the understanding of artefacts exploited by deep models and the development of more robust signal representations.

Massimiliano is a core organiser of the ASVspoof and VoicePrivacy international challenge series—key initiatives that drive standardisation and innovation across the community. His current research explores adversarial learning, explainability, and the use of generative AI to enhance the transparency and reliability of audio and speech technologies.

Lecture

When Voices Lie: Understanding and Defending Against Voice Deepfakes

In recent years, the rapid rise of accessible AI tools has made it alarmingly easy to create highly realistic voice deepfakes. With just a few seconds of recorded audio, malicious actors can clone a person’s voice to commit fraud. One high-profile example involved a UK energy firm losing €220,000 to attackers who used AI-generated speech to impersonate a company executive and authorise a fraudulent transaction. Other scams have mimicked the voices of distressed family members, successfully convincing victims to send emergency funds.

These threats are not limited to deceiving human listeners. Automatic speaker verification systems—now widely used in smartphones, banking, and smart assistants—are also vulnerable to such attacks. Synthetic or adversarially manipulated speech can bypass these systems, granting access to sensitive information and services.

Over the past decade, voice biometrics has matured into a key technology for secure and convenient user authentication. Yet, public trust remains fragile due to persistent vulnerabilities. Research efforts —particularly through the ASVspoof challenge series— have driven significant progress in detecting spoofing attacks, such as those based on replay, text-to-speech and voice conversion.

However, critical questions remain unanswered: What artefacts do systems detect? How are they represented internally? Can they generalise to unseen attacks? And how can subtle, adversarial perturbations fool even the best-performing models?

This lecture provides a comprehensive introduction to the security challenges of modern voice biometric systems, spanning from traditional spoofing to adversarial machine learning. Topics will include:
– how voice deepfakes are created using publicly available tools with minimal data;
– the vulnerability of automated speaker verification to spoofing and adversarial attacks;
– latent artefacts in speech signals and how models exploit them;
– the role of data augmentation, boosting, and perceptual signal analysis in enhancing robustness;
– advanced defensive strategies, including adversarial training, purification, and explainability-driven countermeasures.

The session will combine accessible explanations with live examples to help demystify these complex threats. It aims to equip participants with a foundational understanding of how trust in voice technologies is being challenged—and how current research is working to restore it.

Main Partners

PARTNERS

Other Partners