Science & research
The Thorsten-Voice speech datasets are now used in more than 20 scientific papers by international research institutions – from German universities to MIT and the University of Texas at Austin. All datasets are freely available under the CC0 license on Zenodo and Hugging Face. You'll find technical details on the individual datasets on the recordings page.
Hof University of Applied Sciences
University of Stuttgart
TU Graz — Signal Processing and Speech Communication Laboratory
- Text-to-Speech Data Augmentation for Dysarthric Child Speech Reconstruction (Pfeiler, Mayrhofer, Schuppler & Hagmüller, Forum Acusticum 2026) ↗ — Uses the Thorsten-Voice dataset to pretrain TTS models for data augmentation when reconstructing impaired child speech.
University of Magdeburg
Maastricht University
Fraunhofer Institute for Applied and Integrated Security (AISEC)
- 2024-01-17 MLAAD: The Multi-Language Audio Anti-Spoofing Dataset ↗ — Thorsten Müller is a co-author here.
German Informatics Society (Gesellschaft für Informatik)
University of Augsburg
IEEE Engineering in Medicine & Biology Society
University of Southern California
Yıldız Technical University
Adıyaman University
Massachusetts Institute of Technology (MIT)
- 2023-10-11 Audio-Visual Neural Syntax Acquisition ↗
University of Texas at Austin
Universitat Politècnica de Catalunya
Toyota Technological Institute at Chicago
Virginia Commonwealth University
University of Bucharest
POSTECH, Republic of Korea
University of Lübeck
- Automatische Optimierung von Audiosignalen für Transkription mit Evolutionären Algorithmen und Machine Learning ↗ — Deals with speech technology in the healthcare sector.
HAL Open Science
- Popular Voices: Computational Analysis of Poetry and Song (2026) ↗ — Thorsten-Voice is used as an example here.
Saarland University