← Back to About the project

Technical details on the recordings

Since 2019, the technology behind Thorsten-Voice has evolved continuously – from Coqui TTS to Piper TTS and the current CosyVoice models. Each generation brought more natural intonation, more emotion and new dialect variants. All of this is based on more than 30,000 recordings that I recorded myself over the years.

License & availability

All recordings are freely available under the CC0 license on Zenodo and Hugging Face. Thanks to this permissive license, they are also well suited for science and research.

The datasets

Thorsten-Voice Dataset 2021.02 (Neutral)

Number of recordings
22,668
Audio duration
23+ hours
Sample rate
22,050 Hz
Channels
Mono
Normalization
-24 dB
Sentence length (min/avg/max)
2 / 52 / 180 characters
Speaking rate (avg)
14 characters/second
Questions
2,780
Exclamations
1,840
Citation (BibTeX)
@dataset{muller_thorsten_2021_5525342,
  author    = {Müller, Thorsten and Kreutz, Dominik},
  title     = {Thorsten - Open German Voice (Neutral) Dataset},
  month     = feb,
  year      = 2021,
  publisher = {Zenodo},
  version   = {3.0},
  doi       = {10.5281/zenodo.5525342},
  url       = {https://doi.org/10.5281/zenodo.5525342}
}

Thorsten-Voice Dataset 2021.06 (Emotional)

300 different sentences, each spoken in eight emotions: neutral, disgusted, angry, amused, surprised, sleepy, whispering and drunk (acted only – I was sober during the recordings).

Number of recordings
2,400
Channels
Mono
Normalization
-24 dB
Sentence length (min/max)
59 / 148 characters
Citation (BibTeX)
@dataset{muller_thorsten_2021_5525023,
  author    = {Müller, Thorsten and Kreutz, Dominik},
  title     = {Thorsten - Open German Voice (Emotional) Dataset},
  month     = jun,
  year      = 2021,
  publisher = {Zenodo},
  version   = {2.0},
  doi       = {10.5281/zenodo.5525023},
  url       = {https://doi.org/10.5281/zenodo.5525023}
}

Thorsten-Voice Dataset 2022.10 (Neutral)

Number of recordings
12,432
Audio duration
11 hours
Sample rate
22,050 Hz
Channels
Mono
Normalization
-24 dB
Speaking rate (avg)
17.5 characters/second
Citation (BibTeX)
@dataset{muller_thorsten_2022_7265581,
  author    = {Müller, Thorsten and Kreutz, Dominik},
  title     = {ThorstenVoice Dataset 2022.10},
  month     = oct,
  year      = 2022,
  publisher = {Zenodo},
  version   = {1.0},
  doi       = {10.5281/zenodo.7265581},
  url       = {https://doi.org/10.5281/zenodo.7265581}
}

Thorsten-Voice Dataset 2023.09 (Hessian dialect)

Number of recordings
2,108
Audio duration
approx. 2 hours
Sample rate
22,050 Hz
Channels
Mono
Normalization
-24 dB
Citation (BibTeX)
@dataset{muller_2024_10511260,
  author    = {Müller, Thorsten and Kreutz, Dominik},
  title     = {Thorsten-Voice Dataset 2023.09 Hessisch},
  month     = jan,
  year      = 2024,
  publisher = {Zenodo},
  doi       = {10.5281/zenodo.10511260},
  url       = {https://doi.org/10.5281/zenodo.10511260}
}

Thorsten-Voice Dataset (TV-44kHz-Full)

All recordings in a single dataset, in the original 44 kHz sample rate, logically split into subsets and enriched with metadata on duration, speaking rate, recording month and quality.

Citation (BibTeX)
@misc{thorsten_mueller_2024,
  author    = {Thorsten Müller},
  title     = {TV-44kHz-Full (Revision ff427ec)},
  year      = 2024,
  url       = {https://huggingface.co/datasets/Thorsten-Voice/TV-44kHz-Full},
  doi       = {10.57967/hf/3290},
  publisher = {Hugging Face}
}

Thorsten-Voice Dataset 2025.12 (Mini FineTuning)

With only 60 recordings, much more compact than the other datasets. Intended for fine-tuning existing Thorsten-Voice TTS models.

Number of recordings
60
Sample rate
24 kHz
Normalization
-24 dB

Thorsten-Voice on YouTube

As an enthusiast for free speech technology, I've been running the "Thorsten-Voice" YouTube channel for several years. There I regularly publish step-by-step tutorials on open-source speech technology, news from the field and occasional interviews with fascinating people from the world of free speech synthesis.

Thorsten-Voice @ YouTube