Top AI Repos β open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
π A comprehensive list of open-source datasets for voice and sound computing (95+ datasets).
| Date | Stars |
|---|---|
| 2026-07-24 | 2212 |
| 2026-07-25 | 2212 |
| 2026-07-28 | 2212 |
| 2026-07-30 | 2212 |
| 2026-08-06 | 2212 |
Today
β stars today
This week
β stars this week
This month
β stars this month
Momentum
0.0
growth rate 0.00%/day
# voice_datasets A comprehensive list of open source voice and music datasets. I released this for the talk @ the [VOICE Summit 2019.](https://github.com/jim-schwoebel/voice_gender_detection)  ## Audio datasets There are two main types of audio datasets: speech datasets and audio event/music datasets. ### Speech datasets * [AESDD](http://m3c.web.auth.gr/research/aesdd-speech-emotion-recognition/) - around 500 utterances by a diverse group of actors (over 5 actors) simlating various emotions. * [ANAD](https://www.kaggle.com/suso172/arabic-natural-audio-dataset) - 1384 recording by multiple speakers; 3 emotions: angry, happy, surprised. * [Arabic Speech Corpus](http://en.arabicspeechcorpus.com/) - The Arabic Speech Corpus (1.5 GB) is a Modern Standard Arabic (MSA) speech corpus for speech synthesis. The corpus contains phonetic and orthographic transcriptions of more than 3.7 hours of MSA speech aligned with recorded speech on the phoneme level. The annotations include word stress marks on the individual phonemes. * [ASR datasets](https://github.com/robmsmt/ASR_Audio_Data_Links) - A list of publically available audio data that anyone can download for ASR or other speech activities * [AudioMNIST](https://github.com/soerenab/AudioMNIST) - The dataset consists of 30000 audio samples of spoken digits (0-9) of 60 different speakers * [Awesome_Diarization](https://github.com/jim-schwoebel/awesome-diarization) - A curated list of awesome Speaker Diarization papers, libraries, datasets, and other resources. * [BAVED](https://www.kaggle.com/a13x10/basic-arabic-vocal-emotions-dataset) - 1935 recording by 61 speakers (45 male and 16 female). * [CaFE](https://www.gel.usherbrooke.ca/audio/cafe.htm) - 6 different sentences by 12 speakers (6 fmelaes + 6 males). * [Common Voice](https://commonvoice.mozilla.org/en/datasets) - Common Voice is Mozilla's initiative to help teach machines how real people speak. 12GB in size; spoken text based on text from a number of public domain sources like user-submitted blog posts, old books, movies, and other public speech corpora. * [CHIME](https://archive.org/details/chime-home) - This is a noisy speech recognition challenge dataset (~4GB in size). The dataset contains real simulated and clean voice recordings. Real being actual recordings of 4 speakers in nearly 9000 recordings over 4 noisy locations, simulated is generated by combining multiple environments over speech utterances and clean being non-noisy recordings. * [Coswara](https://github.com/iiscleap/Coswara-Data) - A database that contains respiratory sounds, namely, cough, breath, and speech of healthy and COVID-19 positive individuals. * [CMU-MOSEI](https://www.amir-zadeh.com/datasets) - 65 hours of annotated video from more than 1000 speakers and 250 topics; 6 Emotion (happiness, sadness, anger,fear, disgust, surprise) + Likert scale. * [CMU-MOSI](https://www.amir-zadeh.com/datasets) - 2199 opinion utterances with annotated sentiment; Sentiment annotated between very negative to very positive in seven Likert steps. * [CMU Wilderness](http://festvox.org/cmu_wilderness/) - (noncommercial) - not available but a great speech dataset many accents reciting passages from the Bible. * [CREMA-D](https://github.com/CheyneyComputerScience/CREMA-D) - CREMA-D is a data set of 7,442 original clips from 91 actors. These clips were from 48 male and 43 female actors between the ages of 20 and 74 coming from a variety of races and ethnicities (African America, Asian, Caucasian, Hispanic, and Unspecified). * [DAPS Dataset](https://archive.org/details/daps_dataset) - DAPS consists of 20 speakers (10 female and 10 male) reading 5 excerpts each from public domain books (which provides about 14 minutes of data per speaker). * [Deep Clustering Dataset](https://www.merl.com/demos/deep-clustering) - Training deep discriminative embeddings to solve the cocktail party problem. * [DEMoS](https://zenodo.
Excerpt of 23,438 characters
Read on GitHub42
2
Fabian-Robert StΓΆter Β· Audioshake Β· Germany
1
1
Would you bet a product on this? Bounded 0β100 and slow moving.
matched fp:609b0018aa8b49fa, topic:dataset, topic:datasets, readme:dataset