jordipons/musicnn
quality grade D, 37 out of 100Pronounced as "musician", musicnn is a set of pre-trained deep convolutional neural networks for music audio tagging.
- stars
- 713
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
913 results
Pronounced as "musician", musicnn is a set of pre-trained deep convolutional neural networks for music audio tagging.
Guitar plugin made with JUCE that uses neural networks to emulate a tube amplifier.
Generates video game music using neural networks.
Implementation of Diffusion Convolutional Recurrent Neural Network in Tensorflow
Deep neural networks for voice conversion (voice style transfer) in Tensorflow
Recurrent neural network for audio noise reduction
Inference and training library for high-quality TTS models.
A modular voice assistant application for experimenting with state-of-the-art transcription, response generation, and text-to-speech models. Supports OpenAI, Groq, Elevanlabs, CartesiaAI, and Deepgram APIs, plus local models via Ollama. Ideal for research and development in voice technology.
Use API to call the music generation AI of suno.ai, and easily integrate it into agents like GPTs.
🎵 The Ultimate Open Source Suno Alternative - Professional UI for ACE-Step 1.5 AI Music Generation. Free, local, unlimited. Stop paying for Suno!
A fundamental toolkit designed for music, song, and audio generation
Pronunciation lexicon covering both English and Chinese languages for Automatic Speech Recognition.
Phoneme Recognition using pre-trained models Wav2vec2, HuBERT and WavLM. Throughout this project, we compared specifically three different self-supervised models, Wav2vec (2019, 2020), HuBERT (2021) and WavLM (2022) pretrained on a corpus of English speech that we will use in various ways to perform phoneme recognition for different languages with a network trained with Connectionist Temporal Classification (CTC) algorithm.
A pure python module for reading and writing kaldi ark files
Wav2Vec for speech recognition, classification, and audio classification
Python module for evaluating ASR hypotheses (e.g. word error rate, word recognition rate).
An end-to-end framework for multi-speaker transcription that jointly models who spoke, when, and what.
simple delaysum, MVDR and CGMM-MVDR
This repository is a curated list of awesome Speech Keyword Spotting (Wake-Up Word Detection).
一个基于云端语音识别的智能控制设备,类似于天猫精灵,小爱同学。采用的芯片为stm32f407,wm8978,esp8266。
AudioBench: A Universal Benchmark for Audio Large Language Models
Automatic Speech Recognition (ASR) - German
Aims to create a comprehensive voice toolkit for training, testing, and deploying speaker verification systems.
Speech recognition framework allowing powerful Python-based scripting and extension of Dragon NaturallySpeaking (DNS), Windows Speech Recognition (WSR), Kaldi and CMU Pocket Sphinx
24,524 repositories in the index in total.