KevinWang676/Bark-Voice-Cloning
quality grade C, 59 out of 100Bark Voice Cloning and Voice Cloning for Chinese Speech
- stars
- 2.9k
- stars gained this week
- -1this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Speech recognition, text-to-speech, voice cloning, music generation and audio processing.
Signals: speech-recognition, text-to-speech, tts, stt, asr, whisper, voice-cloning, speech-synthesis
913 results
Bark Voice Cloning and Voice Cloning for Chinese Speech
Apply diffusion models using the new Hugging Face diffusers package to synthesize music instead of images.
Score-based Generative Models (Diffusion Models) for Speech Enhancement and Dereverberation
Generate subtitles for your videos with secure, on-device machine learning models.
Machine Learning applied to sound
Analyze videos using LLMs, Computer Vision and Automatic Speech Recognition
Open dubbing is an AI dubbing system which uses machine learning models to automatically translate and synchronize audio dialogue into different languages.
TorchSig is an open-source signal processing machine learning toolkit based on the PyTorch data handling pipeline.
Collection of notebooks and scripts related to audio processing and machine learning.
Automatically synchronize subtitles with audio using machine learning
pytorch implementation of "Deep Learning-Enabled Semantic Communication Systems with Task-Unaware Transmitter and Dynamic Data"
Deep Learning Networks for Real Time Guitar Effect Emulation using WaveNet with PyTorch
Some Code for Master Thesis - Research on Deep Learning Based Modulation Recognition Technologies
Implementation of "MOSNet: Deep Learning based Objective Assessment for Voice Conversion"
Speech Recognition with the Caffe deep learning framework, migrating to
Code for YouTube series: Deep Learning for Audio Classification
State of the Art of Music Generation with Deep Learning and AI
This is an open source project (formerly named Listen, Attend and Spell - PyTorch Implementation) for end-to-end ASR implemented with Pytorch, the well known deep learning toolkit.
State-of-the-art deep learning based audio codec supporting both mono 24 kHz audio and stereo 48 kHz audio.
Audio Plugin for Audio to MIDI transcription using deep learning.
Resources on Music Generation with Deep Learning
A python package to analyze and compare voices with deep learning
Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.
PyTorch implementations of neural network models for keyword spotting
24,524 repositories in the index in total.