Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Core Engine of Singing Voice Conversion & Singing Voice Clone
| Date | Stars |
|---|---|
| 2026-07-31 | 2863 |
| 2026-08-03 | 2863 |
| 2026-08-05 | 2863 |
| 2026-08-06 | 2863 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<div align="center">
<h1> Variational Inference with adversarial learning for end-to-end Singing Voice Conversion based on VITS </h1>
[](https://huggingface.co/spaces/maxmax20160403/sovits5.0)
<img alt="GitHub Repo stars" src="https://img.shields.io/github/stars/PlayVoice/so-vits-svc-5.0">
<img alt="GitHub forks" src="https://img.shields.io/github/forks/PlayVoice/so-vits-svc-5.0">
<img alt="GitHub issues" src="https://img.shields.io/github/issues/PlayVoice/so-vits-svc-5.0">
<img alt="GitHub" src="https://img.shields.io/github/license/PlayVoice/so-vits-svc-5.0">
[中文文档](./README_ZH.md)
The tree [bigvgan-mix-v2](https://github.com/PlayVoice/whisper-vits-svc/tree/bigvgan-mix-v2) has good audio quality
The tree [RoFormer-HiFTNet](https://github.com/PlayVoice/whisper-vits-svc/tree/RoFormer-HiFTNet) has fast infer speed
No More Upgrade
</div>
- This project targets deep learning beginners, basic knowledge of Python and PyTorch are the prerequisites for this project;
- This project aims to help deep learning beginners get rid of boring pure theoretical learning, and master the basic knowledge of deep learning by combining it with practices;
- This project does not support real-time voice converting; (need to replace whisper if real-time voice converting is what you are looking for)
- This project will not develop one-click packages for other purposes;

- A minimum VRAM requirement of 6GB for training
- Support for multiple speakers
- Create unique speakers through speaker mixing
- It can even convert voices with light accompaniment
- You can edit F0 using Excel
https://github.com/PlayVoice/so-vits-svc-5.0/assets/16432329/6a09805e-ab93-47fe-9a14-9cbc1e0e7c3a
Powered by [@ShadowVap](https://space.bilibili.com/491283091)
## Model properties
| Feature | From | Status | Function |
| :--- | :--- | :--- | :--- |
| whisper | OpenAI | ✅ | strong noise immunity |
| bigvgan | NVIDA | ✅ | alias and snake | The formant is clearer and the sound quality is obviously improved |
| natural speech | Microsoft | ✅ | reduce mispronunciation |
| neural source-filter | Xin Wang | ✅ | solve the problem of audio F0 discontinuity |
| pitch quantization | Xin Wang | ✅ | quantize the F0 for embedding |
| speaker encoder | Google | ✅ | Timbre Encoding and Clustering |
| GRL for speaker | Ubisoft |✅ | Preventing Encoder Leakage Timbre |
| SNAC | Samsung | ✅ | One Shot Clone of VITS |
| SCLN | Microsoft | ✅ | Improve Clone |
| Diffusion | HuaWei | ✅ | Improve sound quality |
| PPG perturbation | this project | ✅ | Improved noise immunity and de-timbre |
| HuBERT perturbation | this project | ✅ | Improved noise immunity and de-timbre |
| VAE perturbation | this project | ✅ | Improve sound quality |
| MIX encoder | this project | ✅ | Improve conversion stability |
| USP infer | this project | ✅ | Improve conversion stability |
| HiFTNet | Columbia University | ✅ | NSF-iSTFTNet for speed up |
| RoFormer | Zhuiyi Technology | ✅ | Rotary Positional Embeddings |
due to the use of data perturbation, it takes longer to train than other projects.
**USP : Unvoice and Silence with Pitch when infer**

## Why mix

## Plug-In-Diffusion

## Setup Environment
1. Install [PyTorch](https://pytorch.org/get-started/locally/).
2. Install project dependencies
```shell
pip install -i https://pypi.tuna.tsinghua.edu.cn/simple -r requirements.txt
```
**Note: whisper is already built-in, do not insExcerpt of 20,496 characters
Read on GitHubMaxMax2016 · UESTC
356
57
12
9
6
5
4
3
3
1
1
1
1
Stardust·减 · 39 AI Inc.
1
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:3a54087d5332dbac, llm:Topics: singing-voice-conversion, singing-voice-clone, svc, sovits, vits, vits2, diffusion-svc; description: 'Core Engine of Singing Voice Conversion & Singing Voice Clone'
matched fp:3a54087d5332dbac, llm:Topics: singing-voice-conversion, singing-voice-clone, svc, sovits, vits, vits2, diffusion-svc; description: 'Core Engine of Singing Voice Conversion & Singing Voice Clone'
matched fp:3a54087d5332dbac, llm:Topics: singing-voice-conversion, singing-voice-clone, svc, sovits, vits, vits2, diffusion-svc; description: 'Core Engine of Singing Voice Conversion & Singing Voice Clone'