2026.01.13 - Soprano-Factory released! You can now train/fine-tune your own Soprano models.
Soprano-Factory is a 600-LOC training script allowing anyone to train/fine-tune Soprano models with their own data, on their own hardware. You can use Soprano-Factory to add new voices, styles, and languages to the original Soprano model.
Soprano is an ultra‑lightweight, on-device text‑to‑speech (TTS) model designed for expressive, high‑fidelity speech synthesis at unprecedented speed. Soprano was designed with the following features:
- Up to 20x real-time generation on CPU and 2000x real-time on GPU
- Lossless streaming with <250 ms latency on CPU, <15 ms on GPU
- <1 GB memory usage with a compact 80M parameter architecture
- Infinite generation length with automatic text splitting
- Highly expressive, crystal clear audio generation at 32kHz
- Widespread support for CUDA, CPU, and MPS devices on Windows, Linux, and Mac
- Supports WebUI, CLI, and OpenAI-compatible endpoint for easy and production-ready inference
git clone https://github.com/ekwek1/soprano-factory.git
cd soprano-factory
pip install -r requirements.txtIf using Windows you may need to reinstall Pytorch to have CUDA support.
Soprano-Factory expects input data in LJSpeech format. Please see example_dataset for how to structure your dataset. The wav files can be in any sample rate; they will be automatically resampled to 32 kHz.
python generate_dataset.py --input-dir path/to/files
Args:
--input-dir: Path to directory of LJSpeech-style dataset. If none is provided this defaults to the provided example dataset.
python train.py --input-dir path/to/files --save-dir path/to/weights
Args:
--input-dir: Path to directory of LJSpeech-style dataset. If none is provided this defaults to the provided example dataset.
--save-dir: Path to directory to save weights
Once the model has been trained, you can then use the custom weights in the Soprano repository! Just pass your save directory into model_path.
I did not originally design Soprano with finetuning in mind. As a result, I cannot guarantee that you will see good results after training.
This project is licensed under the Apache-2.0 license. See LICENSE for details.
