Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Automatically Update Text-to-speech (TTS) Papers Daily using Github Actions (Update Every 12th hours)
| Date | Stars |
|---|---|
| 2026-07-31 | 663 |
| 2026-08-03 | 663 |
| 2026-08-06 | 663 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
## Updated on 2026.08.02
> Usage instructions: [here](./docs/README.md#usage)
> This page is modified from [here](https://github.com/Vincentqyw/cv-arxiv-daily)
<details>
<summary>Table of Contents</summary>
<ol>
<li><a href=#tts>TTS</a></li>
</ol>
</details>
## TTS
|Publish Date|Title|Authors|PDF|Code|
|---|---|---|---|---|
|**2026-07-30**|**Teffic-Audio: Tell Fact from Fiction**|Wan Lin et.al.|[2607.28351](http://arxiv.org/abs/2607.28351)|null|
|**2026-07-30**|**VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition**|Yukun Chen et.al.|[2607.27768](http://arxiv.org/abs/2607.27768)|null|
|**2026-07-29**|**Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model**|Carlos Muñoz-Romero et.al.|[2607.26742](http://arxiv.org/abs/2607.26742)|null|
|**2026-07-29**|**MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation**|Wei-Jaw Lee et.al.|[2607.26698](http://arxiv.org/abs/2607.26698)|**[link](https://github.com/YatingMusic/MPEcho)**|
|**2026-07-28**|**Extracting Voice Styles from Frozen TTS Models via Gradient-Based Inverse Optimization**|Gyeongmin Kim et.al.|[2607.25351](http://arxiv.org/abs/2607.25351)|null|
|**2026-07-27**|**Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis**|Yifan Hu et.al.|[2607.24430](http://arxiv.org/abs/2607.24430)|**[link](https://github.com/walker-hyf/FacialTalker)**|
|**2026-07-27**|**Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection**|Mingrui Liang et.al.|[2607.23961](http://arxiv.org/abs/2607.23961)|null|
|**2026-07-27**|**Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm**|Bajian Xiang et.al.|[2607.23938](http://arxiv.org/abs/2607.23938)|null|
|**2026-07-26**|**Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers**|Dongseong Hwang et.al.|[2607.23811](http://arxiv.org/abs/2607.23811)|null|
|**2026-07-25**|**Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English**|Ivan Kukanov et.al.|[2607.23027](http://arxiv.org/abs/2607.23027)|null|
|**2026-07-24**|**Cycles of Discourse, Speech Dysfluency, and Active Inference**|Thomas Parr et.al.|[2607.22180](http://arxiv.org/abs/2607.22180)|null|
|**2026-07-24**|**VisionPulse: A Virtual Reality System Enabling Accessible Discovery and Navigation for Blind and Low Vision Users**|Samuel Martin et.al.|[2607.21944](http://arxiv.org/abs/2607.21944)|null|
|**2026-07-23**|**Toward Interpretable Speech Deepfake Detection using Artifact-Specific Experts and Calibrated Detection Scores**|Viola Negroni et.al.|[2607.21127](http://arxiv.org/abs/2607.21127)|null|
|**2026-07-23**|**Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs**|Muyang Du et.al.|[2607.21042](http://arxiv.org/abs/2607.21042)|null|
|**2026-07-22**|**ETPDesigner: Multi-Agent Orchestration for Interactive Multimodal Electronic Theater Program**|Mengtian Li et.al.|[2607.19947](http://arxiv.org/abs/2607.19947)|null|
|**2026-07-22**|**Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models**|Pengchao Feng et.al.|[2607.19932](http://arxiv.org/abs/2607.19932)|null|
|**2026-07-22**|**StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis**|Kaicheng Luo et.al.|[2607.19859](http://arxiv.org/abs/2607.19859)|null|
|**2026-07-19**|**Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer**|Sivateja Trikutam et.al.|[2607.18662](http://arxiv.org/abs/2607.18662)|null|
|**2026-07-21**|**CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses**|Sajid Fardin Dipto et.al.|[2607.18629](http://arxiv.org/abs/2607.18629)|null|
|**2026-07-17**|**A Situational Speech Synthesizer for Yoruba: System Design, PhonoloExcerpt of 370,518 characters
Read on GitHub1.3k
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:a5619399aa35440d, desc:text-to-speech