Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
YSDA course in Speech Processing.
| Date | Stars |
|---|---|
| 2026-07-24 | 343 |
| 2026-07-25 | 343 |
| 2026-07-28 | 343 |
| 2026-07-30 | 343 |
| 2026-08-06 | 343 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# YSDA Speech Processing Course
- Materials for each week are in ./week* folders
## Course program
- Week 1: [Slides](https://docs.google.com/presentation/d/1MS_mj4TnXX5emzyulBwElyLLlnw2IG7ZoshXUtK84_8/edit?usp=sharing) | [Lecture](https://disk.yandex.ru/d/T6ZyKv_7bo_lUg) | [Seminar](https://disk.yandex.ru/i/QSG30iIUF3648g)
- Lecture: Intro to Digital Signal Processing (DSP)
- Seminar: Implement DSP pipeline
- Homework (5pt): Implement mel-spectrogram transformations
- Week 2: [Slides](https://docs.google.com/presentation/d/1OKeUW8f7SrKNh5W8LG4rnebqrtPBTibg1sn9-f9fq1I/edit?usp=sharing) | [Lecture](https://disk.yandex.ru/i/g8lCY4vr0voh5Q) | [Seminar](https://disk.yandex.ru/d/XxjSFpDF7V-N0Q)
- Lecture: Introduction to speech NN discriminative models. Voice Activity Detection (VAD) and Sound Event Detection (SED) tasks
- Seminar: Train VAD models, intro to the homework
- Homework (15pt): Train SED models; (3pt bonus) SED models vibecoding
- Week 3: [Slides](https://docs.google.com/presentation/d/1iBHjRY3W0kHvBUhFl5Vf76DK9glwwLmAZvUJ7JzcZyE/edit?usp=sharingg) | [Lecture](https://disk.yandex.ru/i/AgmlIrd0_KZdEA) | [Seminar](https://disk.yandex.ru/i/DWI1hzJ5eZMBLw)
- Lecture: Keyword Spotting and Speech Biometrics tasks
- Seminar: Train Biometrics model and look at embeddings
- Homework (20pt): Train Biometrics model ECAPA-TDNN with contrastive loss
- Week 4 [Slides](https://docs.google.com/presentation/d/1TGAaI4uHM1pCkP8sS-ExIqJxaQyXOfEesOr4dIKYlfY) | [Lecture+Seminar](https://disk.yandex.ru/i/nY4RniGGOxaPFg)
- Lecture: Speech Recognition I
- Seminar: CTC forward-backward, soft alignment
- Homework (10pt): CTC/RNN-T decoding, RNN-T forward-backward
- Week 5 [Slides](https://docs.google.com/presentation/d/10rdROxSQ0N3Kctei8B006We1Ra1n0dS6Hn4PZzJ52gs/edit?usp=sharing) | [Lecture](https://disk.yandex.ru/i/tWqBUIUlQPe22A) | [Seminar](https://disk.yandex.ru/i/aeA_qoYTsFElKQ)
- Lecture: Pretraining in Speech Recognition
- Seminar: Speech Pretraining - quantization and losses
- Homework (5pt): Speech Pretraining
- Week 6 [Slides](https://docs.google.com/presentation/d/1AA8M8-kSMlX4F4qrCnE_COp_vvyL1pO9Gv77s2Rj5dg/edit?usp=sharing) | [Lecture](https://disk.yandex.ru/i/-dfJ7Hnm5VgMCw)
- Leсture: ASR Inference
- Homework (5pt): Implement streaming inference
- Week 7 [Slides](https://docs.google.com/presentation/d/12w0YcIQKkEFElzPJMj1xmxg3LZft-ExP0KilP2izSqA/edit?usp=sharing) | [Lecture](https://disk.yandex.ru/i/nLLBPCDfLwVLtg)
- Lecture: Intro to TTS. Normalisation, Tasks, Metrics
- Week 8 [Slides](https://docs.google.com/presentation/d/1L_bv9B8mvKY7PO86pSnKQQdIqzzrYMnArNi0Nq1rsxw/edit?usp=sharing) | [Lecture](https://disk.yandex.ru/i/ojvxvVZAjd9GQg)
- Lecture: Tacotron2, FastPitch, HiFiGAN
- Seminar (5pt): Implement pitch estimation
- Homework (10pt): Implement FastPitch
- Week 9 [Slides](https://docs.google.com/presentation/d/1MMNb_FKuRE1rq2f1fCrVjnFewmH9qpB1835nWoibVTI/edit?usp=sharing) | [Lecture](https://disk.yandex.ru/i/ZavO7g1SPtSB3Q) | [Seminar](https://disk.yandex.ru/i/LK0jCg2E3r-5VA)
- Lecture: Quantisation and Neural codecs
- Seminar: Implement several quantisation methods
- Homework (10pt): Implement more advanced quantisations and audio codecs
- Week 10 [Slides](https://docs.google.com/presentation/d/17_uJGrqkEsFrNspt0lZHreOi6gSKhFREvhH3GXfn-Mk/edit?usp=sharing) | [Lecture](https://disk.yandex.ru/d/4Uv2tOsWcwzDqA)
- Lecture: Diffusions and transformers for voice cloning
- Homework (10pt): Implement slow-fast transformer inference
- Week 11 [Slides](https://docs.google.com/presentation/d/1zfo-vvYHKIFZniDuh-afKif8v8O3272uLSAXm7cFyL8/edit?usp=sharing) | [Lecture](https://disk.yandex.ru/i/JddYxXDCXiHo-Q)
- Lecture: Spoken dialogue models
- Week 13 [Slides](https://docs.google.com/presentation/d/1eNkZMHYDuFv9unzLW7bidDpkWJythQsICo4wgCyeoA0/edit?usp=sharing) | [Lecture](https://disk.yandex.ru/i/oT5oQQPXSLnlYA)
Excerpt of 13,817 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:4a12ee4668da4a2e, topic:tts, topic:asr, readme:speech recognition
matched fp:4a12ee4668da4a2e, name:course, desc:course, readme:course