Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Tools for handling multimodal data in machine learning projects.
| Date | Stars |
|---|---|
| 2026-07-24 | 1143 |
| 2026-07-25 | 1143 |
| 2026-07-28 | 1143 |
| 2026-07-30 | 1143 |
| 2026-07-31 | 1144 |
| 2026-08-06 | 1145 |
Today
+1 stars today
This week
+2 stars this week
This month
— stars this month
Momentum
11.0
growth rate 0.18%/day
<div align="center"> <img src="https://raw.githubusercontent.com/lhotse-speech/lhotse/master/docs/logo.png" width=376> [](https://badge.fury.io/py/lhotse) [](https://pypi.org/project/lhotse/) [](https://pepy.tech/project/lhotse) [](https://actions-badge.atrox.dev/pzelasko/lhotse/goto?ref=master) [](https://lhotse.readthedocs.io/en/latest/?badge=latest) [](https://codecov.io/gh/lhotse-speech/lhotse) [](https://github.com/psf/black) [](https://colab.research.google.com/github/lhotse-speech/lhotse-speech.github.io/blob/master/notebooks/lhotse-introduction.ipynb) [](https://arxiv.org/abs/2110.12561) </div> # Lhotse Lhotse is a Python library aiming to make multimodal (speech, audio, video, image, text) data preparation flexible and accessible to a wider community. Alongside [k2](https://github.com/k2-fsa/k2), it is a part of the next generation [Kaldi](https://github.com/kaldi-asr/kaldi) speech processing library. ## Tutorial presentations and materials - (_Interspeech 2023_) Tutorial notebook [](https://colab.research.google.com/drive/1obZjUuVwks3A4oFX3gXFtPOM2LtrPQfL?usp=sharing) - (_Interspeech 2023_) [Tutorial slides](https://livejohnshopkins-my.sharepoint.com/:p:/g/personal/mwiesne2_jh_edu/EYqRDl8cIr5BsVDxi1MOW5EBUpdqh10WFkzqixPIFM63hg?e=u3lrmL) - (_Interspeech 2021_) [Recorded lecture (3h)](https://www.youtube.com/watch?v=y6CJLFQlmhc&pp=ygUgaW50ZXJzcGVlY2ggMjAyMSBsaG90c2UgdHV0b3JpYWw%3D) ## About ### Main goals (updated for 2025) - Scale to multimodal data pipelines including audio, text, image, and video modalities. - Provide state-of-the-art dataloading algorithms such as dataset blending and efficient on-the-fly bucketing. - Handle data randomization (or de-duplication) for distributed multi-node training. - Attract a wider community to multimodal processing tasks with a **Python-centric design**. - Provide **standard data preparation recipes** for commonly used corpora. - Flexible data preparation for model training with the notion of **audio/video cuts**. - Support for efficient sequential I/O data formats such as Lhotse Shar (similar to webdataset). ### Tutorials We offer the following tutorials available in `examples` directory: - Basic complete Lhotse workflow [](https://colab.research.google.com/github/lhotse-speech/lhotse/blob/master/examples/00-basic-workflow.ipynb) - Transforming data with Cuts [](https://colab.research.google.com/github/lhotse-speech/lhotse/blob/master/examples/01-cut-python-api.ipynb) - WebDataset integration [](https://colab.research.google.com/github/lhotse-speech/lhotse/blob/master/examples/02-webdataset-integration.ipynb) - How to combine multiple datasets [](https://colab.research.google.com/github/lhotse-speech/lhotse/blob/master/examples/03-combining-datasets.ipynb) - Lhotse Shar: storage format optimized for sequential I/O and modularity [](https://colab.research.google.com/gi
Excerpt of 16,428 characters
Read on GitHubPiotr Żelasko · @NVIDIA
1.9k
352
Jan "yenda" Trmal
34
33
Fangjun Kuang · Xiaomi Corporation
26
25
Wei Kang · Xiaomi Corporation · China
22
Yifan Yang · Shanghai Jiao Tong University
21
21
21
15
15
14
Karel Vesely · Brno University of Technology · Czech Republic
12
11
Dmitrii Mukhutdinov · United Kingdom
11
Yuekai Zhang · @Nvidia · China
10
10
10
彭震东 · WeNet · China
9
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:9f87dd723a072476, topic:deep-learning, topic:pytorch
matched fp:9f87dd723a072476, topic:speech-recognition