rom1504/img2dataset
quality grade C, 56 out of 100Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
- stars
- 4.5k
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Published datasets, dataset tooling and corpora for training and evaluation.
Signals: dataset, datasets, corpus, training-data, open-data
190 results
Easily turn large sets of image urls to an image dataset. Can download, resize and package 100M urls in 20h on one machine.
The Python library for names.
Large-scale Pre-training Corpus for Chinese 100G 中文预训练语料
A multilingual dialog corpus
Language Understanding Evaluation benchmark for Chinese: datasets, baselines, pre-trained models,corpus and leaderboard
中文医学NLP公开资源整理:术语集/语料库/词向量/预训练模型/知识图谱/命名实体识别/QA/信息抽取/模型/论文/etc
Fully open data curation for reasoning models
FMA: A Dataset For Music Analysis
🕵 Artificial Intelligence for social control of public administration | **This repository does not receive frequent updates. Check out the README**
Papers and Datasets about Point Cloud.
A simple PyTorch Implementation of Generative Adversarial Networks, focusing on anime face drawing.
A (PyTorch) imbalanced dataset sampler for oversampling low frequent classes and undersampling high frequent ones.
A central, open resource for data and tools related to chain-of-thought reasoning in large language models. Developed @ Samwald research group: https://samwald.info/
:helicopter: 保险行业语料库,聊天机器人
A powerful tool for creating datasets for LLM fine-tuning 、RAG and Eval
We introduce the Audio Logical Reasoning (ALR) dataset, consisting of 6,446 text-audio annotated samples specifically designed for complex reasoning tasks. Building on this resource, we propose SoundMind, a rule-based reinforcement learning (RL) algorithm tailored to endow audio language models (ALMs) with deep bimodal reasoning abilities.
🔥3D点云目标检测&语义分割(深度学习)-SOTA方法,代码,论文,数据集等
Human Trajectory Prediction Dataset Benchmark (ACCV 2020)
A list of Twitter datasets and related resources.
Dataset of Linus Torvalds' rants classified by negativity using sentiment analysis
Github Language Statistics
Asynchronous Multiple LiDAR-Inertial Odometry using Point-wise Inter-LiDAR Uncertainty Propagation
UrbanLoco: A Full Sensor Suite Dataset for Mapping and Localization in Urban Scenes
UrbanNav: an Open-Sourcing Localization Data Collected in Asian Urban Canyons, Including Tokyo and Hong Kong
24,539 repositories in the index in total.