mlabonne/llm-datasets
quality grade D, 49 out of 100Curated list of datasets and tools for post-training.
- stars
- 4.7k
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Published datasets, dataset tooling and corpora for training and evaluation.
Signals: dataset, datasets, corpus, training-data, open-data
186 results
Curated list of datasets and tools for post-training.
tracking papers, datasets, and models of "large language model (LLM) for time series"
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Language Models [NeurIPS 2024 Datasets and Benchmarks Track]
Tools for curating biomedical training data for large-scale language modeling
The RedPajama-Data repository contains code for preparing large datasets for training large language models.
DataComp for Language Models
ChainKnowledgeGraph, 产业链知识图谱包括A股上市公司、行业和产品共3类实体,包括上市公司所属行业关系、行业上级关系、产品上游原材料关系、产品下游产品关系、公司主营产品、产品小类共6大类。 上市公司4,654家,行业511个,产品95,559条、上游材料56,824条,上级行业480条,下游产品390条,产品小类52,937条,所属行业3,946条。
史上最大规模1.4亿知识图谱数据免费下载,知识图谱,通用知识图谱,融合了两千五百多万的实体,拥有亿级别的实体属性关系。
史上最大规模1.4亿中文知识图谱开源下载
A minimal PyTorch implementation of probabilistic diffusion models for 2D datasets.
List of datasets to apply stats/machine learning/technology to the world of social good.
A list of papers, blogs, datasets and software in the field of lifelong/continual machine learning
No description
Extension to edit dataset captions for SD web UI by AUTOMATIC1111
The AMLSim project is intended to provide a multi-agent based simulator that generates synthetic banking transaction data together with a set of known money laundering patterns - mainly for the purpose of testing machine learning models and graph algorithms. We welcome you to enhance this effort since the data set related to money laundering is critical to advance detection capabilities of money laundering activities.
汉字拆字库,可以将汉字拆解成偏旁部首,在机器学习中作为汉字的字形特征 | Hanzi Decomposition Library allows Chinese characters to be broken down into radicals and components, which can be used as character shape features in machine learning.
Explore my diverse collection of projects showcasing machine learning, data analysis, and more. Organized by project, each directory contains code, datasets, documentation, and resources. Dive in, to discover insights and techniques in data science. Reach out for collaborations and feedback.
Models to perform neural summarization (extractive and abstractive) using machine learning transformers and a tool to convert abstractive summarization datasets to the extractive task.
A list of databases, datasets and books/handbooks where you can find materials properties for machine learning applications.
🌳 A curated list of ground-truth forest datasets for the machine learning and forestry community.
The NES Music Database: use machine learning to compose music for the Nintendo Entertainment System!
source{d} datasets ("big code") for source code analysis and machine learning on source code
CircuitNet: An Open-Source Dataset for Machine Learning Applications in Electronic Design Automation (EDA)
An all-atom protein structure dataset for machine learning.
24,523 repositories in the index in total.