Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A powerful tool for creating datasets for LLM fine-tuning 、RAG and Eval
| Date | Stars |
|---|---|
| 2026-07-24 | 14692 |
| 2026-07-25 | 14698 |
| 2026-07-28 | 14698 |
| 2026-07-30 | 14698 |
| 2026-07-31 | 14729 |
| 2026-08-06 | 14729 |
Today
— stars today
This week
+31 stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.21%/day
<div align="center">  <img alt="GitHub Repo stars" src="https://img.shields.io/github/stars/ConardLi/easy-dataset"> <img alt="GitHub Downloads (all assets, all releases)" src="https://img.shields.io/github/downloads/ConardLi/easy-dataset/total"> <img alt="GitHub Release" src="https://img.shields.io/github/v/release/ConardLi/easy-dataset"> <img src="https://img.shields.io/badge/license-AGPL--3.0-green.svg" alt="AGPL 3.0 License"/> <img alt="GitHub contributors" src="https://img.shields.io/github/contributors/ConardLi/easy-dataset"> <img alt="GitHub last commit" src="https://img.shields.io/github/last-commit/ConardLi/easy-dataset"> <a href="https://arxiv.org/abs/2507.04009v1" target="_blank"> <img src="https://img.shields.io/badge/arXiv-2507.04009-b31b1b.svg" alt="arXiv:2507.04009"> </a> <a href="https://trendshift.io/repositories/13944" target="_blank"><img src="https://trendshift.io/api/badge/repositories/13944" alt="ConardLi%2Feasy-dataset | Trendshift" style="width: 250px; height: 55px;" width="250" height="55"/></a> **A powerful tool for creating fine-tuning datasets for Large Language Models** [简体中文](./README.zh-CN.md) | [English](./README.md) | [Türkçe](./README.tr.md) [Features](#features) • [Quick Start](#local-run) • [Documentation](https://docs.easy-dataset.com/ed/en) • [Contributing](#contributing) • [License](#license) If you like this project, please give it a Star⭐️, or buy the author a coffee => [Donate](./public/imgs/aw.jpg) ❤️! </div> ## Overview Easy Dataset is an application specifically designed for building large language model (LLM) datasets. It features an intuitive interface, along with built-in powerful document parsing tools, intelligent segmentation algorithms, data cleaning and augmentation capabilities. The application can convert domain-specific documents in various formats into high-quality structured datasets, which are applicable to scenarios such as model fine-tuning, retrieval-augmented generation (RAG), and model performance evaluation.  ## News 🎉🎉 Easy Dataset Version 1.7.0 launches brand-new evaluation capabilities! You can effortlessly convert domain-specific documents into evaluation datasets (test sets) and automatically run multi-dimensional evaluation tasks. Additionally, it comes with a human blind test system, enabling you to easily meet needs such as vertical domain model evaluation, post-fine-tuning model performance assessment, and RAG recall rate evaluation. Tutorial: [https://www.bilibili.com/video/BV1CRrVB7Eb4/](https://www.bilibili.com/video/BV1CRrVB7Eb4/) ## Features ### 📄 Document Processing & Data Generation - **Intelligent Document Processing**: Supports PDF, Markdown, DOCX, TXT, EPUB and more formats with intelligent recognition - **Intelligent Text Splitting**: Multiple splitting algorithms (Markdown structure, recursive separators, fixed length, code-aware chunking), with customizable visual segmentation - **Intelligent Question Generation**: Auto-extract relevant questions from text segments, with question templates and batch generation - **Domain Label Tree**: Intelligently builds global domain label trees based on document structure, with auto-tagging capabilities - **Answer Generation**: Uses LLM API to generate comprehensive answers and Chain of Thought (COT), with AI optimization - **Data Cleaning**: Intelligent text cleaning to remove noise and improve data quality ### 🔄 Multiple Dataset Types - **Single-Turn QA Datasets**: Standard question-answer pairs for basic fine-tuning - **Multi-Turn Dialogue Datasets**: Customizable roles and scenarios for conversational format - **Image QA Datasets**: Generate visual QA data from images, with multiple import methods (directory, PDF, ZIP) - **Data Distillation**: Generate label trees and questions directly from domain topics without uploading documents ### 📊 Model Evaluation System - **Evaluation Datasets**: Generate true/false, single-choice, mul
Excerpt of 12,368 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:41b973401aff3979, topic:dataset, name:dataset, readme:dataset
matched fp:41b973401aff3979, topic:fine-tuning, desc:fine-tuning, readme:fine-tuning
matched fp:41b973401aff3979, topic:llm
matched fp:41b973401aff3979, topic:rag, readme:retrieval-augmented generation, readme:retrieval augmented