huggingface/dataset-viewer
quality grade B, 66 out of 100Backend that powers the dataset viewer on Hugging Face dataset pages through a public API.
- stars
- 900
- stars gained this week
- +1this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Published datasets, dataset tooling and corpora for training and evaluation.
Signals: dataset, datasets, corpus, training-data, open-data
190 results
Backend that powers the dataset viewer on Hugging Face dataset pages through a public API.
EEGdenoiseNet, a benchmark dataset, that is suited for training and testing deep learning-based EEG denoising models, as well as for comparing the performance across different models.
Custom APRL machine-learning algorithm + a curated ML knowledge base of 923 papers, lectures, and explainers. NumPy classifier/regressor, tests, benchmarks, provenance, Obsidian, and agent-ready Markdown.
Neural-network consensus polishing and variant calling for Oxford Nanopore sequencing data
official implementation of the spatial-temporal attention neural network (STANet) for remote sensing image change detection
Official repo of Toucan: Synthesizing 1.5M Tool-Agentic Data from Real-World MCP Environments
Matrix (Multi-Agent daTa geneRation Infra and eXperimentation framework) is a versatile engine for multi-agent conversational data generation.
ToolQA, a new dataset to evaluate the capabilities of LLMs in answering challenging questions with external tools. It offers two levels (easy/hard) across eight real-life scenarios.
A ChatGPT(GPT-3.5) & GPT-4 Workload Trace to Optimize LLM Serving Systems
Create an open source toy dataset for finetuning LLMs with reasoning abilities
Do-Not-Answer: A Dataset for Evaluating Safeguards in LLMs
A curated list of popular Datasets, Models and Papers for LLMs in Medical/Healthcare
Financial datasets for LLMs 🧪
This repository contains a summary of knowledge cut-off dates for various large language models (LLMs), such as GPT, Claude, Gemini, Llama, and more.
[Neurips'24 Spotlight] Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
[NeurIPS 2023 Datasets and Benchmarks Track] LAMM: Multi-Modal Large Language Models and Applications as AI Agents
[ICLR 2024] Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language Models
A curated collection of papers, datasets, and resources on Scientific Datasets and Large Language Models (LLMs)
This repository contains related work, benchmarks and datasets for the paper "Large Language Models in Finance (FinLLMs)".
Formerly known as code.google.com/p/1-billion-word-language-modeling-benchmark
[NIPS2023] Code and Model for VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset
A comprehensive list of PAPERS, CODEBASES, and, DATASETS on Decision Making using Foundation Models including LLMs and VLMs.
Dr. Zero Self-Evolving Search Agents without Training Data
Datasets for Instruction Tuning of Large Language Models
24,537 repositories in the index in total.