virattt/financial-datasets
quality grade D, 41 out of 100Financial datasets for LLMs 🧪
- stars
- 429
- stars gained this week
- —this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Published datasets, dataset tooling and corpora for training and evaluation.
Signals: dataset, datasets, corpus, training-data, open-data
186 results
Financial datasets for LLMs 🧪
This repository contains a summary of knowledge cut-off dates for various large language models (LLMs), such as GPT, Claude, Gemini, Llama, and more.
[Neurips'24 Spotlight] Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning
[NeurIPS 2023 Datasets and Benchmarks Track] LAMM: Multi-Modal Large Language Models and Applications as AI Agents
[ICLR 2024] Mol-Instructions: A Large-Scale Biomolecular Instruction Dataset for Large Language Models
A curated collection of papers, datasets, and resources on Scientific Datasets and Large Language Models (LLMs)
This repository contains related work, benchmarks and datasets for the paper "Large Language Models in Finance (FinLLMs)".
Formerly known as code.google.com/p/1-billion-word-language-modeling-benchmark
[NIPS2023] Code and Model for VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and Dataset
A comprehensive list of PAPERS, CODEBASES, and, DATASETS on Decision Making using Foundation Models including LLMs and VLMs.
Dr. Zero Self-Evolving Search Agents without Training Data
Datasets for Instruction Tuning of Large Language Models
Github repository for "RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models"
Datasets for Knowledge Graph Completion with textual information about the entities
农业领域知识图谱的构建,包括数据爬取(百度百科)、数据分类、利用结构化数据生成三元组、非结构化数据的分句(LTP),分词(jieba),命名实体识别(LTP)、基于依存句法分析(主谓关系等)的关系抽取和利用neo4j生成可视化知识图谱
Unlock the Power of LLM: Explore These Datasets to Train Your Own ChatGPT!
Generic rag framework to apply the power of LLMs on any given dataset
A collection of benchmarks and datasets for evaluating LLM.
No description
A curated list of medical LLMs, multimodal systems, datasets, benchmarks, and more. 🏥
Summarize existing representative LLMs text datasets.
A minimal yet resourceful implementation of diffusion models (along with pretrained models + synthetic images for nine datasets)
心理健康大模型 (LLM x Mental Health), Pre & Post-training & Dataset & Evaluation & Depoly & RAG, with InternLM / Qwen / Baichuan / DeepSeek / Mixtral / LLama / GLM series models
Toolkit for linearizing PDFs for LLM datasets/training
24,523 repositories in the index in total.