Ayanami0730/deep_research_bench
quality grade B, 65 out of 100DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
- stars
- 806
- stars gained this week
- —
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Tokenization, parsing, classical NLP pipelines, translation and information extraction.
Signals: nlp, natural-language-processing, tokenizer, named-entity-recognition, text-classification, machine-translation, sentiment-analysis, spacy
637 results
DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
A personal semantic search engine capable of surfacing relevant bookmarks, journal entries, notes, blogs, contacts, and more, built on an efficient document embedding algorithm and Monocle's personal search index.
Fuzzy and semantic search for captioned YouTube videos.
Semantic search for competitive programming problems
AraVec is a pre-trained distributed word representation (word embedding) open source project which aims to provide the Arabic NLP research community with free to use and powerful word embedding models.
Make function calling with LLM easier
A curated list of resources for Japanese natural language processing (NLP): Python libraries, LLMs, dictionaries, corpora, and datasets. Includes Claude Code skills to search resources.
A series of top performing Text to SQL LLMs
Analyzing Hacker News discussions from a decade ago in hindsight with LLMs
State-of-the-art LLM-based translation models.
LLM-based ontological extraction tools, including SPIRES
No description
OO for LLMs
ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Minimal, clean code for the Byte Pair Encoding (BPE) algorithm commonly used in LLM tokenization.
SoTA LLM for converting natural language questions to SQL queries
适配轻小说/Galgame的日中翻译大模型
This repository collects an extensive list of awesome papers about Story Generation / Storytelling, exclusively focusing on the era of Large Language Models (LLMs).
[中文法律大模型] DISC-LawLLM: an intelligent legal system powered by large language models (LLMs) to provide a wide range of legal services.
This is the code for the SpeechTokenizer presented in the SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models. Samples are presented on
[EMNLP 2023] Enabling Large Language Models to Generate Text with Citations. Paper: https://arxiv.org/abs/2305.14627
[EMNLP 2022] Unifying and multi-tasking structured knowledge grounding with language models
A novel method to tune language models. Codes and datasets for paper ``GPT understands, too''.
Awesome papers on Language-Model-as-a-Service (LMaaS)
24,523 repositories in the index in total.