CLUEbenchmark/CLUEDatasetSearch
quality grade D, 42 out of 100搜索所有中文NLP数据集,附常用英文NLP数据集
- stars
- 4.5k
- stars gained this week
- -1this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Tokenization, parsing, classical NLP pipelines, translation and information extraction.
Signals: nlp, natural-language-processing, tokenizer, named-entity-recognition, text-classification, machine-translation, sentiment-analysis, spacy
637 results
搜索所有中文NLP数据集,附常用英文NLP数据集
all kinds of text classification models and more with deep learning
SharpToken is a C# library for tokenizing natural language text. It's based on the tiktoken Python library and designed to be fast and accurate.
Contains source-code for viewers following along with my Beginners Guide To Building Interpreters series on my Youtube Channel.
🌸 Possibly the smallest Lua compiler ever
🛥 Vaporetto: Very accelerated pointwise prediction based tokenizer
Bitextor generates translation memories from multilingual websites
Text2Text Language Modeling Toolkit
NLP tokenizers written in Go language
Fast and customizable text tokenization library with BPE and SentencePiece support
Ready-made tokenizer library for working with GPT and tiktoken
Lex machinary for go.
Juman++ (a Morphological Analyzer Toolkit)
🎤 vibrato: Viterbi-based accelerated tokenizer
A Japanese tokenizer based on recurrent neural networks
EFFICIENT AND OPTIMIZED TOKENIZER ENGINE FOR LLM INFERENCE SERVING
A multilingual command line sentence tokenizer in Golang
CogComp's Natural Language Processing Libraries and Demos: Modules include lemmatizer, ner, pos, prep-srl, quantifier, question type, relation-extraction, similarity, temporal normalizer, tokenizer, transliteration, verb-sense, and more.
Python port of Moses tokenizer, truecaser and normalizer
High performance Chinese tokenizer with both GBK and UTF-8 charset support based on MMSEG algorithm developed by ANSI C. Completely based on modular implementation and can be easily embedded in other programs, like: MySQL, PostgreSQL, PHP, etc.
VSCode extension to highlight nested code blocks
A Cython MeCab wrapper for fast, pythonic Japanese tokenization and morphological analysis.
Tiny JavaScript tokenizer.
Ungreedy subword tokenizer and vocabulary trainer for Python, Go, C++ & Javascript
24,535 repositories in the index in total.