daac-tools/vaporetto
quality grade B, 67 out of 100🛥 Vaporetto: Very accelerated pointwise prediction based tokenizer
- stars
- 297
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Tokenization, parsing, classical NLP pipelines, translation and information extraction.
Signals: nlp, natural-language-processing, tokenizer, named-entity-recognition, text-classification, machine-translation, sentiment-analysis, spacy
637 results
🛥 Vaporetto: Very accelerated pointwise prediction based tokenizer
Bitextor generates translation memories from multilingual websites
Text2Text Language Modeling Toolkit
NLP tokenizers written in Go language
Fast and customizable text tokenization library with BPE and SentencePiece support
Ready-made tokenizer library for working with GPT and tiktoken
Lex machinary for go.
Juman++ (a Morphological Analyzer Toolkit)
🎤 vibrato: Viterbi-based accelerated tokenizer
A Japanese tokenizer based on recurrent neural networks
EFFICIENT AND OPTIMIZED TOKENIZER ENGINE FOR LLM INFERENCE SERVING
A multilingual command line sentence tokenizer in Golang
CogComp's Natural Language Processing Libraries and Demos: Modules include lemmatizer, ner, pos, prep-srl, quantifier, question type, relation-extraction, similarity, temporal normalizer, tokenizer, transliteration, verb-sense, and more.
Python port of Moses tokenizer, truecaser and normalizer
High performance Chinese tokenizer with both GBK and UTF-8 charset support based on MMSEG algorithm developed by ANSI C. Completely based on modular implementation and can be easily embedded in other programs, like: MySQL, PostgreSQL, PHP, etc.
VSCode extension to highlight nested code blocks
A Cython MeCab wrapper for fast, pythonic Japanese tokenization and morphological analysis.
Tiny JavaScript tokenizer.
Ungreedy subword tokenizer and vocabulary trainer for Python, Go, C++ & Javascript
The fast scanner generator for Java™ with full Unicode support
Open Korean Text Processor - An Open-source Korean Text Processor
Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).
🌭 Mustard is a Swift library for tokenizing strings when splitting by whitespace doesn't cut it.
数据标注是一款专门对文本数据进行处理和标注的工具,通过简化快捷的文本标注流程和动态的算法反馈,支持用户快速标注关键词并能通过算法持续减少人工标注的成本和时间。数据标注的过程先由人工标注构建基础,再由自动标注反哺人工标注,最后由人工标注进行纠偏,从而大幅度提高标注的精准度和高效性。数据标注需要依赖开源的数字底座进行人员岗位管控。
24,523 repositories in the index in total.