jflex-de/jflex
quality grade D, 42 out of 100The fast scanner generator for Java™ with full Unicode support
- stars
- 634
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Tokenization, parsing, classical NLP pipelines, translation and information extraction.
Signals: nlp, natural-language-processing, tokenizer, named-entity-recognition, text-classification, machine-translation, sentiment-analysis, spacy
637 results
The fast scanner generator for Java™ with full Unicode support
Open Korean Text Processor - An Open-source Korean Text Processor
Ekphrasis is a text processing tool, geared towards text from social networks, such as Twitter or Facebook. Ekphrasis performs tokenization, word normalization, word segmentation (for splitting hashtags) and spell correction, using word statistics from 2 big corpora (english Wikipedia, twitter - 330mil english tweets).
🌭 Mustard is a Swift library for tokenizing strings when splitting by whitespace doesn't cut it.
数据标注是一款专门对文本数据进行处理和标注的工具,通过简化快捷的文本标注流程和动态的算法反馈,支持用户快速标注关键词并能通过算法持续减少人工标注的成本和时间。数据标注的过程先由人工标注构建基础,再由自动标注反哺人工标注,最后由人工标注进行纠偏,从而大幅度提高标注的精准度和高效性。数据标注需要依赖开源的数字底座进行人员岗位管控。
支持中文和拼音的 SQLite fts5 全文搜索扩展 | A SQLite3 fts5 tokenizer which supports Chinese and PinYin
Optimised tokenizer/lexer generator! 🐄 Uses /y for performance. Moo.
한국어 자연어처리를 위한 파이썬 라이브러리입니다. 단어 추출/ 토크나이저 / 품사판별/ 전처리의 기능을 제공합니다.
Solves basic Russian NLP tasks, API for lower level Natasha projects
Persian NLP Toolkit
A small library for converting tokenized PHP source code into XML (and potentially other formats)
This program goes thru reddit, finds the most mentioned tickers and uses Vader SentimentIntensityAnalyzer to calculate the ticker compound value.
A language detection library for PHP. Detects the language from a given text string.
skweak: A software toolkit for weak supervision applied to NLP tasks
Summarization, translation, sentiment-analysis, text-generation and more at blazing speed using a T5 version implemented in ONNX.
Easy to use NLP library built on PyTorch and TorchText
텐서플로2와 머신러닝으로 시작하는 자연어처리 (로지스틱회귀부터 BERT와 GPT3까지) 실습자료
All NLP you Need Here. 目前包含15个NLP demo的pytorch实现(大量代码借鉴于其他开源项目,原先是自己玩的,后来干脆也开源出来)
Research and Materials on Hardware implementation of Transformer Model
The official implementation for ICLR23 spotlight paper "DIFFormer: Scalable (Graph) Transformers Induced by Energy Constrained Diffusion"
Repository for Project Insight: NLP as a Service
[ACL'20] HAT: Hardware-Aware Transformers for Efficient Natural Language Processing
Rust-tokenizer offers high-performance tokenizers for modern language models, including WordPiece, Byte-Pair Encoding (BPE) and Unigram (SentencePiece) models
Exploring attention weights in transformer-based models with linguistic knowledge.
24,535 repositories in the index in total.