thunlp/THUOCL
quality grade D, 40 out of 100THUOCL(THU Open Chinese Lexicon)中文词库
- stars
- 1.1k
- stars gained this week
- +6this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Tokenization, parsing, classical NLP pipelines, translation and information extraction.
Signals: nlp, natural-language-processing, tokenizer, named-entity-recognition, text-classification, machine-translation, sentiment-analysis, spacy
637 results
THUOCL(THU Open Chinese Lexicon)中文词库
Plug and Play Language Model implementation. Allows to steer topic and attributes of GPT-2 models.
Multilingual word vectors in 78 languages
Pre-trained subword embeddings in 275 languages, based on Byte-Pair Encoding (BPE)
📖 A curated list of awesome resources dedicated to Relation Extraction, one of the most important tasks in Natural Language Processing (NLP).
TextRank implementation for Python 3.
基于医药知识图谱的智能问答系统
Trained models & code to predict toxic comments on all 3 Jigsaw Toxic Comment Challenges. Built using ⚡ Pytorch Lightning and 🤗 Transformers. For access to our API, please email us at [email protected].
基于知识图谱的《红楼梦》人物关系可视化及问答系统
The most accurate natural language detection library for Go, suitable for short text and mixed-language text
"結巴"中文分詞:做最好的 PHP 中文分詞、中文斷詞組件。 / "Jieba" (Chinese for "to stutter") Chinese text segmentation: built to be the best PHP Chinese word segmentation module.
Evaluation code for various unsupervised automated metrics for Natural Language Generation.
:memo: Подборка ресурсов по машинному обучению
Python package for Korean natural language processing.
这个项目是一个基本包.封装了大多数nlp项目中常用工具
Code repo for "WebArena: A Realistic Web Environment for Building Autonomous Agents"
A collection of corpora for named entity recognition (NER) and entity recognition tasks. These annotated datasets cover a variety of languages, domains and entity types.
Must-read Papers on Textual Adversarial Attack and Defense
similarity: Text similarity calculation Toolkit for Java. 文本相似度计算工具包,java编写,可用于文本相似度计算、情感分析等任务,开箱即用。
This repository contains the code for "Exploiting Cloze Questions for Few-Shot Text Classification and Natural Language Inference"
:us: a python library for parsing unstructured United States address strings into address components
🧠 A study guide to learn about Transformers
Tika-Python is a Python binding to the Apache Tika™ REST services allowing Tika to be called natively in the Python community.
A fast, efficient universal vector embedding utility package.
24,535 repositories in the index in total.