Cuimao777/cuimao-translator
quality grade D, 38 out of 100一键将英文PDF翻译为流畅中文的Claude Code技能
- stars
- 480
- stars gained this week
- —this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Tokenization, parsing, classical NLP pipelines, translation and information extraction.
Signals: nlp, natural-language-processing, tokenizer, named-entity-recognition, text-classification, machine-translation, sentiment-analysis, spacy
637 results
一键将英文PDF翻译为流畅中文的Claude Code技能
使用Claude Code,Cursor, codex, gemini 写微信公众号文章工具。
WeChat chat history analyzer — runs in Claude Code, generates dual-user visualization + AI personality comparison report. No API key needed.
LLM-Powered Semi-Structured Table Question Answering
Data extraction with LLM on CPU
🤖 State-of-the-art, production ready LLM apps made mega-easy, so you don't have to build them from scratch 🤯 Create a bot, now 🫵
🐍 TOON for Python (Token-Oriented Object Notation) Encoder/Decoder - Reduce LLM token costs by 30-60% with structured data.
基于LLM和证据检索的虚假新闻自动检测系统
Classify and extract structured data with LLMs
No description
A customer segmentation project can be approached in multiple ways. In this repository, we will explore advanced techniques for defining clusters and analyzing the results.
PyNLPl, pronounced as 'pineapple', is a Python library for Natural Language Processing. It contains various modules useful for common, and less common, NLP tasks. PyNLPl can be used for basic tasks such as the extraction of n-grams and frequency lists, and to build simple language model. There are also more complex data types and algorithms. Moreover, there are parsers for file formats common in NLP (e.g. FoLiA/Giza/Moses/ARPA/Timbl/CQL). There are also clients to interface with various NLP specific servers. PyNLPl most notably features a very extensive library for working with FoLiA XML (Format for Linguistic Annotation).
This repository contains the official release of the model "BanglaBERT" and associated downstream finetuning code and datasets introduced in the paper titled "BanglaBERT: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in Bangla" accpeted in Findings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics: NAACL-2022.
PICARD - Parsing Incrementally for Constrained Auto-Regressive Decoding from Language Models. PICARD is a ServiceNow Research project that was started at Element AI.
The project page for "LOGIC-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning"
[EMNLP 2020] Obtain Word Alignments using Pretrained Language Models (e.g., mBERT)
[EMNLP 2023] Adapting Language Models to Compress Long Contexts
Large Language Models Are Reasoning Teachers (ACL 2023)
Code for the paper "Deep Entity Matching with Pre-trained Language Models"
Large Language Models: In this repository Language models are introduced covering both theoretical and practical aspects.
Student version of Assignment 2 for Stanford CS336 - Language Modeling From Scratch
Practical 6: LSTM language models
Tensorflow implementation of "Language Modeling with Gated Convolutional Networks"
[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
24,535 repositories in the index in total.