xaviviro/python-toon
quality grade C, 53 out of 100🐍 TOON for Python (Token-Oriented Object Notation) Encoder/Decoder - Reduce LLM token costs by 30-60% with structured data.
- stars
- 344
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Tokenization, parsing, classical NLP pipelines, translation and information extraction.
Signals: nlp, natural-language-processing, tokenizer, named-entity-recognition, text-classification, machine-translation, sentiment-analysis, spacy
637 results
🐍 TOON for Python (Token-Oriented Object Notation) Encoder/Decoder - Reduce LLM token costs by 30-60% with structured data.
基于LLM和证据检索的虚假新闻自动检测系统
Classify and extract structured data with LLMs
No description
A customer segmentation project can be approached in multiple ways. In this repository, we will explore advanced techniques for defining clusters and analyzing the results.
PyNLPl, pronounced as 'pineapple', is a Python library for Natural Language Processing. It contains various modules useful for common, and less common, NLP tasks. PyNLPl can be used for basic tasks such as the extraction of n-grams and frequency lists, and to build simple language model. There are also more complex data types and algorithms. Moreover, there are parsers for file formats common in NLP (e.g. FoLiA/Giza/Moses/ARPA/Timbl/CQL). There are also clients to interface with various NLP specific servers. PyNLPl most notably features a very extensive library for working with FoLiA XML (Format for Linguistic Annotation).
This repository contains the official release of the model "BanglaBERT" and associated downstream finetuning code and datasets introduced in the paper titled "BanglaBERT: Language Model Pretraining and Benchmarks for Low-Resource Language Understanding Evaluation in Bangla" accpeted in Findings of the Annual Conference of the North American Chapter of the Association for Computational Linguistics: NAACL-2022.
PICARD - Parsing Incrementally for Constrained Auto-Regressive Decoding from Language Models. PICARD is a ServiceNow Research project that was started at Element AI.
The project page for "LOGIC-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning"
[EMNLP 2020] Obtain Word Alignments using Pretrained Language Models (e.g., mBERT)
[EMNLP 2023] Adapting Language Models to Compress Long Contexts
Large Language Models Are Reasoning Teachers (ACL 2023)
Code for the paper "Deep Entity Matching with Pre-trained Language Models"
Large Language Models: In this repository Language models are introduced covering both theoretical and practical aspects.
Student version of Assignment 2 for Stanford CS336 - Language Modeling From Scratch
Practical 6: LSTM language models
Tensorflow implementation of "Language Modeling with Gated Convolutional Networks"
[NeurIPS 2022] 🛒WebShop: Towards Scalable Real-World Web Interaction with Grounded Language Agents
An open platform for artificial intelligence, chat bots, virtual agents, social media automation, and live chat automation.
A platform for building conversational interfaces with intelligent agents (chatbots)
Language model fine-tuning on NER with an easy interface and cross-domain evaluation. "T-NER: An All-Round Python Library for Transformer-based Named Entity Recognition, EACL 2021"
The code to reproduce results from paper "MultiFiT: Efficient Multi-lingual Language Model Fine-tuning" https://arxiv.org/abs/1909.04761
Query your data in plain English with a fine-tuned Text2SQL model
autonomous trading agent powered by social sentiment analysis and ai that learns, grows, and adapts
24,523 repositories in the index in total.