lynote-ai/humanize-text
quality grade A, 84 out of 100Open-source pipeline and reference implementations for improving the readability and natural cadence of AI-assisted drafts.
- stars
- 3.0k
- stars gained this week
- +1.4kthis week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
Open-source pipeline and reference implementations for improving the readability and natural cadence of AI-assisted drafts.
LLM inference in C/C++
Self-hosted LLM router with a managed safety net. OpenAI-compatible. BYOK. Single-workspace. Streaming. For more advanced routing choose hosted OrcaRouter
Terraform Skill for Claude Code and Codex. LLMs hallucinate a lot with Terraform - TerraShark fixes this. It eliminates hallucinations, is designed for modular and secure code and grounds your IaC in the official Hashicorp Terraform best practices.
The fastest way to run Qwen 3.8 Flash Next and Qwen 3.8 27B on a Mac: 125 tok/s in OpenCode on an M5 Max. Native MTP speculative decoding on Apple Silicon, exact at any temperature. OpenAI and Anthropic compatible local server.
An all-in-one, pure C++ inference engine for audio models, powered by ggml. Supports TTS, STT, VAD, voice conversion, music generation, and more, with highly optimized performance. No Python dependency.
Production-grade C++ edge AI engine for video analytics and on-device VLM across Sophon, Rockchip RKNN, and x86, with visual orchestration, real-time OSD, events, and reproducible benchmarks.
Efficient Long-Context LLM Serving with Head-Aware KV Reuse and SegPagedAttention
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python. Includes MLX Core macOS app with chat, agent mode, and tool calling.
DeepSeek 4 Flash and PRO local inference engine for Metal, CUDA and ROCm
A framework for efficient model inference with omni-modality models
原汁原昧 Claude Code 可运行,可构建, 可调试版; 生产级工程化, 企业级可靠性; 安全无毒, 内存泄露修复
A high-performance inference engine for LLM, VLM, DiT and REC models, optimized for diverse AI accelerators. It is hosted in OpenAtom Foundation.
Your AI second brain. Self-hostable. Get answers from the web or your docs. Build custom agents, schedule automations, do deep research. Turn any online or local LLM into your personal, autonomous AI (gpt, claude, gemini, llama, qwen, mistral). Get started - free.
Silero VAD: pre-trained enterprise-grade Voice Activity Detector
Virtualized Elastic KV Cache for Dynamic GPU Sharing and Beyond
A Datacenter Scale Distributed Inference Serving Framework
LMCache: Supercharge Your LLM with the Fastest KV Cache Layer
Tiny, no-nonsense, self-contained, Tensorflow and ONNX inference
LLMRouter: An Open-Source Library for LLM Routing
Run LLMs with MLX
Distribute and run LLMs with a single file.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
A programmable Mixture-of-Models router for heterogeneous LLM inference
24,535 repositories in the index in total.