TingsongYu/PyTorch-Tutorial-2nd
quality grade C, 56 out of 100《Pytorch实用教程》(第二版)无论是零基础入门,还是CV、NLP、LLM项目应用,或是进阶工程化部署落地,在这里都有。相信在本书的帮助下,读者将能够轻松掌握 PyTorch 的使用,成为一名优秀的深度学习工程师。
- stars
- 4.6k
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Runtimes and servers that execute model inference at speed and scale.
Signals: llm-inference, inference, inference-engine, model-serving, llm-serving, llama-cpp, vllm, triton-inference-server
410 results
《Pytorch实用教程》(第二版)无论是零基础入门,还是CV、NLP、LLM项目应用,或是进阶工程化部署落地,在这里都有。相信在本书的帮助下,读者将能够轻松掌握 PyTorch 的使用,成为一名优秀的深度学习工程师。
LiteRT and LiteRT-LM sample apps, model recipes, agent skills and utilities.
A lite C++ AI toolkit: 100+ models with MNN, ORT and TRT, including Det, Seg, Stable-Diffusion, Face-Fusion.
Accurate, large-scale, and extensible simulator for LLM inference Systems
Minimal yet performant LLM examples in pure JAX
Open-source LLM/VLM load balancer and serving platform for self-hosting LLMs (and VLMs) at scale 🏓🦙 Alternative to projects like llm-d, Docker Model Runner, etc but with less moving parts and simple deployments built around ggml ecosystem. Runs on CPU and GPU.
Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.
Espressif deep-learning library for AIoT applications
LLM.swift is a simple and readable library that allows you to interact with large language models locally with ease for macOS, iOS, watchOS, tvOS, and visionOS.
High-performance Inference and Deployment Toolkit for LLMs and VLMs based on PaddlePaddle
Everything you need to know about LLM inference
Examples for using ONNX Runtime for machine learning inferencing.
No description
An open-source project for Windows developers to learn how to add AI with local models and APIs to Windows apps.
📚 Jupyter notebook tutorials for OpenVINO™
Pre-trained Deep Learning models and demos (high quality and extremely fast)
Ollama alternative for Rockchip NPU: An efficient solution for running AI and Deep learning models on Rockchip devices with optimized NPU support ( rkllm )
✍🏻 Source Code Deep Dives, System Design & Engineering Blogs | Halfrost-Field 冰霜之地:源码解析、系统设计与工程实践笔记
Zant simplifies the deployment and optimization of neural networks on microprocessors
Serverless LLM Serving for Everyone.
CacheRoute is an innovative LLM scheduling scheme dedicated to enabling flexible KV cache reuse across LLM systems, improving task performance and system efficiency.
☸️ Easy, advanced inference platform for large language models on Kubernetes. 🌟 Star to support our work!
Dynamic Memory Management for Serving LLMs without PagedAttention
[NeurIPS'24 Spotlight, ICLR'25, ICML'25] To speed up Long-context LLMs' inference, approximate and dynamic sparse calculate the attention, which reduces inference latency by up to 10x for pre-filling on an A100 while maintaining accuracy.
24,535 repositories in the index in total.