kyegomez/BitNet
quality grade B, 77 out of 100Implementation of "BitNet: Scaling 1-bit Transformers for Large Language Models" in pytorch
- stars
- 1.9k
- stars gained this week
- -1this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
340 results
Implementation of "BitNet: Scaling 1-bit Transformers for Large Language Models" in pytorch
[ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.
GenAI Processors is a lightweight Python library that enables efficient, parallel content processing.
[ECCV2024] Video Foundation Models & Data for Multimodal Understanding
mPLUG-DocOwl: Modularized Multimodal Large Language Model for Document Understanding
Easily compute clip embeddings and build a clip retrieval system with them
A modular library built on top of Keras and TensorFlow to generate a caption in natural language for any input image.
Image Captioning using InceptionV3 and beam search
TensorFlow Implementation of "Show, Attend and Tell"
ZMJImageEditor is a picture editing component like WeChat. It is powerful and easy to integrate, supporting rendering, text, rotation, tailoring, mapping and other functions. (ZMJImageEditor 是一个和微信一样图片编辑的组件,功能强大,极易集成,支持绘制、文字、旋转、剪裁、贴图等功能)
Image Captioning Using Transformer
Transformer-based image captioning extension for pytorch/fairseq
多模态情感分析——基于BERT+ResNet的多种融合方法
[CVPR 2020] Meshed-Memory Transformer for Image Captioning
No description
Optimized local inference for LLMs with HuggingFace-like APIs for quantization, vision/language models, multimodal agents, speech, vector DB, and RAG.
Pocket-Sized Multimodal AI for content understanding and generation across multilingual texts, images, and 🔜 video, up to 5x faster than OpenAI CLIP and LLaVA 🖼️ & 🖋️
A simple, unified multimodal models training engine. Lean, flexible, and built for hacking at scale.
From scratch implementation of a vision language model in pure PyTorch
VLM-RL: A Unified Vision Language Models and Reinforcement Learning Framework for Safe Autonomous Driving
[NeurIPS 2025] Official implementation of "RoboRefer: Towards Spatial Referring with Reasoning in Vision-Language Models for Robotics"
[ICML'25][TPAMI'26] Official implementation of paper "SparseVLM" and "SparseVLM+".
Codes for VPGTrans: Transfer Visual Prompt Generator across LLMs. VL-LLaMA, VL-Vicuna.
Archived snapshot of Thinking-with-Visual-Primitives
24,538 repositories in the index in total.