barefootford/buttercut
quality grade C, 62 out of 100Edit Video with Claude Code
- stars
- 581
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
344 results
Edit Video with Claude Code
The code used to train and run inference with the ColVision models, e.g. ColPali, ColQwen2, and ColSmol.
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Mobile-Agent: The Powerful GUI Agent Family
A collection of resources on applications of multi-modal learning in medical imaging.
[Awesome-Spatial-VLMs] This repository is the official, community-maintained resource for the survey paper: Spatial Intelligence in Vision-Language Models: A Comprehensive Survey;
TapCanvas — 无限画布 ❌,沉浸式画布 ✅。线上版功能基本对齐TOP级别。agents-cli 智能体蒸馏自 Claude code + codex + hermes-agent。 可使用 cli 工具直接调试 & 调用画布全部功能。
GenAI Processors is a lightweight Python library that enables efficient, parallel content processing.
SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool.
🔍 OpenSearch-VL provides a fully open recipe for training strong multimodal deep search agents through high-quality data curation, diverse visual/search tools, and fatal-aware agentic reinforcement learning.
🔮 Libraries & tools for enabling Machine Learning driven user-experiences on the web
A novel Multimodal Large Language Model (MLLM) architecture, designed to structurally align visual and textual embeddings.
Implementation of "BitNet: Scaling 1-bit Transformers for Large Language Models" in pytorch
Inductive graph-based matrix completion (IGMC) from "M. Zhang and Y. Chen, Inductive Matrix Completion Based on Graph Neural Networks, ICLR 2020 spotlight".
This repository contains the code for the paper `End-to-End Multimodal Emotion Recognition using Deep Neural Networks`.
Generates poetry from images using convolutional and recurrent neural networks
[ICLR 2026] MetaSpatial leverages reinforcement learning to enhance 3D spatial reasoning in vision-language models (VLMs), enabling more structured, realistic, and adaptive scene generation for applications in the metaverse, AR/VR, and game development.
VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
[ICLR 2026] VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
Swift Core ML Examples
No description
[ICLR'25 Oral] UGround: Universal GUI Visual Grounding for GUI Agents
Towards Large Multimodal Models as Visual Foundation Agents
Official repo for "GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization"
24,523 repositories in the index in total.