basketikun/infinite-canvas
quality grade A, 83 out of 100面向 AI 创作的开源无限画布工作台,集成 AI 生图、参考图编辑、视频生成、Agent 智能助手、画布编排、对话创作、提示词库与素材管理等能力,支持可视化创作流程与多 Agent 协同工作。兼容 OpenAI 接口生态,支持 chatgpt2api、grok2api、flow2api、newapi 等渠道接入。
- stars
- 6.8k
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
340 results
面向 AI 创作的开源无限画布工作台,集成 AI 生图、参考图编辑、视频生成、Agent 智能助手、画布编排、对话创作、提示词库与素材管理等能力,支持可视化创作流程与多 Agent 协同工作。兼容 OpenAI 接口生态,支持 chatgpt2api、grok2api、flow2api、newapi 等渠道接入。
Give Claude the ability to watch any video. /watch downloads, extracts frames, transcribes, hands it all to Claude.
【影策】面向 AI 影视创作的开源无限画布工作台,集成多模态生成、分镜编排、素材管理与 Agent 工作流。
No description
Visualize, query, and stream to train on multimodal robotics data.
The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra
Let Claude (or any LLM) actually watch a video — scene-aware, deduplicated frames + transcript, from a URL or local file. Runs locally, MIT.
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
(ECCV 2026 oral) LingBot-Map: Geometric Context Transformer for Streaming 3D Reconstruction
An open-weight 11B model series for long-form and real-time video understanding
Nurture your own AI assistant on your own computer. Local-first and MIT: memories, data and apps stay as plain files in your workspace. Chat summons the right GUI — wiki, spreadsheet, chart, form, 3D. Build small apps for an audience of one, no programming required.
Edit Video with Claude Code
Next-Gen AI Translation Tool Powered by LLM. Support Office documents, PDF, TXT, and more format with just one click.
An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation.
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
独一无二的沉浸式画布,完全开源
SciPilot Skills family - Publication-grade scientific figure copilot for Claude Code
Production ready toolkit to run AI locally
Visual intelligence for your home.
Claude support for Apple Foundation Models
Houdini Agent - DCC Asset Manager with AI capabilities
The code used to train and run inference with the ColVision models, e.g. ColPali, ColQwen2, and ColSmol.
Build realtime multimodal AI agents with Node.js
Cross-platform on-device AI toolkit
24,537 repositories in the index in total.