ddcat-ai/open-ai-canvas
quality grade B, 66 out of 100面向 AI 影视创作的开源无限画布工作台,集成多模态生成、分镜编排、素材管理与 Agent 工作流。
- stars
- 374
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
344 results
面向 AI 影视创作的开源无限画布工作台,集成多模态生成、分镜编排、素材管理与 Agent 工作流。
This is the official code of VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding (ECCV 2024)
Build realtime voice and video agents with Google's new Gemini 2.0 (API is free for now)
Official Code Release of SAGE: Scalable Agentic 3D Scene Generation for Embodied AI
An AI agent that automates the creation of CVPR/NeurIPS standard academic diagrams. Implements a strict "Logic (Architect) -> Vision (Renderer)" workflow to transform paper abstracts into high-fidelity scientific illustrations.
Local AI filmmaking studio — skills, canvas, timeline — driven from your coding agent.
OpenCode/Claude Code 技能库 。特色技能:视频生成、图片生成、AI Agent 互联、智能问数等等各场景,持续开发更新中
AI-powered Instagram carousel builder. Chat with Claude to design slides; export as PNGs at exact Instagram dimensions. Type /start in Claude Code to bootstrap.
为访谈视频添加综艺风格视觉特效的 Claude Skill。AI 分析字幕内容生成建议,用户审批后自动渲染。A Claude Skill for adding variety-show-style visual effects to interview videos. The AI analyzes subtitle content to generate suggestions
SenseCAP Watcher: Intelligent ESP32S3-based device with Himax WiseEye2 AI, capable of seeing, hearing, and interacting using advanced AI and the LLM-enabled SenseCraft suite. Perfect for environment-aware applications in automation and security.
[ACL 2025 🔥] Rethinking Step-by-step Visual Reasoning in LLMs
General video interaction platform based on LLMs, including Video ChatGPT
Interesting physics-sims generated via LLM prompting.
chat to visualization with LLM
[CVPR'2024 Highlight] Official PyTorch implementation of the paper "VTimeLLM: Empower LLM to Grasp Video Moments".
🎨 点击式探索的知识画册,长按图片生成带标注子图 | Flipbook Canvas — click-to-explore knowledge picture-book. Long-press any image to spawn an annotated child diagram via a pluggable multimodal pipeline (LLM + image gen + web search + OCR).
✨First Open-Source R1-like Video-LLM [2025/02/18]
引入LLM增强输入体验的小狼毫输入法 | LLM-powered IME for Chinese text completion
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale (CVPR 2025)
[ICLR 2024 Spotlight] DreamLLM: Synergistic Multimodal Comprehension and Creation
No description
[CVPR 2025] This is a model aggregated with CLIP and SAM version of SkySense for remote sensing interpretation described in SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling.
Cognitive runtime for language models with memory, metacognition, multimodal channels, native plugins, and a self-evolving Executive.
Official Implementation of ECCV2024 paper: Chat Edit 3D: Interactive 3D Scene Editing via Large Language Model
24,523 repositories in the index in total.