Haojae/scipilot-figure-skill
quality grade C, 60 out of 100SciPilot Skills family - Publication-grade scientific figure copilot for Claude Code
- stars
- 1.6k
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
343 results
SciPilot Skills family - Publication-grade scientific figure copilot for Claude Code
An Asynchronous Reinforcement Learning Engine for Omni-Modal Post-Training at Scale
Build realtime multimodal AI agents with Node.js
面向 AI 创作的开源无限画布工作台,集成 AI 生图、参考图编辑、视频生成、Agent 智能助手、画布编排、对话创作、提示词库与素材管理等能力,支持可视化创作流程与多 Agent 协同工作。兼容 OpenAI 接口生态,支持 chatgpt2api、grok2api、flow2api、newapi 等渠道接入。
Visual intelligence for your home.
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac using MLX.
Multimodal RL training framework for diffusion & omni models
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.
Next-Gen AI Translation Tool Powered by LLM. Support Office documents, PDF, TXT, and more format with just one click.
Claude support for Apple Foundation Models
Open-source framework for developing real-time multimodal conversational AI agents.
azooKey-Desktop is an open-source Japanese input method for macOS, written in Swift and powered by the Zenzai neural kana-kanji converter. It provides live conversion, optional LLM-based “Magic Conversions”, and Tuner-backed personalization for a smooth, desktop typing experience.
Open Source project using LLMs to translate subtitles (SRT, SSA/ASS, VTT)
AI powered open source recommender system engine supports classical/LLM rankers and multimodal content via embedding
An extensible, state-of-the-art framework for columnar compression, and the fastest FOSS columnar file format. Formerly at @spiraldb, now an Incubation Stage project at LFAI&Data, part of the Linux Foundation.
🖨️ Automated scanner document processor with AI-powered naming and WebDav integration. Receives scans via FTP, extracts text using Vision AI, generates intelligent filenames with Ollama AI, and uploads to your cloud storage.
Ptera Software is a fast, easy-to-use, and open-source software package for analyzing flapping-wing flight.
Production ready toolkit to run AI locally
TongFlow : An Open-Source Multi-Modal GenAI Workflow Studio
Visualize, query, and stream to train on multimodal robotics data.
Mobile-Agent: The Powerful GUI Agent Family
AI-native video platform powered by Proof of Physical AI — agents and humans create, curate, and engage on hardware verified by physics. Part of the RustChain DePIN ecosystem.
ComfyUI nodes for vision-language models: Qwen3-VL, Moondream 3, Florence-2, SmolVLM2, InternVL, Gemma 3, MiniCPM-V. Plus open-vocabulary detection, SAM2/SAM3 segmentation, video temporal reasoning, GGUF via llama.cpp, and hosted LLM/VLM APIs.
Edit Video with Claude Code
24,520 repositories in the index in total.