tong-io/tongflow
quality grade B, 72 out of 100Modality-First GenAI Platform
- stars
- 1.0k
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
340 results
Modality-First GenAI Platform
🖨️ Automated scanner document processor with AI-powered naming and WebDav integration. Receives scans via FTP, extracts text using Vision AI, generates intelligent filenames with Ollama AI, and uploads to your cloud storage.
👀 Train a 65M-parameter VLM from scratch in just 2h!
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
NVIDIA AI Blueprint for video search and summarization (VSS) is a GPU-accelerated reference architecture for building video analytics agents with real-time verified alerts, visual Q&A, and automated reporting. The VSS Blueprint uses vision language models (VLMs) such as NVIDIA Cosmos, LLMs such as NVIDIA Nemotron, RAG, and NVIDIA NIMs.
Clone any .pptx into your own deck — OpenAI gpt-image-2 mimics the layout, you supply the content. 10 bundled styles. | 把任何 .pptx 模板"抄"成你的 PPT:gpt-image-2 仿版式、你换内容,另含 10 套精选风格。Claude Code / OpenClaw skill.
Mobile-Agent: The Powerful GUI Agent Family
Fully Open Framework for Democratized Multimodal Training
AI powered open source recommender system engine supports classical/LLM rankers and multimodal content via embedding
AI-native video platform where autonomous agents create, publish, and earn. Multi-agent content economy on Sybil-resistant, hardware-verified infrastructure. Part of the RustChain DePIN ecosystem.
Open Source project using LLMs to translate subtitles (SRT, SSA/ASS, VTT)
Ptera Software is a fast, easy-to-use, and open-source software package for analyzing flapping-wing flight.
An open-source implementaion for fine-tuning Qwen-VL series by Alibaba Cloud.
azooKey-Desktop is an open-source Japanese input method for macOS, written in Swift and powered by the Zenzai neural kana-kanji converter. It provides live conversion, optional LLM-based “Magic Conversions”, and Tuner-backed personalization for a smooth, desktop typing experience.
Aircraft design optimization made fast through computational graph transformations (e.g., automatic differentiation). Composable analysis tools for aerodynamics, propulsion, structures, trajectory design, and much more.
🧠 Web Neural Network API
SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool.
Translate EPUB books using Large Language Models while preserving the original text. The translated content is displayed side-by-side with the original, creating bilingual books perfect for language learning and cross-reference reading.
Real-time Vision Language Model interaction via webcam - WebRTC-based web interface
A Scientific Multimodal Foundation Model
A most Frontend Collection and survey of vision-language model papers, and models GitHub repository. Continuous updates.
🥂 Gracefully face hCaptcha challenge with multimodal large language model.
Fine-tune Gemma 4 and 3n with audio, images and text on Apple Silicon, using PyTorch and Metal Performance Shaders.
Make text LLMs listen and speak
24,537 repositories in the index in total.