IBM/terramind
quality grade B, 65 out of 100TerraMind is the first any-to-any generative foundation model for Earth Observation, built by IBM and ESA.
- stars
- 298
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
344 results
TerraMind is the first any-to-any generative foundation model for Earth Observation, built by IBM and ESA.
[CVPR 2026] Scaling Spatial Intelligence with Multimodal Foundation Models
Foundation Models and Data for Human-Human and Human-AI interactions.
Developing Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
Multimodal Whole Slide Foundation Model for Pathology - Nature Medicine
Falcon: A Remote Sensing Vision-Language Foundation Model
[TMLR 2024] repository for VLN with foundation models
Vircadia open source agent-based metaverse ecosystem.
Pathology Language and Image Pre-Training (PLIP) is the first vision and language foundation model for Pathology AI (Nature Medicine). PLIP is a large-scale pre-trained model that can be used to extract visual and language features from pathology images and text description. The model is a fine-tuned version of the original CLIP model.
[ICML'24] SeeAct is a system for generalist web agents that autonomously carry out tasks on any given website, with a focus on large multimodal models (LMMs) such as GPT-4V(ision).
Code for "WebVoyager: WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models"
An open source implementation of "Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning", an all-new multi modal AI that uses just a decoder to generate both text and images
This is the official repository for Retrieval Augmented Visual Question Answering
🤖 AI-native Visual Analytics framework build for agents.
ScholarMind - 面向大模型Agent领域的多模态学术 Agent | Multimodal Academic Research Agent with Knowledge Graph & Learning Path Planning
A curated list of AWESOME papers, datasets and tutorials within Multimodal Knowledge Graph.
Clone any .pptx into your own deck — OpenAI gpt-image-2 mimics the layout, you supply the content. 10 bundled styles. | 把任何 .pptx 模板"抄"成你的 PPT:gpt-image-2 仿版式、你换内容,另含 10 套精选风格。Claude Code / OpenClaw skill.
A Claude skill for developing WebGPU applications with Three.js
分镜助手 - 基于节点画布的 AI 分镜工作台,一站式完成图片生成、编辑与分镜流程
NeurIPS 2024 Paper: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
An Open-Source Alternative to MATLAB: LLM-Driven & Scilab-Based AI Mathematical Modeling Tool with Natural Language Interface
Ming - facilitating advanced multimodal understanding and generation capabilities built upon the Ling LLM.
An Open-Source Multimodal AIGC Solution based on ComfyUI + MCP + LLM https://pixelle.ai
Local Video-LLM powered AI Baby Monitor
24,523 repositories in the index in total.