IBM/terramind
quality grade B, 65 out of 100TerraMind is the first any-to-any generative foundation model for Earth Observation, built by IBM and ESA.
- stars
- 321
- stars gained this week
- +2this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
340 results
TerraMind is the first any-to-any generative foundation model for Earth Observation, built by IBM and ESA.
Official repository for "VideoPrism: A Foundational Visual Encoder for Video Understanding" (ICML 2024)
This repo contains the code for "VLM2Vec / MMEB" [ICLR 2025], "VLM2Vec-V2 / MMEB-V2" [TMLR 2026], and "MMEB-V3" [COLM 2026]
GLM-4.6V/4.5V/4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Local AI filmmaking studio — skills, canvas, timeline — driven from your coding agent.
Multimodal Large Language Models for Remote Sensing (RS-MLLMs): A Survey
😎 up-to-date & curated list of awesome Attacks on Large-Vision-Language-Models papers, methods & resources.
This repository contains the official implementation of "FastVLM: Efficient Vision Encoding for Vision Language Models" - CVPR 2025
A unified multimodal language model based on discrete sequence modeling
[IEEE TPAMI 2026] Simulating the Real World: Survey & Resources, which contains our survey "Simulating the Real World: A Unified Survey of Multimodal Generative Models" (IEEE TPAMI, 2026) and Awesome-Text2X-Resources. Watch this repository for the latest updates! 🔥
Get clean data from tricky documents, powered by vision-language models ⚡
Inductive graph-based matrix completion (IGMC) from "M. Zhang and Y. Chen, Inductive Matrix Completion Based on Graph Neural Networks, ICLR 2020 spotlight".
Code for our SIGKDD'22 paper Pre-training-Enhanced Spatial-Temporal Graph Neural Network For Multivariate Time Series Forecasting.
This repository contains the code for the paper `End-to-End Multimodal Emotion Recognition using Deep Neural Networks`.
Generates poetry from images using convolutional and recurrent neural networks
[ICLR 2026] MetaSpatial leverages reinforcement learning to enhance 3D spatial reasoning in vision-language models (VLMs), enabling more structured, realistic, and adaptive scene generation for applications in the metaverse, AR/VR, and game development.
🔍 OpenSearch-VL provides a fully open recipe for training strong multimodal deep search agents through high-quality data curation, diverse visual/search tools, and fatal-aware agentic reinforcement learning.
VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
Swift Core ML Examples
No description
[ICLR'25 Oral] UGround: Universal GUI Visual Grounding for GUI Agents
Towards Large Multimodal Models as Visual Foundation Agents
Official repo for "GeoVista: Web-Augmented Agentic Visual Reasoning for Geolocalization"
This is the official code of VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding (ECCV 2024)
24,537 repositories in the index in total.