facebookresearch/seamless_interaction
quality grade D, 42 out of 100Foundation Models and Data for Human-Human and Human-AI interactions.
- stars
- 413
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
340 results
Foundation Models and Data for Human-Human and Human-AI interactions.
Developing Generalist Foundation Models from a Multimodal Dataset for 3D Computed Tomography
Multimodal Whole Slide Foundation Model for Pathology - Nature Medicine
Falcon: A Remote Sensing Vision-Language Foundation Model
[TMLR 2024] repository for VLN with foundation models
Vircadia open source agent-based metaverse ecosystem.
Pathology Language and Image Pre-Training (PLIP) is the first vision and language foundation model for Pathology AI (Nature Medicine). PLIP is a large-scale pre-trained model that can be used to extract visual and language features from pathology images and text description. The model is a fine-tuned version of the original CLIP model.
[ICML'24] SeeAct is a system for generalist web agents that autonomously carry out tasks on any given website, with a focus on large multimodal models (LMMs) such as GPT-4V(ision).
Code for "WebVoyager: WebVoyager: Building an End-to-End Web Agent with Large Multimodal Models"
An open source implementation of "Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning", an all-new multi modal AI that uses just a decoder to generate both text and images
This is the official repository for Retrieval Augmented Visual Question Answering
🤖 AI-native Visual Analytics framework build for agents.
ScholarMind - 面向大模型Agent领域的多模态学术 Agent | Multimodal Academic Research Agent with Knowledge Graph & Learning Path Planning
A curated list of AWESOME papers, datasets and tutorials within Multimodal Knowledge Graph.
A Claude skill for developing WebGPU applications with Three.js
分镜助手 - 基于节点画布的 AI 分镜工作台,一站式完成图片生成、编辑与分镜流程
NeurIPS 2024 Paper: A Unified Pixel-level Vision LLM for Understanding, Generating, Segmenting, Editing
Ming - facilitating advanced multimodal understanding and generation capabilities built upon the Ling LLM.
An Open-Source Multimodal AIGC Solution based on ComfyUI + MCP + LLM https://pixelle.ai
Local Video-LLM powered AI Baby Monitor
Recent LLM-based CV and related works. Welcome to comment/contribute!
[CVPR 2024] OneLLM: One Framework to Align All Modalities with Language
LLM2CLIP significantly improves already state-of-the-art CLIP models.
VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
24,538 repositories in the index in total.