xzf-thu/Audio-Reasoner
quality grade D, 37 out of 100The first Large Audio Language Model that enables native in-depth thinking, which is trained on large-scale audio Chain-of-Thought data.
- stars
- 297
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Vision-language models, document understanding and any-to-any architectures.
Signals: multimodal, vision-language-model, vlm, clip, multimodal-llm, visual-question-answering, document-understanding, image-captioning
344 results
The first Large Audio Language Model that enables native in-depth thinking, which is trained on large-scale audio Chain-of-Thought data.
Official Implementation of "MMaDA-Parallel: Multimodal Large Diffusion Language Models for Thinking-Aware Editing and Generation"
Official repository of "GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing"
The Codes and Data of A Comprehensive Benchmark for Multimodal Large Language Models in Industrial Anomaly Detection [ICLR'25]
Code for the paper "LLark: A Multimodal Instruction-Following Language Model for Music" by Josh Gardner, Simon Durand, Daniel Stoller, and Rachel Bittner.
This is the official repository for the LENS (Large Language Models Enhanced to See) system.
Codes for Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model
Official repo for NeurIPS 2023 paper "LayoutGPT: Compositional Visual Planning and Generation with Large Language Models"
SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
[CVPR 2024] VCoder: Versatile Vision Encoders for Multimodal Large Language Models
a multiscale multimodal large language models for radiology report generation (RRG) tasks
E5-V: Universal Embeddings with Multimodal Large Language Models
[CVPR 2024] TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding
Implementation of the image-sentence embedding method described in "Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models"
Language Models Can See: Plugging Visual Controls in Text Generation
Align 3D Point Cloud with Multi-modalities for Large Language Models
MU-LLaMA: Music Understanding Large Language Model
Symphony Generation with Permutation Invariant Language Model
Implementation of "PaLM-E: An Embodied Multimodal Language Model"
Bridging Vision and Language Model
Multimodal Large Language Models for Remote Sensing (RS-MLLMs): A Survey
A General, Accurate, Long-Horizon, and Efficient Mobile Agent driven by Multimodal Foundation Models
Official Repository of ChartX & ChartVLM: A Versatile Benchmark and Foundation Model for Complicated Chart Reasoning
24,523 repositories in the index in total.