Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
A curated list of academic papers and resources on Physical AI — focusing on Vision-Language-Action (VLA) models, world models, embodied ai, and robotic foundation models.
| Date | Stars |
|---|---|
| 2026-07-31 | 371 |
| 2026-08-04 | 374 |
| 2026-08-06 | 374 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
# Awesome Physical AI [](https://awesome.re) A curated list of academic papers and resources on **Physical AI** — focusing on Vision-Language-Action (VLA) models, world models, embodied ai, and robotic foundation models. > **Physical AI** refers to AI systems that interact with and manipulate the physical world through robotic embodiments, combining perception, reasoning, and action in real-world environments. --- ## Table of Contents - [Foundations](#foundations) - [Vision-Language Backbones](#vision-language-backbones) - [Visual Representations](#visual-representations) - [VLA Architectures](#vla-architectures) - [End-to-End VLAs](#end-to-end-vlas) - [Modular VLAs](#modular-vlas) - [Compact & Efficient VLAs](#compact--efficient-vlas) - [Action Representation](#action-representation) - [Discrete Tokenization](#discrete-tokenization) - [Discrete Diffusion VLAs](#discrete-diffusion-vlas) - [Continuous & Diffusion Policies](#continuous--diffusion-policies) - [World Models](#world-models) - [JEPA & Latent Prediction](#jepa--latent-prediction) - [Generative World Models](#generative-world-models) - [Embodied World Models](#embodied-world-models) - [Reasoning & Planning](#reasoning--planning) - [Chain-of-Thought & Deliberation](#chain-of-thought--deliberation) - [Error Detection & Recovery](#error-detection--recovery) - [Learning Paradigms](#learning-paradigms) - [Imitation Learning](#imitation-learning) - [Reinforcement Learning](#reinforcement-learning) - [Reward Design](#reward-design) - [Scaling & Generalization](#scaling--generalization) - [Scaling Laws](#scaling-laws) - [Cross-Embodiment Transfer](#cross-embodiment-transfer) - [Open-Vocabulary Generalization](#open-vocabulary-generalization) - [Deployment](#deployment) - [Quantization & Compression](#quantization--compression) - [Real-Time Control](#real-time-control) - [Safety & Alignment](#safety--alignment) - [Lifelong Learning](#lifelong-learning) - [Applications](#applications) - [Humanoid Robots](#humanoid-robots) - [Manipulation](#manipulation) - [Navigation](#navigation) - [Sim-to-Real Transfer](#sim-to-real-transfer) - [Surveys](#surveys) - [Resources](#resources) - [Datasets & Benchmarks](#datasets--benchmarks) - [Simulation Platforms](#simulation-platforms) - [Companies & Projects](#companies--projects) - [Related Works](#related-works) --- ## Foundations ### Vision-Language Backbones > Core vision-language models that serve as pretrained backbones for Physical AI systems. - **CLIP**: "Learning Transferable Visual Models From Natural Language Supervision", *ICML 2021*. [[Paper](https://arxiv.org/abs/2103.00020)] [[Code](https://github.com/openai/CLIP)] - Foundational model aligning vision and language that underlies most VLA perception systems. - **SigLIP**: "Sigmoid Loss for Language Image Pre-Training", *ICCV 2023*. [[Paper](https://arxiv.org/abs/2303.15343)] - **PaLI-X**: "PaLI-X: On Scaling up a Multilingual Vision and Language Model", *CVPR 2024*. [[Paper](https://arxiv.org/abs/2305.18565)] - **LLaVA**: "Visual Instruction Tuning", *NeurIPS 2023*. [[Paper](https://arxiv.org/abs/2304.08485)] [[Project](https://llava-vl.github.io/)] - **Prismatic VLMs**: "Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models", *ICML 2024*. [[Paper](https://arxiv.org/abs/2402.07865)] [[Code](https://github.com/TRI-ML/prismatic-vlms)] - Systematic study of VLM design choices informing OpenVLA and other robotics VLMs. ### Visual Representations > Self-supervised visual encoders and perception models used in robotics. - **DINOv2**: "DINOv2: Learning Robust Visual Features without Supervision", *arXiv, Apr 2023*. [[Paper](https://arxiv.org/abs/2304.07193)] [[Code](https://github.com/facebookresearch/dinov2)] - **SAM**: "Segment Anything", *ICCV 2023*. [[Paper](https://arxiv.org/abs/2304.02643)] [[Project](https://segment-anything.com/)] -
Excerpt of 103,950 characters
Read on GitHubWould you bet a product on this? Bounded 0–100 and slow moving.
matched fp:62e3e2316a663988, topic:robotics, topic:embodied-ai, desc:embodied ai
matched fp:62e3e2316a663988, topic:awesome, desc:curated list