Top AI Repos — open-source AI, indexed and scored
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
HY-Embodied: Embodied Foundation Models for Real-World Agents
| Date | Stars |
|---|---|
| 2026-07-31 | 840 |
| 2026-08-02 | 841 |
| 2026-08-06 | 841 |
Today
— stars today
This week
— stars this week
This month
— stars this month
Momentum
0.0
growth rate 0.00%/day
<div align="center"> <h1>Hy-Embodied-VLM-1.0</h1> <p><b>Efficient Physical-World Agents</b></p> <p><i>Tencent Robotics X × Hy Vision Team × Futian Laboratory</i></p> <a href="https://arxiv.org/abs/2607.12894"><img src="https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv" alt="arXiv"></a> <a href="hy_embodied_vlm_1_0_tech_report.pdf"><img src="https://img.shields.io/badge/PDF-Report-green?logo=report" alt="Tech Report"></a> <a href="https://huggingface.co/tencent/Hy-Embodied-VLM-1.0"><img src="https://img.shields.io/badge/Models-HuggingFace-yellow?logo=huggingface" alt="Models"></a> <a href="./LICENSE"><img src="https://img.shields.io/badge/License-Apache_2.0-blue" alt="License"></a> </div> ## 🔥 Updates * **`[2026-07-15]`** 🚀 We have released **Hy-Embodied-VLM-1.0**! An efficient Mixture-of-Experts vision–language foundation model for embodied agents in the physical world, activating only **~3B parameters** per token (~30B total) for high inference efficiency. Weights are available on [Hugging Face](https://huggingface.co/tencent/Hy-Embodied-VLM-1.0), together with inference code for both HuggingFace `transformers` and vLLM. * **`[2026-06-15]`** 🤖 We have released **HY-VLA-0.5**! The [official code](https://github.com/Tencent-Hunyuan/Hy-Embodied-0.5-VLA), UMI-trained [weights](https://huggingface.co/tencent/Hy-Embodied-0.5-VLA-UMI) and 2000+ hours of high-fidelity UMI [data](https://huggingface.co/datasets/tencent/Hy-Embodied-0.5-VLA-Data) are now available. * **`[2026-04-09]`** 🚀 We have released **HY-Embodied-0.5**, featuring the open-sourced `HY-Embodied-0.5 MoT-2B` weights on [Hugging Face](https://huggingface.co/tencent/HY-Embodied-0.5/tree/main) along with the official inference code! Documentation is now available under [`Hy-Embodied-0.5/`](./Hy-Embodied-0.5/). ## 📖 Abstract Building capable embodied agents requires not only multimodal perception and understanding, but also agentic capabilities for reasoning about actions, adapting to evolving situations, and interacting with the physical world. In this report, we introduce **Hy-Embodied-VLM-1.0**, an efficient and powerful embodied foundation model specifically designed for embodied agents operating in the physical world. To cultivate such capabilities from the pre-training stage onward, we define an **action-centric capability taxonomy** comprising three progressive dimensions: **Action-Relevant State Understanding**, **Action–Transition Reasoning**, and **Sequential and Adaptive Reasoning**. Guided by this taxonomy, we develop a systematic data pipeline and curate data mixtures spanning both pre-training and post-training. To deliver strong physical-world understanding and interaction capabilities while supporting latency-sensitive deployment, we build our model on the **Hy3-A3B language backbone** and the **Hy-ViT2 vision encoder**. Its efficient Mixture-of-Experts architecture combines strong model capacity with high inference efficiency. We evaluate Hy-Embodied-VLM-1.0 on a comprehensive suite of **38 benchmarks** covering embodied perception, physical-world understanding, and embodied reasoning. The model achieves the best performance among similarly sized models on **19 of the 38 benchmarks** and substantially outperforms strong competitors, including **Qwen3.6-A3B** and **Cosmos 3**. Compared with the previous-generation Hy-Embodied-0.5 MoT-2B, Hy-Embodied-VLM-1.0 improves average performance by **8.4%**. Despite activating only **3B parameters**, it achieves performance close to that of the previous-generation model with 32B activated parameters. Beyond static benchmark evaluation, Hy-Embodied-VLM-1.0 also demonstrates strong performance on embodied agentic tasks requiring multi-turn interaction and long-horizon reasoning. <div align="center"> <img src="Hy-Embodied-VLM-1.0/figures/1-teaser.png" alt="Hy-Embodied-VLM-1.0 Performance" width="85%"> </div> ## ⭐️ Key Features * 🧠 **Efficient MoE, ~3B activated** — Combines the Hy3-A3B lan
Excerpt of 17,704 characters
Read on GitHubZuyan Liu · Tsinghua University · China
9
Xumin Yu · Tsinghua University · China
3
3
Yongming Rao · Tencent
2
2
2
Would you bet a product on this? Bounded 0–100 and slow moving.
matched fp:66db8c1cbf9f5e08, llm:Repository title and description: 'HY-Embodied: Embodied Foundation Models for Real-World Agents' — suggests foundation models for embodied agents/real-world robotics/agents.
matched fp:66db8c1cbf9f5e08, llm:Repository title and description: 'HY-Embodied: Embodied Foundation Models for Real-World Agents' — suggests foundation models for embodied agents/real-world robotics/agents.
matched fp:66db8c1cbf9f5e08, llm:Repository title and description: 'HY-Embodied: Embodied Foundation Models for Real-World Agents' — suggests foundation models for embodied agents/real-world robotics/agents.
matched fp:66db8c1cbf9f5e08, llm:Repository title and description: 'HY-Embodied: Embodied Foundation Models for Real-World Agents' — suggests foundation models for embodied agents/real-world robotics/agents.