OpenGVLab/CaFo
quality grade D, 43 out of 100[CVPR 2023] Prompt, Generate, then Cache: Cascade of Foundation Models makes Strong Few-shot Learners
- stars
- 379
- stars gained this week
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Detection, segmentation, tracking, OCR models, 3D reconstruction and classical vision.
Signals: computer-vision, object-detection, image-segmentation, yolo, opencv, image-classification, pose-estimation, object-tracking
1,479 results
[CVPR 2023] Prompt, Generate, then Cache: Cascade of Foundation Models makes Strong Few-shot Learners
Official implementation of "Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation".
[CVPR 2024] Probing the 3D Awareness of Visual Foundation Models
[CVPR 2025] DEFOM-Stereo: Depth foundation model based stereo matching
[ECCV 2024] Official repository of Agent Attention
VIGA: Vision-as-Inverse-Graphics Agent
Our method reconstructs 3D worlds from video diffusion models using non-rigid alignment to resolve inherent 3D inconsistencies in the generated sequences.
[CVPR'25 Oral] Official implementation for "DiffusionRenderer: Neural Inverse and Forward Rendering with Video Diffusion Models"
Official pytorch implementation for "LightenDiffusion: Unsupervised Low-Light Image Enhancement with Latent-Retinex Diffusion Models"
This is the official Pytorch implementation of the paper "Diffusion Models for Implicit Image Segmentation Ensembles".
[CVPR'24] Scaling Diffusion Models to Real-World 3D LiDAR Scene Completion
[CVPR 2023] DiffusioNeRF: Regularizing Neural Radiance Fields with Denoising Diffusion Models
Code for "Restoring Vision in Adverse Weather Conditions with Patch-Based Denoising Diffusion Models" [TPAMI 2023]
[SIGGRAPH Asia 2026] AnyRecon: Arbitrary-View 3D Reconstruction with Video Diffusion Model
[AAAI 2025] MV-VTON: Multi-View Virtual Try-On with Diffusion Models
3D Visualization of an GPT-style LLM
AI-Video-Cropper is a Python-based tool that leverages the power of GPT-4 (OpenAI's language model) to automatically analyze videos, extract the most interesting sections, and crop them for improved viewing experience. This project combines the capabilities of GPT-4, FFmpeg, and OpenCV to automate the process of identifying highlights in videos
A powerful AI library that bridges Large Language Models (LLMs) with video processing frameworks like YOLO, enabling seamless integration of vision and language for advanced video understanding, object detection, and contextual analysis.
(CVPR 2025 highlight✨) Official repository of paper "LLMDet: Learning Strong Open-Vocabulary Object Detectors under the Supervision of Large Language Models"
This is the pytorch implement of our paper "RSPrompter: Learning to Prompt for Remote Sensing Instance Segmentation based on Visual Foundation Model"
🌟A curated list of DUSt3R-related papers and resources, tracking recent advancements using this geometric foundation model.
[NeurIPS'23 Spotlight] Segment Any Point Cloud Sequences by Distilling Vision Foundation Models
The repo for "Metric3D: Towards Zero-shot Metric 3D Prediction from A Single Image" and "Metric3Dv2: A Versatile Monocular Geometric Foundation Model..."
Official implementation of "Lotus: Diffusion-based Visual Foundation Model for High-quality Dense Prediction"
24,538 repositories in the index in total.