YingqingHe/LVDM
quality grade D, 39 out of 100LVDM: Latent Video Diffusion Models for High-Fidelity Long Video Generation
- stars
- 503
- stars gained this week
- —this week
- forks, open issues and contributors
Top AI Repos tracks AI repositories on GitHub and answers two different questions about each one: is it moving right now, and would you bet a product on it.
Video synthesis and editing, talking heads, avatars, animation and 3D asset generation.
Signals: video-generation, text-to-video, video-editing, animation, talking-head, avatar, deepfake, lip-sync
333 results
LVDM: Latent Video Diffusion Models for High-Fidelity Long Video Generation
Vchitect-2.0: Parallel Transformer for Scaling Up Video Diffusion Models
[ICCV 2025, Oral] TrajectoryCrafter: Redirecting Camera Trajectory for Monocular Videos via Diffusion Models
A collection of papers on diffusion models for 3D generation.
ViViD: Video Virtual Try-on using Diffusion Models
(CVPR 2025) From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
MotionDiffuse: Text-Driven Human Motion Generation with Diffusion Model
Official implementation of "DreamPose: Fashion Image-to-Video Synthesis via Stable Diffusion"
🔥🔥🔥自定义Android相机(仿抖音 TikTok),其中功能包括视频人脸识别贴纸,美颜,分段录制,视频裁剪,视频帧处理,获取视频关键帧,视频旋转,添加滤镜,添加水印,合成Gif到视频,文字转视频,图片转视频,音视频合成,音频变声处理,SoundTouch,Fmod音频处理。 Android camera(imitation Tik Tok), which includes video editor,audio editor,video face recognition stickers, segment recording,video cropping, video frame processing, get the first video frame, key frame, video rotation, add filter Mirror ,add watermark ,add gif to video, add text to video, picture to video, audio and video synthesis, audio change processing, SoundTouch, Fmod audio processing.
Allegro is a powerful text-to-video model that generates high-quality videos up to 6 seconds at 15 FPS and 720p resolution from simple text input.
Industry-level video foundation model for unified Text-to-Video (T2V) and Image-to-Video (I2V) generation.
[IJCV] Show-1: Marrying Pixel and Latent Diffusion Models for Text-to-Video Generation
[ICCV 2023] Tune-A-Video: One-Shot Tuning of Image Diffusion Models for Text-to-Video Generation
Puffer is a free live TV streaming website and a research study at Stanford using machine learning to improve video streaming
A deep learning library for video understanding research.
Official implementation for NIPS'17 paper: PredRNN: Recurrent Neural Networks for Predictive Learning Using Spatiotemporal LSTMs.
TransNet V2: Shot Boundary Detection Neural Network
Skill for Agent automating JianYing (CapCut Chinese version) video editing.
🎬 Open source, transcript-based video/audio editor that lives in the browser.
A PIC/FLIP fluid simulation based on the methods found in Robert Bridson's "Fluid Simulation for Computer Graphics"
Official implementation of "MoMask: Generative Masked Modeling of 3D Human Motions (CVPR2024)"
[SIGGRAPH 2022 Journal Track] AvatarCLIP: Zero-Shot Text-Driven Generation and Animation of 3D Avatars
More than 130+ pages in this beautiful app and more than 45 developers has contributed to it.
Adversarial skill embeddings for training reusable controllers for physically simulated characters.
24,523 repositories in the index in total.