Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System Paper • 2609.01607 • Published Sep 1 • 24
PAWBench: How Far Are We from Probabilistically Aligned World Modeling? Paper • 2608.27345 • Published Aug 27 • 77
PAWBench: How Far Are We from Probabilistically Aligned World Modeling? Paper • 2608.27345 • Published Aug 27 • 77
DynEval: Holistic Evaluations of T2I Generative Models in the Wild Paper • 2607.11199 • Published Jul 13 • 1
4KLSDB: A Large-Scale Dataset for 4K Image Restoration and Generation Paper • 2605.24762 • Published May 23 • 1
Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models Paper • 2607.04461 • Published Jul 5 • 11
Flash-BoN: Instant Drafts for Inference-Time Scaling in Diffusion Models Paper • 2607.04461 • Published Jul 5 • 11
RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space Paper • 2606.14700 • Published Jun 12 • 18
Video Probabilistic Diffusion Models in Projected Latent Space Paper • 2302.07685 • Published Feb 15, 2023
Controllable Human Image Generation with Personalized Multi-Garments Paper • 2411.16801 • Published Nov 25, 2024 • 3
Enhancing Motion Dynamics of Image-to-Video Models via Adaptive Low-Pass Guidance Paper • 2506.08456 • Published Jun 10, 2025 • 2
ContextVLA: Vision-Language-Action Model with Amortized Multi-Frame Context Paper • 2510.04246 • Published Oct 5, 2025 • 1
Decoupled MeanFlow: Turning Flow Models into Flow Maps for Accelerated Sampling Paper • 2510.24474 • Published Oct 28, 2025 • 1
DreamDojo: A Generalist Robot World Model from Large-Scale Human Videos Paper • 2602.06949 • Published Feb 6 • 37