openai/whisper-large-v3-turbo Automatic Speech Recognition • 0.8B • Updated Oct 4, 2024 • 7.34M • • 3.27k
srinivas2002/rollout_eval_act_v2v3_20260729_20260729_213447 Viewer • Updated 24 days ago • 49.3k • 109 • 1
OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data Paper • 2606.13432 • Published Jun 11 • 113
STREAM: A Data-Centric Framework for Mining High-Value Task-Oriented Dialogues from Streaming Media Paper • 2605.25162 • Published May 24 • 4
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Paper • 2605.28816 • Published May 27 • 433
AR-VLA: True Autoregressive Action Expert for Vision-Language-Action Models Paper • 2603.10126 • Published May 11 • 2
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining Paper • 2605.14747 • Published May 14 • 147