AutoGUIWorld: Image Generators as Visual World Models for GUI Agent
Abstract
GUI agents require high-quality interaction trajectories to learn how software environments respond to actions, maintain state, and support multi-step workflows. However, the diversity of available trajectories is constrained by the applications, interface states, and workflows accessible in the underlying environments. Expanding this coverage requires deploying increasingly diverse and complex software, with specialized applications imposing additional installation, configuration, and runtime costs. We introduce AutoGUIWorld, a data generation framework that combines the visual priors of image generators with the task knowledge of a planner to synthesize GUI interaction trajectories without deploying or running the corresponding software environments. AutoGUIWorld samples initial GUI scenes from structured specifications of operating-system context, visual appearance, and interface state, and generates tasks conditioned on those scenes. A planner then specifies atomic actions and their intended visual consequences, while an image generator iteratively edits the current screenshot to produce subsequent observations. Action grounding and transition-level quality filtering yield 79,266 spatially annotated step-level training samples across Ubuntu, Windows, macOS, and Chrome. Fine-tuning Qwen3.5-35B-A3B on AutoGUIWorld trajectories improves the mean task score on OSWorld from 33.0% to 40.8% and the task success rate on ScienceBoard from 14.0% to 32.2%. These results show that generated trajectories improve GUI-agent performance on real desktop and scientific tasks.
Community
Image Generators as Visual World Models for GUI Agent
github: https://github.com/ImYangC7/AutoGUIWorld
full paper: https://huggingface.co/YangC777/AGW-35B/blob/main/AutoGUIWorld_Report.pdf
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories (2026)
- StepReflect: Structured UI Transition Reflection for Mobile GUI Agents (2026)
- UI-Mate: Advancing Open-Weight Foundation GUI Agents with In-Context Demonstrations (2026)
- Reflect, Revise, Reuse: Training-Free Skill Evolution for GUI Agents (2026)
- Reflection with Action-Induced Visual Differences for Desktop GUI Agents (2026)
- UI-Venus-2 Technical Report (2026)
- HybridCUA: Learning to Orchestrate GUI and CLI for Computer-Use Agents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 1
Datasets citing this paper 0
No dataset linking this paper