The authors propose giving vision-language models a visual workspace for robot action, extending approaches that generate constraints or programs or provide only a view of the scene.

The proposed approach introduces World Action Agent (WAA), a multi-agent system that utilizes VLMs to control robots using basic tools. It features a visual action workspace with contact views, action rehearsal, and in-view correction mechanisms to refine actions based on observations and feedback.

The authors report acquiring embodied procedural knowledge from expert videos and human teaching, with skills transferring from LIBERO to robosuite. The abstract includes success rates and baseline comparisons. These reported results have not been independently verified for this brief.