less than 1 minute read

Maowei Jiang, Qi Wang, Hongfeng Ai, Zhiyong Dong, Yusong Hu, Ao Liang, Yifan Wang, Ruiqi Li, Quangao Liu, Moquan Chen, Peter Búš, Long Zeng.

Information Fusion, 127, Part C, 103923 (2026).

Read the paper · Official code

Prompt2Act uses vision-language models, visual grounding, and a hybrid execution agent to translate multimodal prompts into sequences of robotic actions.

Updated: