Prompt2Act: Mapping prompts into sequence of robotic actions with large foundation models
Maowei Jiang, Qi Wang, Hongfeng Ai, Zhiyong Dong, Yusong Hu, Ao Liang, Yifan Wang, Ruiqi Li, Quangao Liu, Moquan Chen, Peter Búš, Long Zeng.
Information Fusion, 127, Part C, 103923 (2026).
Read the paper · Official code
Prompt2Act uses vision-language models, visual grounding, and a hybrid execution agent to translate multimodal prompts into sequences of robotic actions.