跳到正文

StageGuard: Learning Stage Transitions for Long-Horizon Robot Tasks via Agentic Distillation

机器人 0 次浏览

原文StageGuard: Learning Stage Transitions for Long-Horizon Robot Tasks via Agentic Distillation

作者:Jinbang Huang, Yuanzhao Hu, Zhiyuan Li, Ran Qi, Yixin Xiao, Yangzheng Wu, Tengyue Ba, Zhanguang Zhang, Yingxue Zhang

来源:arXiv cs.RO(机器人)

正文

Computer Science > Robotics

arXiv:2609.20791v1 (cs)

[Submitted on 17 Sep 2026]

Title:StageGuard: Learning Stage Transitions for Long-Horizon Robot Tasks via Agentic Distillation

Authors:Jinbang Huang, Yuanzhao Hu, Zhiyuan Li, Ran Qi, Yixin Xiao, Yangzheng Wu, Tengyue Ba, Zhanguang Zhang, Yingxue Zhang

View a PDF of the paper titled StageGuard: Learning Stage Transitions for Long-Horizon Robot Tasks via Agentic Distillation, by Jinbang Huang and 8 other authors

View PDF

HTML (experimental)

Abstract:Hierarchical planning frameworks combine skills from multiple robot control policies for long-horizon task execution, where determining when to terminate the current skill and advance to the next subtask is essential. Existing approaches often rely on pre-designed completion signal checkers that are hard to obtain in real-world execution. Large-scale vision-language models (VLMs) offer strong reasoning capabilities, but their decision boundaries are not inherently aligned with task completion criteria, while cloud deployment and lengthy reasoning introduce substantial latency, limiting real-time monitoring. We propose StageGuard, an agentic distillation framework for accurate and efficient stage-transition decisions. StageGuard combines teacher-model reasoning with demonstration trajectories to generate structured explanations of subtask completion and policy switching. A lightweight student VLM uses these explanations to generate compact self-explanations, which are used for supervised fine-tuning. We evaluate stage-transition prediction on trajectories from two benchmarks and assess closed-loop task success through integration into hierarchical robot control on BEHAVIOR-1K, with further validation on real robots. Results show substantial improvements in stage-transition prediction while supporting efficient online monitoring.

Comments:

8 pages, 2 figures

Subjects:

Robotics (cs.RO)

Cite as:

arXiv:2609.20791 [cs.RO]

(or

arXiv:2609.20791v1 [cs.RO] for this version)

https://doi.org/10.48550/arXiv.2609.20791

Focus to learn more

arXiv-issued DOI via DataCite (pending registration)

主题

机器人


由「前沿雷达」于 2026-09-20 采集。正文取自原文页面,已保留出处链接。标题与正文版权归原作者所有。

评论

加载中…