Real-Time Action Chunking
Teaching Robots to Listen and Think Harder

Teaching Robots to Listen and Think Harder

2/26/2025

What this post added

Introduces the 'Hi Robot' system, a hierarchical inference process for VLA models. This system uses a high-level VLM ('System 2') to reason through complex prompts and language, generating intermediate steps for a low-level VLA model ('System 1'). The high-level policy is trained using synthetic data and can incorporate real-time user feedback. The post presents quantitative results showing improved instruction-following accuracy and task progress compared to flat VLA policies and GPT-4o.

Read the original post ↗