Language-Driven Multi-Task Manipulation with
Action-Mask-Enhanced Multimodal Learning

Abstract

For long-horizon multi-task robotic manipulation, hierarchical approaches provide an effective way to combine high-level language-based task planning with low-level vision- language based sub-task execution. Then, we propose a frame- work that integrates a two-stage task planner with a multimodal low-level action planner incorporating an explicit action-mask policy. At high level, a Vision-Language Model (VLM) first perceives object and scene information from observations, and a Large Language Model (LLM) then reasons over this together with a task library and human instruction to generate a textual task plan. This two-stage design mitigates modality bias between perception and planning. At low level, an asymmetric multimodal encoder, SigLIP2 with Weight-Decomposed Low- Rank Adaptation (DoRA) for text and multi-view ResNets for vision, feeds into an Action Chunking with Transformer (ACT)- based policy enhanced by Temperature-Scaled Spatial Attention and Bidirectional Cross-Attention for language-vision fusion. We further introduce an explicit action-mask policy that jointly predicts actions and their validity, enabling real-time sub-task termination detection and robust switching across multiple sub-tasks without additional inference overhead. Experiments on weighing and multi-object manipulation tasks demonstrate performance, with ablation studies validating the contribution of each component. Generalization experiments under out- of-distribution conditions highlight both the strengths and limitations. Deployment on a distinct dual-arm robotic platform in a new scenario validates transferability.


Human-Language-Guided Multimodal Autonomous Bimanual Multi-Task Robotic Manipulation

Demos include weighing, multi-object, and drawer manipulation scenarios

Covering human instructions of varying difficulty levels

Approach Overview

Approach Overview

Fig. 1. Overview of the proposed hierarchical framework for language-guided, multimodal, multi-task robotic manipulation. The framework operates in three processes: (I) Instruction Recognition, where natural language commands are transcribed using a Whisper-based voice-to-text module; (II) High-Level Task Planning, where a VLM first perceives the scene to extract objects and information, and then an LLM, conditioned on the human instruction, the VLM output, and a library of predefined skills, decomposes the command into a sequence of sub-tasks; (III) Low-Level Action Generation and Sub-task Switching, where an ACT-based multimodal language-vision policy executes each sub-task by leveraging textual input, visual observations, and robot joint states, while a Task Completion Checker detects termination and triggers the transition to the next sub-task. The right shows two examples of decomposed human instructions and their sub-task executions.

Two-stage high-level task planner

Fig. 2. Two-stage high-level planner: the VLM parses the scene, and the LLM uses it with the human instruction and sub-task library to generate sub-tasks paln list with a brief reply.

llmsact

Fig. 3. Architecture of the Low-Level Action Planner. For a sub-task (e.g., “open the balance”), we adopt an asymmetric encoder: the language input is encoded by a frozen SigLIP2 augmented with trainable DoRA adapters, while multi-view images at time t are processed by a fine-tuned ResNet. The resulting features are fused via (i) a temperature-scaled spatial attention for visual enhancement and (ii) a bidirectional cross-attention for language-vision integration. These fused features, together with the robot joint state, are passed to an ACT-based policy, whose output is normalized and augmented with two additional linear heads predicting the next k timesteps of joint actions and their corresponding validity masks. A Task-Completion Checker monitors the masks and terminates the current sub-task if more than n consecutive invalid actions are detected, switching to the next sub-task.