Events

Qualifying Exam

Instruction-conditioned RL

 

Download as iCal file

Wednesday, April 30, 2025, 02:00pm - 03:30pm

 

Speaker: Wensen Mao

Bio

Location : CoRE 305

Committee

Professor Yongfeng Zhang

Professor  Dong Deng

Professor  Zihan Tan

Professor  He Zhu

Event Type: Qualifying Exam

Abstract: Reinforcement learning (RL) holds great promise in enabling autonomous systems to learn complex behaviors through interaction with their environment. However, specifying reward functions for RL is challenging due to the need for precise and often complex definitions. To address this, we propose instruction-guided reinforcement learning, where users specify tasks through instructions, eliminating the need for explicit reward functions. In the first part of this work, we consider Linear Temporal Logic (LTL) formulas to formally define instructions. Existing approaches for finding LTL-satisfying policies rely on sampling a large set of LTL instructions during training to adapt to unseen tasks at inference time, but they do not guarantee generalization to out-of-distribution LTL objectives of increased complexity. We introduce a novel approach to address this challenge, demonstrating that simple goal-conditioned RL agents can follow arbitrary LTL specifications without additional training over the LTL task space. Our approach generalizes to ω-regular expressions. Experimental results show the effectiveness of our strategy in enabling goal-conditioned RL agents to satisfy complex temporal logic task specifications zero-shot. Specifying tasks as LTLs is challenging because it requires deep knowledge of formal logic. To address this, we consider natural language for task instructions in the second part of our work. Previous studies have shown that large language model (LLM) agents can plan long-horizon tasks using a predefined set of skills from task descriptions. However, the need for prior knowledge of the required skill set limits applicability and flexibility. Our approach leverages LLMs to decompose natural language task descriptions into reusable skills, defined by LLM-generated dense reward functions and termination conditions. This facilitates effective skill policy training and chaining for task execution. To address uncertainty in the parameters used by LLMs in the generated reward and termination functions, we train parameter-conditioned skill policies that perform well across a broad spectrum of parameter values. As the impact of these parameters for one skill on the overall task becomes apparent only when subsequent skills are trained, we optimize the most suitable parameter values during the training of subsequent skills to mitigate the risk associated with incorrect parameter choices. Our experimental results show that our method is capable of generating reusable skills to solve a wide range of robot manipulation tasks.List of publicationsCompositional Reinforcement Learning from Logical Specifications. Jothimurugan et al. Learning to Reach Goals via Iterated Supervised Learning. Ghosh et al. Hindsight Experience Replay. Andrychowicz et al. Text2Reward: Automated Dense Reward Function Generation for Reinforcement Learning. Xie et al.Compositional Reinforcement Learning from Logical Specifications. Kishor et al.

Organization

Contact  Professor Yongfeng Zhang

Zoom link
https://rutgers.zoom.us/my/wm300?pwd=RXpFeTlmVjZuOGdEZGVydUl3anV1dz09