Events
Qualifying ExamInstruction-conditioned RL |
|
||
Wednesday, April 30, 2025, 02:00pm - 03:30pm |
|||
Speaker: Wensen Mao
Bio
Location : CoRE 305
Committee:
Professor Yongfeng Zhang
Professor Dong Deng
Professor Zihan Tan
Professor He Zhu
Event Type: Qualifying Exam
Abstract: Reinforcement learning (RL) holds great promise in enabling autonomous systems to learn complex behaviors through interaction with their environment. However, specifying reward functions for RL is challenging due to the need for precise and often complex definitions. To address this, we propose instruction-guided reinforcement learning, where users specify tasks through instructions, eliminating the need for explicit reward functions. In the first part of this work, we consider Linear Temporal Logic (LTL) formulas to formally define instructions. Existing approaches for finding LTL-satisfying policies rely on sampling a large set of LTL instructions during training to adapt to unseen tasks at inference time, but they do not guarantee generalization to out-of-distribution LTL objectives of increased complexity. We introduce a novel approach to address this challenge, demonstrating that simple goal-conditioned RL agents can follow arbitrary LTL specifications without additional training over the LTL task space. Our approach generalizes to ω-regular expressions. Experimental results show the effectiveness of our strategy in enabling goal-conditioned RL agents to satisfy complex temporal logic task specifications zero-shot. Specifying tasks as LTLs is challenging because it requires deep knowledge of formal logic. To address this, we consider natural language for task instructions in the second part of our work. Previous studies have shown that large language model (LLM) agents can plan long-horizon tasks using a predefined set of skills from task descriptions. However, the need for prior knowledge of the required skill set limits applicability and flexibility. Our approach leverages LLMs to decompose natural language task descriptions into reusable skills, defined by LLM-generated dense reward functions and termination conditions. This facilitates effective skill policy training and chaining for task execution. To address uncertainty in the parameters used by LLMs in the generated reward and termination functions, we train parameter-conditioned skill policies that perform well across a broad spectrum of parameter values. As the impact of these parameters for one skill on the overall task becomes apparent only when subsequent skills are trained, we optimize the most suitable parameter values during the training of subsequent skills to mitigate the risk associated with incorrect parameter choices. Our experimental results show that our method is capable of generating reusable skills to solve a wide range of robot manipulation tasks.List of publicationsCompositional Reinforcement Learning from Logical Specifications. Jothimurugan et al. Learning to Reach Goals via Iterated Supervised Learning. Ghosh et al. Hindsight Experience Replay. Andrychowicz et al. Text2Reward: Automated Dense Reward Function Generation for Reinforcement Learning. Xie et al.Compositional Reinforcement Learning from Logical Specifications. Kishor et al.
Organization:
Contact Professor Yongfeng Zhang
Zoom link
https://rutgers.zoom.us/my/wm300?pwd=RXpFeTlmVjZuOGdEZGVydUl3anV1dz09