Abstract
Intrinsically motivated skill acquisition is an area of computational reinforcement learning (RL) focused on learning multiple skills via the pursuit of self-generated goals. In this context, skills are defined as goal-conditioned policies, and how to learn a diverse set of skills is a key subproblem of intrinsically motivated skill acquisition. In answer, one class of methods uses intrinsic rewards computed as functions of the current goal’s discriminability. To successfully discriminate between different goals, pursuing one goal should typically result in the agent observing regions of the observation space distinct from those observed while pursuing other goals. In theory, this encourages the agent to learn a diverse set of skills without extrinsic rewards. Many such discriminator-based systems select goals uniformly at random during training; yet, non-uniform goal selection has been shown to improve learning, as it enables the agent to focus on developing those skills that need more training. Here, we introduce a method for learning a goal-selection policy: "Diversity Progress" (DP). The method centres around the learner forming a curriculum based on observed improvement in discriminability over its set of goals. Our formalism is related to Learning Progress, designed to select goals that are of learning-optimal difficulty with respect to the agent’s current capabilities. Our proposed method is applicable broadly to the class of discriminability-motivated agents, and we demonstrate empirically that a DP-motivated agent can learn a set of distinguishable skills faster than a similar agent selecting goals uniformly at random, and do so without suffering from a collapse of the goal distribution—a known issue with some prior approaches. The latest write-up of the work is available at https://arxiv.org/abs/2411.01521.
| Original language | English |
|---|---|
| Publication status | Unpublished - 13 Jun 2025 |
| MoE publication type | Not Eligible |
| Event | Multi-disciplinary Conference on Reinforcement Learning and Decision Making - Trinity College Dublin, Dublin, Ireland Duration: 11 Jun 2025 → 14 Jun 2025 https://rldm.org/ |
Conference
| Conference | Multi-disciplinary Conference on Reinforcement Learning and Decision Making |
|---|---|
| Abbreviated title | RLDM |
| Country/Territory | Ireland |
| City | Dublin |
| Period | 11/06/2025 → 14/06/2025 |
| Internet address |
Fingerprint
Dive into the research topics of 'What To Practise Next? Using Diversity Progress To Select From Discriminability-Motivated Goals'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver