Skip to main navigation Skip to search Skip to main content

What To Practise Next? Using Diversity Progress To Select From Discriminability-Motivated Goals

Research output: Contribution to conferencePosterScientificpeer-review

Abstract

Intrinsically motivated skill acquisition is an area of computational reinforcement learning (RL) focused on learning multiple skills via the pursuit of self-generated goals. In this context, skills are defined as goal-conditioned policies, and how to learn a diverse set of skills is a key subproblem of intrinsically motivated skill acquisition. In answer, one class of methods uses intrinsic rewards computed as functions of the current goal’s discriminability. To successfully discriminate between different goals, pursuing one goal should typically result in the agent observing regions of the observation space distinct from those observed while pursuing other goals. In theory, this encourages the agent to learn a diverse set of skills without extrinsic rewards. Many such discriminator-based systems select goals uniformly at random during training; yet, non-uniform goal selection has been shown to improve learning, as it enables the agent to focus on developing those skills that need more training. Here, we introduce a method for learning a goal-selection policy: "Diversity Progress" (DP). The method centres around the learner forming a curriculum based on observed improvement in discriminability over its set of goals. Our formalism is related to Learning Progress, designed to select goals that are of learning-optimal difficulty with respect to the agent’s current capabilities. Our proposed method is applicable broadly to the class of discriminability-motivated agents, and we demonstrate empirically that a DP-motivated agent can learn a set of distinguishable skills faster than a similar agent selecting goals uniformly at random, and do so without suffering from a collapse of the goal distribution—a known issue with some prior approaches. The latest write-up of the work is available at https://arxiv.org/abs/2411.01521.
Original languageEnglish
Publication statusUnpublished - 13 Jun 2025
MoE publication typeNot Eligible
EventMulti-disciplinary Conference on Reinforcement Learning and Decision Making - Trinity College Dublin, Dublin, Ireland
Duration: 11 Jun 202514 Jun 2025
https://rldm.org/

Conference

ConferenceMulti-disciplinary Conference on Reinforcement Learning and Decision Making
Abbreviated titleRLDM
Country/TerritoryIreland
CityDublin
Period11/06/202514/06/2025
Internet address

Fingerprint

Dive into the research topics of 'What To Practise Next? Using Diversity Progress To Select From Discriminability-Motivated Goals'. Together they form a unique fingerprint.

Cite this