Patent · US Active

Evaluating reinforcement learning policies

US10445653B1 · kind B1 · utility

6Cited by

0References

20Claims

0Family size

Assignee

DeepMind Technologies Limited · GB

Inventors

Joel William Veness · London, GB
Marc Gendron-Bellemare · London, GB

Key dates

Filing date	Aug 7, 2015
Grant date	Oct 15, 2019
Priority date	—
Expiry date	May 20, 2037

Classification

Technology area (CPC G)Physics
CPC primaryG06N5/022
WIPO fieldComputer technology
WIPO sectorElectrical engineering

Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for evaluating reinforcement learning policies. One of the methods includes receiving a plurality of training histories for a reinforcement learning agent; determining a total reward for each training observation in the training histories; partitioning the training observations into a plurality of partitions; determining, for each partition and from the partitioned training observations, a probability that the reinforcement learning agent will receive the total reward for the partition if the reinforcement learning agent performs the action for the partition in response to receiving the current observation; determining, from the probabilities and for each total reward, a respective estimated value of performing each action in response to receiving the current observation; and selecting an action from the pre-determined set of actions from the estimated values in accordance with an action selection policy.

Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.