Patent · US Active

Device and method to improve learning of a policy for robots

US12246450B2 · kind B2 · utility

0Cited by

1References

10Claims

0Family size

Assignee

Robert Bosch GmbH · DE

Inventors

Felix Berkenkamp · Munich, DE
Lukas Froehlich · Freiburg (Elbe), DE
Maksym Lefarov · Stuttgart, DE
Andreas Doerr · Stuttgart, DE

Key dates

Filing date	Mar 1, 2022
Grant date	Mar 11, 2025
Priority date	—
Expiry date	Jul 1, 2043

Classification

Technology area (CPC G)Physics
CPC primaryG06N3/08
WIPO fieldComputer technology
WIPO sectorElectrical engineering

Abstract

A computer-implemented method for for learning a policy. The method includes: recording at least an episode of interactions of the agent with its environment following policy and adding the recorded episode to a set of training data; optimizing a transition dynamics model based on the training data such that the transition dynamics model predicts the next states of the environment depending on the states and actions contained in the training data; optimizing policy parameters based on the training data and the transition dynamics model by optimizing a reward. In the method, the transition dynamics model comprises a first model characterizing the global model and a second model characterizing a correction model, which is configured to correct outputs of the first model.

Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.