Patent · US Active

Visual language models for perception

US12158762B1 · kind B1 · utility

2Cited by
0References
20Claims
0Family size

Assignee

Inventors

Key dates

Filing dateSep 11, 2023
Grant dateDec 3, 2024
Priority date
Expiry dateSep 11, 2043

Classification

  • Technology area (CPC G)Physics
  • CPC primaryG06V20/588
  • WIPO fieldComputer technology
  • WIPO sectorElectrical engineering

Abstract

A method is provided, that includes: receiving camera data from a perception system of an autonomous vehicle; and providing the camera data to a visual language model, where the visual language model includes a mapping of a corpus of images and a corpus of text to a common parameter space. The method further includes: receiving from the visual language model an output corresponding to one or more text tokens; accessing a configuration file comprising a plurality of text tokens representing a plurality of objects or events of interest to the autonomous vehicle; and identifying a respective object or event of interest in an environment of the autonomous vehicle by determining that a text token of the output matches a respective one of the plurality of text tokens in the configuration file. The autonomous vehicle can then be controlled based at least in part on the respective object or event of interest.

Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.