Patent · US Active

Custom display post processing in speech recognition

US12061861B2 · kind B2 · utility

0Cited by

1References

20Claims

0Family size

Assignee

MICROSOFT TECHNOLOGY LICENSING, LLC · US

Inventors

Wei-Min Liu · Dublin, US
Padma Varadharajan · Palo Alto, US
Piyush BEHRE · Santa Clara, US
Nicholas Kibre · Redwood City, US
Edward C. Lin · Beijing, CN
Shuangyu Chang · Fremont, US
Che ZHAO · Beijing, CN
Khuram Shahid · Seattle, US
Heiko Rahmel · Bellevue, US

Key dates

Filing date	Jul 26, 2022
Grant date	Aug 13, 2024
Priority date	—
Expiry date	Jul 26, 2042

Classification

Technology area (CPC G)Physics
CPC primaryG10L15/26
WIPO fieldComputer technology
WIPO sectorElectrical engineering

Abstract

Solutions for custom display post processing (DPP) in speech recognition (SR) use a customized multi-stage DPP pipeline that transforms a stream of SR tokens from lexical form to display form. A first transformation stage of the DPP pipeline receives the stream of tokens, in turn, by an upstream filter, a base model stage, and a downstream filter, and transforms a first aspect of the stream of tokens (e.g., disfluency, inverse text normalization (ITN), capitalization, etc.) from lexical form into display form. The upstream filter and/or the downstream filter alter the stream of tokens to change the default behavior of the DPP pipeline into custom behavior. Additional transformation stages of the DPP pipeline perform further transforms, allowing for outputting final text in a display format that is customized for a specific user. This permits each user to efficiently leverage a common baseline DPP pipeline to produce a custom output.

Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.