Patent · US Active

Identifying multiple languages in a content item

US10180935B2 · kind B2 · utility

0Cited by
24References
19Claims
0Family size

Assignee

Inventors

Key dates

Filing dateFeb 2, 2017
Grant dateJan 15, 2019
Priority date
Expiry dateFeb 2, 2037

Classification

  • Technology area (CPC G)Physics
  • CPC primaryG06F40/163
  • WIPO fieldComputer technology
  • WIPO sectorElectrical engineering

Abstract

A system for identifying language(s) for content items is disclosed. The system can identify different languages for content item words segments by identifying segment languages that maximize a probability across the segments. The probability can be a combination of: an author's likelihood for the language identified for the first word; a combination of transition frequencies for selected languages identified for words, the transition frequencies indicating likelihoods that a transition occurred to the selected language from the previous word's language; and a combination of observation probabilities indicating, for a given word in the content item, a likelihood the given word is in the identified language. For an in-vocabulary word, the observation probabilities can be based on learned probability for that word. For an out-of-vocabulary word, the probability can be computed by breaking the word into overlapping n-grams and computing combined learned probabilities that each n-gram is in the given language.

Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.