Identifying multiple languages in a content item
US10180935B2 · kind B2 · utility
Assignee
Inventors
Key dates
| Filing date | Feb 2, 2017 |
| Grant date | Jan 15, 2019 |
| Priority date | — |
| Expiry date | Feb 2, 2037 |
Classification
- Technology area (CPC G)Physics
- CPC primaryG06F40/163
- WIPO fieldComputer technology
- WIPO sectorElectrical engineering
Abstract
A system for identifying language(s) for content items is disclosed. The system can identify different languages for content item words segments by identifying segment languages that maximize a probability across the segments. The probability can be a combination of: an author's likelihood for the language identified for the first word; a combination of transition frequencies for selected languages identified for words, the transition frequencies indicating likelihoods that a transition occurred to the selected language from the previous word's language; and a combination of observation probabilities indicating, for a given word in the content item, a likelihood the given word is in the identified language. For an in-vocabulary word, the observation probabilities can be based on learned probability for that word. For an out-of-vocabulary word, the probability can be computed by breaking the word into overlapping n-grams and computing combined learned probabilities that each n-gram is in the given language.
Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.