Patent · US Expired

Text-classification system and method

US7016895B2 · kind B2 · utility

52Cited by
28References
26Claims
0Family size

Assignee

Inventors

Key dates

Filing dateFeb 25, 2003
Grant dateMar 21, 2006
Priority date
Expiry dateMay 6, 2024

Classification

  • Technology area (CPC Y)Emerging Cross-Sectional Technologies
  • CPC primaryY10S707/99935
  • WIPO fieldComputer technology
  • WIPO sectorElectrical engineering

Abstract

Disclosed are a computer-readable code, system and method for classifying a target document in the form of a digitally encoded natural-language text as belonging to one or more of two or more different classes. Each of a plurality of non-generic words and optionally, words groups characterizing the target document is selected as a descriptive term if the term has an above-threshold selectivity value in at least one library of texts in a field, where the selectivity value of a term is a measure of the field-specificity of that term. There is then determined, for each of the plurality of sample texts having associated classification identifiers, a match score related to the number of descriptive terms present in or derived from that text that match those in the target text. From the selected matched texts, and the associated classification identifiers, a classification determination of the target document is made.

Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.