Patent · US Expired

Method for learning to infer the topical content of documents based upon their lexical content

US5687364A · kind A · utility

84Cited by

16References

19Claims

0Family size

Assignee

Xerox Corporation · US

Inventors

Eric Saund · Redwood City, US
Marti Hearst · Kensington, US

Key dates

Filing date	Sep 16, 1994
Grant date	Nov 11, 1997
Priority date	—
Expiry date	Sep 16, 2014

Classification

Technology area (CPC Y)Emerging Cross-Sectional Technologies
CPC primaryY10S707/99936
WIPO fieldComputer technology
WIPO sectorElectrical engineering

Abstract

An unsupervised method of learning the relationships between words and unspecified topics in documents using a computer is described. The computer represents the relationships between words and unspecified topics via word clusters and association strength values, which can be used later during topical characterization of documents. The computer learns the relationships between words and unspecified topics in an iterative fashion from a set of learning documents. The computer preprocesses the training documents by generating an observed feature vector for each document of the set of training documents and by setting association strengths to initial values. The computer then determines how well the current association strength values predict the topical content of all of the learning documents by generating a cost for each document and summing the individual costs together to generate a total cost. If the total cost is excessive, the association strength values are modified and the total cost recalculated. The computer continues calculating total cost and modifying association strength values until a set of association strength values are discovered that adequately predict the topica…

Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.