Patent · US Expired

Methods, systems, and articles of manufacture for soft hierarchical clustering of co-occurring objects

US7644102B2 · kind B2 · utility

13Cited by
11References
23Claims
0Family size

Assignee

Inventors

Key dates

Filing dateOct 19, 2001
Grant dateJan 5, 2010
Priority date
Expiry dateFeb 17, 2024

Classification

  • Technology area (CPC Y)Emerging Cross-Sectional Technologies
  • CPC primaryY10S707/99945
  • WIPO fieldComputer technology
  • WIPO sectorElectrical engineering

Abstract

Methods, systems, and articles of manufacture consistent with certain principles related to the present invention enable a computing system to perform hierarchical topical clustering of text data based on statistical modeling of co-occurrences of (document, word) pairs. The computing system may be configured to receive a collection of documents, each document including a plurality of words, and perform a modified deterministic annealing Expectation-Maximization (EM) process on the collection to produce a softly assigned hierarchy of nodes. The process may involve assigning documents and document fragments to multiple nodes in the hierarchy based on words included in the documents, such that a document may be assigned to any ancestor node included in the hierarchy, thus eliminating the hard assignment of documents in the hierarchy.

Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.