Patent · US Active

Preprocessing of text

US8620836B2 · kind B2 · utility

26Cited by

21References

21Claims

0Family size

Assignee

ACCENTURE GLOBAL SERVICES LIMITED · IE

Inventors

Rayid Ghani · Chicago, US
Chad Cumby · Chicago, US
Marko Krema · Evanston, US

Key dates

Filing date	Jan 10, 2011
Grant date	Dec 31, 2013
Priority date	—
Expiry date	Oct 24, 2031

Classification

Technology area (CPC G)Physics
CPC primaryG06F40/205
WIPO fieldComputer technology
WIPO sectorElectrical engineering

Abstract

Performance of statistical machine learning techniques, particularly classification techniques applied to the extraction of attributes and values concerning products, is improved by preprocessing a body of text to be analyzed to remove extraneous information. The body of text is split into a plurality of segments. In an embodiment, sentence identification criteria are applied to identify sentences as the plurality of segments. Thereafter, the plurality of segments are clustered to provide a plurality of clusters. One or more of the resulting clusters are then analyzed to identify segments having low relevance to their respective clusters. Such low relevance segments are then removed from their respective clusters and, consequently, from the body of text. As the resulting relevance-filtered body of text no longer includes portions of the body of text containing mostly extraneous information, the reliability of any subsequent statistical machine learning techniques may be improved.

Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.