Patent · US Active

Method and system for simplifying implicit rhetorical relation prediction in large scale annotated corpus

US9355372B2 · kind B2 · utility

3Cited by
4References
28Claims
0Family size

Assignee

Inventors

Key dates

Filing dateJul 3, 2014
Grant dateMay 31, 2016
Priority date
Expiry dateJul 25, 2034

Classification

  • Technology area (CPC G)Physics
  • CPC primaryG06N20/00
  • WIPO fieldComputer technology
  • WIPO sectorElectrical engineering

Abstract

The present invention provides a method and system directed to predicting implicit rhetorical relations between two spans of text, e.g., in a large annotated corpus, such as the Penn Discourse Treebank (“PDTB”), Rhetorical Structure Theory corpus, and the Discourse Graph Bank, and particularly directed to determining a rhetorical relation in the absence of an explicit discourse marker. Surface level features may be used to capture pragmatic information encoded in the absent marker. In one manner a simplified feature set based only on raw text and semantic dependencies is used to improve performance for all relations. By using surface level features to predict implicit rhetorical relations for the large annotated corpus the invention approaches a theoretical maximum performance, suggesting that more data will not necessarily improve performance based on these and similarly situated features.

Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.