Preprocessing of string inputs in natural language processing
US10372816B2 · kind B2 · utility
Assignee
Inventors
Key dates
| Filing date | Dec 13, 2016 |
| Grant date | Aug 6, 2019 |
| Priority date | — |
| Expiry date | Dec 13, 2036 |
Classification
- Technology area (CPC G)Physics
- CPC primaryG06F40/30
- WIPO fieldComputer technology
- WIPO sectorElectrical engineering
Abstract
Natural language processing of raw text data for optimal sentence boundary placement. Raw text is extracted from a document and subject to cleaning. The extracted raw text is examined to identify preliminary sentence boundaries, which are used to identify potential sentences in the raw text. One or more potential sentences are assigned a well-formedness score. A value of the score correlates to whether the potential sentence is a truncated/ill-formed sentence or a well-formed sentence. One or more preliminary sentence boundaries are optimized depending on the value of the score of the potential sentence(s). Accordingly, the processing herein is an optimization that creates a sentence boundary optimized output.
Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.