Device, method and program for generating accurate corpus data for presentation target for searching
US9645979B2 · kind B2 · utility
Assignee
Inventor
Key dates
| Filing date | Sep 30, 2013 |
| Grant date | May 9, 2017 |
| Priority date | — |
| Expiry date | Sep 30, 2033 |
Classification
- Technology area (CPC G)Physics
- CPC primaryG06F40/30
- WIPO fieldComputer technology
- WIPO sectorElectrical engineering
Abstract
A corpus generation device according to an embodiment includes a web page acquisition unit, a reference word acquisition unit, an attachment unit and an output unit. The web page acquisition unit acquires a web page including description sentence data regarding a presentation target. The reference word acquisition unit acquires a reference word that is an attribute value regarding the presentation target from the web page. The attachment unit extracts a broader word belonging to a layer above the reference word acquired by the reference word acquisition unit from a storage unit that stores hierarchical relationship information indicating a hierarchical relationship between attribute values, and attaches an attribute tag corresponding to the reference word to the broader word included in the description sentence data. The output unit outputs, as corpus data, the description sentence data to which the attribute tag is attached by the attachment unit.
Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.