Patent · US Active

Linking data elements based on similarity data values and semantic annotations

US10229200B2 · kind B2 · utility

4Cited by
11References
21Claims
0Family size

Assignee

Inventors

Key dates

Filing dateJun 8, 2012
Grant dateMar 12, 2019
Priority date
Expiry dateFeb 28, 2034

Classification

  • Technology area (CPC G)Physics
  • CPC primaryG06F16/951
  • WIPO fieldComputer technology
  • WIPO sectorElectrical engineering

Abstract

Data elements from data sources and having a data value set are linked by using hash functions to determine a dimensionally reduced instance signature for each data element based on all data values associated with that data element to yield a plurality of dimensionally reduced instance signatures of equivalent fixed size such that similarities among the data values in the data value sets across all data elements is maintained among the plurality of instance signatures. Candidate pairs of data elements to link are identified using the plurality of instance signatures in locality sensitive hash functions, and a similarity index is generated for each candidate pair using a pre-determined measure of similarity. Candidate pairs of data elements having a similarity index above a given threshold are linked.

Source: USPTO / EPO open patent data. Objective bibliographic and citation counts.