Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Austin Nevins

GRAM: Generative Retrieval Augmented Matching of Data Schemas in the Context of Data Security

Jun 04, 2024

Xuanqing Liu, Luyang Kong, Runhui Wang, Patrick Song, Austin Nevins, Henrik Johnson, Nimish Amlathe, Davor Golac

Figure 1 for GRAM: Generative Retrieval Augmented Matching of Data Schemas in the Context of Data Security

Figure 2 for GRAM: Generative Retrieval Augmented Matching of Data Schemas in the Context of Data Security

Figure 3 for GRAM: Generative Retrieval Augmented Matching of Data Schemas in the Context of Data Security

Figure 4 for GRAM: Generative Retrieval Augmented Matching of Data Schemas in the Context of Data Security

Abstract:Schema matching constitutes a pivotal phase in the data ingestion process for contemporary database systems. Its objective is to discern pairwise similarities between two sets of attributes, each associated with a distinct data table. This challenge emerges at the initial stages of data analytics, such as when incorporating a third-party table into existing databases to inform business insights. Given its significance in the realm of database systems, schema matching has been under investigation since the 2000s. This study revisits this foundational problem within the context of large language models. Adhering to increasingly stringent data security policies, our focus lies on the zero-shot and few-shot scenarios: the model should analyze only a minimal amount of customer data to execute the matching task, contrasting with the conventional approach of scrutinizing the entire data table. We emphasize that the zero-shot or few-shot assumption is imperative to safeguard the identity and privacy of customer data, even at the potential cost of accuracy. The capability to accurately match attributes under such stringent requirements distinguishes our work from previous literature in this domain.

* KDD 2024 Camera Ready; 11 pages, 8 figures

Via

Access Paper or Ask Questions