Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:DART: Open-Domain Structured Data Record to Text Generation

Jul 06, 2020

Dragomir Radev, Rui Zhang, Amrit Rau, Abhinand Sivaprasad, Chiachun Hsieh, Nazneen Fatema Rajani, Xiangru Tang, Aadit Vyas, Neha Verma, Pranav Krishna(+13 more)

Figure 1 for DART: Open-Domain Structured Data Record to Text Generation

Figure 2 for DART: Open-Domain Structured Data Record to Text Generation

Figure 3 for DART: Open-Domain Structured Data Record to Text Generation

Figure 4 for DART: Open-Domain Structured Data Record to Text Generation

Share this with someone who'll enjoy it:

Abstract:We introduce DART, a large dataset for open-domain structured data record to text generation. We consider the structured data record input as a set of RDF entity-relation triples, a format widely used for knowledge representation and semantics description. DART consists of 82,191 examples across different domains with each input being a semantic RDF triple set derived from data records in tables and the tree ontology of the schema, annotated with sentence descriptions that cover all facts in the triple set. This hierarchical, structured format with its open-domain nature differentiates DART from other existing table-to-text corpora. We conduct an analysis of DART on several state-of-the-art text generation models, showing that it introduces new and interesting challenges compared to existing datasets. Furthermore, we demonstrate that finetuning pretrained language models on DART facilitates out-of-domain generalization on the WebNLG 2017 dataset. DART is available at https://github.com/Yale-LILY/dart.

View paper on

Share this with someone who'll enjoy it:

Title:DART: Open-Domain Structured Data Record to Text Generation

Paper and Code