Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Open-Domain Text Evaluation via Meta Distribution Modeling

Jun 20, 2023

Sidi Lu, Asli Celikyilmaz, Tianlu Wang, Nanyun Peng

Figure 1 for Open-Domain Text Evaluation via Meta Distribution Modeling

Figure 2 for Open-Domain Text Evaluation via Meta Distribution Modeling

Figure 3 for Open-Domain Text Evaluation via Meta Distribution Modeling

Figure 4 for Open-Domain Text Evaluation via Meta Distribution Modeling

Share this with someone who'll enjoy it:

Abstract:Recent advances in open-domain text generation models powered by large pre-trained language models (LLMs) have achieved remarkable performance. However, evaluating and controlling these models for desired attributes remains a challenge, as traditional reference-based metrics such as BLEU, ROUGE, and METEOR are insufficient for open-ended generation tasks. Similarly, while trainable discriminator-based evaluation metrics show promise, obtaining high-quality training data is a non-trivial task. In this paper, we introduce a novel approach to evaluate open-domain generation - the Meta-Distribution Methods (MDM). Drawing on the correlation between the rising parameter counts and the improving performance of LLMs, MDM creates a mapping from the contrast of two probabilistic distributions -- one known to be superior to the other -- to quality measures, which can be viewed as a distribution of distributions i.e. Meta-Distribution. We investigate MDM for open-domain text generation evaluation under two paradigms: 1) \emph{Generative} MDM, which leverages the Meta-Distribution Methods to generate in-domain negative samples for training discriminator-based metrics; 2) \emph{Discriminative} MDM, which directly uses distribution discrepancies between two language models for evaluation. Our experiments on multi-turn dialogue and factuality in abstractive summarization demonstrate that MDMs correlate better with human judgment than existing automatic evaluation metrics on both tasks, highlighting the strong performance and generalizability of such methods.

View paper on

Share this with someone who'll enjoy it:

Title:Open-Domain Text Evaluation via Meta Distribution Modeling

Paper and Code