Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Stefano Mezza

OpenDebateEvidence: A Massive-Scale Argument Mining and Summarization Dataset

Jun 20, 2024

Allen Roush, Yusuf Shabazz, Arvind Balaji, Peter Zhang, Stefano Mezza, Markus Zhang, Sanjay Basu, Sriram Vishwanath, Mehdi Fatemi, Ravid Schwartz-Ziv

Figure 1 for OpenDebateEvidence: A Massive-Scale Argument Mining and Summarization Dataset

Figure 2 for OpenDebateEvidence: A Massive-Scale Argument Mining and Summarization Dataset

Figure 3 for OpenDebateEvidence: A Massive-Scale Argument Mining and Summarization Dataset

Figure 4 for OpenDebateEvidence: A Massive-Scale Argument Mining and Summarization Dataset

Abstract:We introduce OpenDebateEvidence, a comprehensive dataset for argument mining and summarization sourced from the American Competitive Debate community. This dataset includes over 3.5 million documents with rich metadata, making it one of the most extensive collections of debate evidence. OpenDebateEvidence captures the complexity of arguments in high school and college debates, providing valuable resources for training and evaluation. Our extensive experiments demonstrate the efficacy of fine-tuning state-of-the-art large language models for argumentative abstractive summarization across various methods, models, and datasets. By providing this comprehensive resource, we aim to advance computational argumentation and support practical applications for debaters, educators, and researchers. OpenDebateEvidence is publicly available to support further research and innovation in computational argumentation. Access it here: https://huggingface.co/datasets/Yusuf5/OpenCaselist

* Accepted for Publication to ARGMIN 2024 at ACL2024

Via

Access Paper or Ask Questions

ISO-Standard Domain-Independent Dialogue Act Tagging for Conversational Agents

Jun 12, 2018

Stefano Mezza, Alessandra Cervone, Giuliano Tortoreto, Evgeny A. Stepanov, Giuseppe Riccardi

Figure 1 for ISO-Standard Domain-Independent Dialogue Act Tagging for Conversational Agents

Figure 2 for ISO-Standard Domain-Independent Dialogue Act Tagging for Conversational Agents

Figure 3 for ISO-Standard Domain-Independent Dialogue Act Tagging for Conversational Agents

Figure 4 for ISO-Standard Domain-Independent Dialogue Act Tagging for Conversational Agents

Abstract:Dialogue Act (DA) tagging is crucial for spoken language understanding systems, as it provides a general representation of speakers' intents, not bound to a particular dialogue system. Unfortunately, publicly available data sets with DA annotation are all based on different annotation schemes and thus incompatible with each other. Moreover, their schemes often do not cover all aspects necessary for open-domain human-machine interaction. In this paper, we propose a methodology to map several publicly available corpora to a subset of the ISO standard, in order to create a large task-independent training corpus for DA classification. We show the feasibility of using this corpus to train a domain-independent DA tagger testing it on out-of-domain conversational data, and argue the importance of training on multiple corpora to achieve robustness across different DA categories.

Via

Access Paper or Ask Questions