Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Pawan Wirawarn

Problem-Oriented Segmentation and Retrieval: Case Study on Tutoring Conversations

Nov 12, 2024

Rose E. Wang, Pawan Wirawarn, Kenny Lam, Omar Khattab, Dorottya Demszky

Figure 1 for Problem-Oriented Segmentation and Retrieval: Case Study on Tutoring Conversations

Figure 2 for Problem-Oriented Segmentation and Retrieval: Case Study on Tutoring Conversations

Figure 3 for Problem-Oriented Segmentation and Retrieval: Case Study on Tutoring Conversations

Figure 4 for Problem-Oriented Segmentation and Retrieval: Case Study on Tutoring Conversations

Abstract:Many open-ended conversations (e.g., tutoring lessons or business meetings) revolve around pre-defined reference materials, like worksheets or meeting bullets. To provide a framework for studying such conversation structure, we introduce Problem-Oriented Segmentation & Retrieval (POSR), the task of jointly breaking down conversations into segments and linking each segment to the relevant reference item. As a case study, we apply POSR to education where effectively structuring lessons around problems is critical yet difficult. We present LessonLink, the first dataset of real-world tutoring lessons, featuring 3,500 segments, spanning 24,300 minutes of instruction and linked to 116 SAT math problems. We define and evaluate several joint and independent approaches for POSR, including segmentation (e.g., TextTiling), retrieval (e.g., ColBERT), and large language models (LLMs) methods. Our results highlight that modeling POSR as one joint task is essential: POSR methods outperform independent segmentation and retrieval pipelines by up to +76% on joint metrics and surpass traditional segmentation methods by up to +78% on segmentation metrics. We demonstrate POSR's practical impact on downstream education applications, deriving new insights on the language and time use in real-world lesson structures.

* EMNLP 2024 Findings. Our code and dataset are open-sourced at https://github.com/rosewang2008/posr

Via

Access Paper or Ask Questions

Backtracing: Retrieving the Cause of the Query

Mar 06, 2024

Rose E. Wang, Pawan Wirawarn, Omar Khattab, Noah Goodman, Dorottya Demszky

Figure 1 for Backtracing: Retrieving the Cause of the Query

Figure 2 for Backtracing: Retrieving the Cause of the Query

Figure 3 for Backtracing: Retrieving the Cause of the Query

Figure 4 for Backtracing: Retrieving the Cause of the Query

Abstract:Many online content portals allow users to ask questions to supplement their understanding (e.g., of lectures). While information retrieval (IR) systems may provide answers for such user queries, they do not directly assist content creators -- such as lecturers who want to improve their content -- identify segments that _caused_ a user to ask those questions. We introduce the task of backtracing, in which systems retrieve the text segment that most likely caused a user query. We formalize three real-world domains for which backtracing is important in improving content delivery and communication: understanding the cause of (a) student confusion in the Lecture domain, (b) reader curiosity in the News Article domain, and (c) user emotion in the Conversation domain. We evaluate the zero-shot performance of popular information retrieval methods and language modeling methods, including bi-encoder, re-ranking and likelihood-based methods and ChatGPT. While traditional IR systems retrieve semantically relevant information (e.g., details on "projection matrices" for a query "does projecting multiple times still lead to the same point?"), they often miss the causally relevant context (e.g., the lecturer states "projecting twice gets me the same answer as one projection"). Our results show that there is room for improvement on backtracing and it requires new retrieval approaches. We hope our benchmark serves to improve future retrieval systems for backtracing, spawning systems that refine content generation and identify linguistic triggers influencing user queries. Our code and data are open-sourced: https://github.com/rosewang2008/backtracing.

* Code: https://github.com/rosewang2008/backtracing; EACL 2024 Findings, Long Paper

Via

Access Paper or Ask Questions

SIGHT: A Large Annotated Dataset on Student Insights Gathered from Higher Education Transcripts

Jun 15, 2023

Rose E. Wang, Pawan Wirawarn, Noah Goodman, Dorottya Demszky

Figure 1 for SIGHT: A Large Annotated Dataset on Student Insights Gathered from Higher Education Transcripts

Figure 2 for SIGHT: A Large Annotated Dataset on Student Insights Gathered from Higher Education Transcripts

Figure 3 for SIGHT: A Large Annotated Dataset on Student Insights Gathered from Higher Education Transcripts

Figure 4 for SIGHT: A Large Annotated Dataset on Student Insights Gathered from Higher Education Transcripts

Abstract:Lectures are a learning experience for both students and teachers. Students learn from teachers about the subject material, while teachers learn from students about how to refine their instruction. However, online student feedback is unstructured and abundant, making it challenging for teachers to learn and improve. We take a step towards tackling this challenge. First, we contribute a dataset for studying this problem: SIGHT is a large dataset of 288 math lecture transcripts and 15,784 comments collected from the Massachusetts Institute of Technology OpenCourseWare (MIT OCW) YouTube channel. Second, we develop a rubric for categorizing feedback types using qualitative analysis. Qualitative analysis methods are powerful in uncovering domain-specific insights, however they are costly to apply to large data sources. To overcome this challenge, we propose a set of best practices for using large language models (LLMs) to cheaply classify the comments at scale. We observe a striking correlation between the model's and humans' annotation: Categories with consistent human annotations (>$0.9$ inter-rater reliability, IRR) also display higher human-model agreement (>$0.7$), while categories with less consistent human annotations ($0.7$-$0.8$ IRR) correspondingly demonstrate lower human-model agreement ($0.3$-$0.5$). These techniques uncover useful student feedback from thousands of comments, costing around $\$0.002$ per comment. We conclude by discussing exciting future directions on using online student feedback and improving automated annotation techniques for qualitative research.

* First two authors contributed equally. In the Proceedings of Innovative Use of NLP for Building Educational Applications 2023. The code and data are open-sourced here: https://github.com/rosewang2008/sight

Via

Access Paper or Ask Questions