Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Myeong Cheol Shin

OutFlip: Generating Out-of-Domain Samples for Unknown Intent Detection with Natural Language Attack

May 12, 2021

DongHyun Choi, Myeong Cheol Shin, EungGyun Kim, Dong Ryeol Shin

Figure 1 for OutFlip: Generating Out-of-Domain Samples for Unknown Intent Detection with Natural Language Attack

Figure 2 for OutFlip: Generating Out-of-Domain Samples for Unknown Intent Detection with Natural Language Attack

Figure 3 for OutFlip: Generating Out-of-Domain Samples for Unknown Intent Detection with Natural Language Attack

Figure 4 for OutFlip: Generating Out-of-Domain Samples for Unknown Intent Detection with Natural Language Attack

Abstract:Out-of-domain (OOD) input detection is vital in a task-oriented dialogue system since the acceptance of unsupported inputs could lead to an incorrect response of the system. This paper proposes OutFlip, a method to generate out-of-domain samples using only in-domain training dataset automatically. A white-box natural language attack method HotFlip is revised to generate out-of-domain samples instead of adversarial examples. Our evaluation results showed that integrating OutFlip-generated out-of-domain samples into the training dataset could significantly improve an intent classification model's out-of-domain detection performance.

* 9 pages, 3 figures; to be appear in ACL Findings of ACL-IJCNLP 2021

Via

Access Paper or Ask Questions

Auxiliary Sequence Labeling Tasks for Disfluency Detection

Oct 24, 2020

Dongyub Lee, Byeongil Ko, Myeong Cheol Shin, Taesun Whang, Daniel Lee, Eun Hwa Kim, EungGyun Kim, Jaechoon Jo

Figure 1 for Auxiliary Sequence Labeling Tasks for Disfluency Detection

Figure 2 for Auxiliary Sequence Labeling Tasks for Disfluency Detection

Figure 3 for Auxiliary Sequence Labeling Tasks for Disfluency Detection

Figure 4 for Auxiliary Sequence Labeling Tasks for Disfluency Detection

Abstract:Detecting disfluencies in spontaneous speech is an important preprocessing step in natural language processing and speech recognition applications. In this paper, we propose a method utilizing named entity recognition (NER) and part-of-speech (POS) as auxiliary sequence labeling (SL) tasks for disfluency detection. First, we show that training a disfluency detection model with auxiliary SL tasks can improve its F-score in disfluency detection. Then, we analyze which auxiliary SL tasks are influential depending on baseline models. Experimental results on the widely used English Switchboard dataset show that our method outperforms the previous state-of-the-art in disfluency detection.

* 5 pages, 3 figures, 3 tables

Via

Access Paper or Ask Questions

Integrated Eojeol Embedding for Erroneous Sentence Classification in Korean Chatbots

Apr 13, 2020

DongHyun Choi, IlNam Park, Myeong Cheol Shin, EungGyun Kim, Dong Ryeol Shin

Figure 1 for Integrated Eojeol Embedding for Erroneous Sentence Classification in Korean Chatbots

Figure 2 for Integrated Eojeol Embedding for Erroneous Sentence Classification in Korean Chatbots

Figure 3 for Integrated Eojeol Embedding for Erroneous Sentence Classification in Korean Chatbots

Figure 4 for Integrated Eojeol Embedding for Erroneous Sentence Classification in Korean Chatbots

Abstract:This paper attempts to analyze the Korean sentence classification system for a chatbot. Sentence classification is the task of classifying an input sentence based on predefined categories. However, spelling or space error contained in the input sentence causes problems in morphological analysis and tokenization. This paper proposes a novel approach of Integrated Eojeol (Korean syntactic word separated by space) Embedding to reduce the effect that poorly analyzed morphemes may make on sentence classification. It also proposes two noise insertion methods that further improve classification performance. Our evaluation results indicate that the proposed system classifies erroneous sentences more accurately than the baseline system by 17%p.0

* 9 pages, 2 figures

Via

Access Paper or Ask Questions

RYANSQL: Recursively Applying Sketch-based Slot Fillings for Complex Text-to-SQL in Cross-Domain Databases

Apr 07, 2020

DongHyun Choi, Myeong Cheol Shin, EungGyun Kim, Dong Ryeol Shin

Figure 1 for RYANSQL: Recursively Applying Sketch-based Slot Fillings for Complex Text-to-SQL in Cross-Domain Databases

Figure 2 for RYANSQL: Recursively Applying Sketch-based Slot Fillings for Complex Text-to-SQL in Cross-Domain Databases

Figure 3 for RYANSQL: Recursively Applying Sketch-based Slot Fillings for Complex Text-to-SQL in Cross-Domain Databases

Figure 4 for RYANSQL: Recursively Applying Sketch-based Slot Fillings for Complex Text-to-SQL in Cross-Domain Databases

Abstract:Text-to-SQL is the problem of converting a user question into an SQL query, when the question and database are given. In this paper, we present a neural network approach called RYANSQL (Recursively Yielding Annotation Network for SQL) to solve complex Text-to-SQL tasks for cross-domain databases. State-ment Position Code (SPC) is defined to trans-form a nested SQL query into a set of non-nested SELECT statements; a sketch-based slot filling approach is proposed to synthesize each SELECT statement for its corresponding SPC. Additionally, two input manipulation methods are presented to improve generation performance further. RYANSQL achieved 58.2% accuracy on the challenging Spider benchmark, which is a 3.2%p improvement over previous state-of-the-art approaches. At the time of writing, RYANSQL achieves the first position on the Spider leaderboard.

* 10 pages, 1 figure

Via

Access Paper or Ask Questions