Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Darren Cook

Optimising Human-Machine Collaboration for Efficient High-Precision Information Extraction from Text Documents

Feb 18, 2023

Bradley Butcher, Miri Zilka, Darren Cook, Jiri Hron, Adrian Weller

Figure 1 for Optimising Human-Machine Collaboration for Efficient High-Precision Information Extraction from Text Documents

Figure 2 for Optimising Human-Machine Collaboration for Efficient High-Precision Information Extraction from Text Documents

Figure 3 for Optimising Human-Machine Collaboration for Efficient High-Precision Information Extraction from Text Documents

Figure 4 for Optimising Human-Machine Collaboration for Efficient High-Precision Information Extraction from Text Documents

Abstract:While humans can extract information from unstructured text with high precision and recall, this is often too time-consuming to be practical. Automated approaches, on the other hand, produce nearly-immediate results, but may not be reliable enough for high-stakes applications where precision is essential. In this work, we consider the benefits and drawbacks of various human-only, human-machine, and machine-only information extraction approaches. We argue for the utility of a human-in-the-loop approach in applications where high precision is required, but purely manual extraction is infeasible. We present a framework and an accompanying tool for information extraction using weak-supervision labelling with human validation. We demonstrate our approach on three criminal justice datasets. We find that the combination of computer speed and human understanding yields precision comparable to manual annotation while requiring only a fraction of time, and significantly outperforms fully automated baselines in terms of precision.

Via

Access Paper or Ask Questions

Can We Automate the Analysis of Online Child Sexual Exploitation Discourse?

Sep 25, 2022

Darren Cook, Miri Zilka, Heidi DeSandre, Susan Giles, Adrian Weller, Simon Maskell

Figure 1 for Can We Automate the Analysis of Online Child Sexual Exploitation Discourse?

Figure 2 for Can We Automate the Analysis of Online Child Sexual Exploitation Discourse?

Figure 3 for Can We Automate the Analysis of Online Child Sexual Exploitation Discourse?

Figure 4 for Can We Automate the Analysis of Online Child Sexual Exploitation Discourse?

Abstract:Social media's growing popularity raises concerns around children's online safety. Interactions between minors and adults with predatory intentions is a particularly grave concern. Research into online sexual grooming has often relied on domain experts to manually annotate conversations, limiting both scale and scope. In this work, we test how well-automated methods can detect conversational behaviors and replace an expert human annotator. Informed by psychological theories of online grooming, we label $6772$ chat messages sent by child-sex offenders with one of eleven predatory behaviors. We train bag-of-words and natural language inference models to classify each behavior, and show that the best performing models classify behaviors in a manner that is consistent, but not on-par, with human annotation.

Via

Access Paper or Ask Questions