Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

M. Iftekhar Tanveer

Ehsan

Unsupervised Speaker Diarization that is Agnostic to Language, Overlap-Aware, and Tuning Free

Jul 25, 2022

M. Iftekhar Tanveer, Diego Casabuena, Jussi Karlgren, Rosie Jones

Figure 1 for Unsupervised Speaker Diarization that is Agnostic to Language, Overlap-Aware, and Tuning Free

Figure 2 for Unsupervised Speaker Diarization that is Agnostic to Language, Overlap-Aware, and Tuning Free

Figure 3 for Unsupervised Speaker Diarization that is Agnostic to Language, Overlap-Aware, and Tuning Free

Figure 4 for Unsupervised Speaker Diarization that is Agnostic to Language, Overlap-Aware, and Tuning Free

Abstract:Podcasts are conversational in nature and speaker changes are frequent -- requiring speaker diarization for content understanding. We propose an unsupervised technique for speaker diarization without relying on language-specific components. The algorithm is overlap-aware and does not require information about the number of speakers. Our approach shows 79% improvement on purity scores (34% on F-score) against the Google Cloud Platform solution on podcast data.

* Published at Interspeech 2022

Via

Access Paper or Ask Questions

Automated Analysis and Prediction of Job Interview Performance

Apr 14, 2015

Iftekhar Naim, M. Iftekhar Tanveer, Daniel Gildea, Mohammed, Hoque

Figure 1 for Automated Analysis and Prediction of Job Interview Performance

Figure 2 for Automated Analysis and Prediction of Job Interview Performance

Figure 3 for Automated Analysis and Prediction of Job Interview Performance

Figure 4 for Automated Analysis and Prediction of Job Interview Performance

Abstract:We present a computational framework for automatically quantifying verbal and nonverbal behaviors in the context of job interviews. The proposed framework is trained by analyzing the videos of 138 interview sessions with 69 internship-seeking undergraduates at the Massachusetts Institute of Technology (MIT). Our automated analysis includes facial expressions (e.g., smiles, head gestures, facial tracking points), language (e.g., word counts, topic modeling), and prosodic information (e.g., pitch, intonation, and pauses) of the interviewees. The ground truth labels are derived by taking a weighted average over the ratings of 9 independent judges. Our framework can automatically predict the ratings for interview traits such as excitement, friendliness, and engagement with correlation coefficients of 0.75 or higher, and can quantify the relative importance of prosody, language, and facial expressions. By analyzing the relative feature weights learned by the regression models, our framework recommends to speak more fluently, use less filler words, speak as "we" (vs. "I"), use more unique words, and smile more. We also find that the students who were rated highly while answering the first interview question were also rated highly overall (i.e., first impression matters). Finally, our MIT Interview dataset will be made available to other researchers to further validate and expand our findings.

* 14 pages, 8 figures, 6 tables

Via

Access Paper or Ask Questions