Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Iva Vukojević

You Are What You Talk About: Inducing Evaluative Topics for Personality Analysis

Feb 01, 2023

Josip Jukić, Iva Vukojević, Jan Šnajder

Abstract:Expressing attitude or stance toward entities and concepts is an integral part of human behavior and personality. Recently, evaluative language data has become more accessible with social media's rapid growth, enabling large-scale opinion analysis. However, surprisingly little research examines the relationship between personality and evaluative language. To bridge this gap, we introduce the notion of evaluative topics, obtained by applying topic models to pre-filtered evaluative text from social media. We then link evaluative topics to individual text authors to build their evaluative profiles. We apply evaluative profiling to Reddit comments labeled with personality scores and conduct an exploratory study on the relationship between evaluative topics and Big Five personality facets, aiming for a more interpretable, facet-level analysis. Finally, we validate our approach by observing correlations consistent with prior research in personality psychology.

* Accepted at EMNLP 2022 (Findings), NLP+CSS

Via

Access Paper or Ask Questions

PANDORA Talks: Personality and Demographics on Reddit

Apr 27, 2020

Matej Gjurković, Mladen Karan, Iva Vukojević, Mihaela Bošnjak, Jan Šnajder

Figure 1 for PANDORA Talks: Personality and Demographics on Reddit

Figure 2 for PANDORA Talks: Personality and Demographics on Reddit

Figure 3 for PANDORA Talks: Personality and Demographics on Reddit

Figure 4 for PANDORA Talks: Personality and Demographics on Reddit

Abstract:Personality and demographics are important variables in social sciences, while in NLP they can aid in interpretability and removal of societal biases. However, datasets with both personality and demographic labels are scarce. To address this, we present PANDORA, the first large-scale dataset of Reddit comments labeled with three personality models (including the well-established Big 5 model) and demographics (age, gender, and location) for more than 10k users. We showcase the usefulness of this dataset on three experiments, where we leverage the more readily available data from other personality models to predict the Big 5 traits, analyze gender classification biases arising from psycho-demographic variables, and carry out a confirmatory and exploratory analysis based on psychological theories. Finally, we present benchmark prediction models for all personality and demographic variables.

Via

Access Paper or Ask Questions