Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Barbara Kovačić

MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank

Mar 15, 2024

Verena Blaschke, Barbara Kovačić, Siyao Peng, Hinrich Schütze, Barbara Plank

Figure 1 for MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank

Figure 2 for MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank

Figure 3 for MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank

Figure 4 for MaiBaam: A Multi-Dialectal Bavarian Universal Dependency Treebank

Abstract:Despite the success of the Universal Dependencies (UD) project exemplified by its impressive language breadth, there is still a lack in `within-language breadth': most treebanks focus on standard languages. Even for German, the language with the most annotations in UD, so far no treebank exists for one of its language varieties spoken by over 10M people: Bavarian. To contribute to closing this gap, we present the first multi-dialect Bavarian treebank (MaiBaam) manually annotated with part-of-speech and syntactic dependency information in UD, covering multiple text genres (wiki, fiction, grammar examples, social, non-fiction). We highlight the morphosyntactic differences between the closely-related Bavarian and German and showcase the rich variability of speakers' orthographies. Our corpus includes 15k tokens, covering dialects from all Bavarian-speaking areas spanning three countries. We provide baseline parsing and POS tagging results, which are lower than results obtained on German and vary substantially between different graph-based parsers. To support further research on Bavarian syntax, we make our dataset, language-specific guidelines and code publicly available.

* LREC-COLING 2024

Via

Access Paper or Ask Questions

MaiBaam Annotation Guidelines

Mar 09, 2024

Verena Blaschke, Barbara Kovačić, Siyao Peng, Barbara Plank

Abstract:This document provides the annotation guidelines for MaiBaam, a Bavarian corpus annotated with part-of-speech (POS) tags and syntactic dependencies. MaiBaam belongs to the Universal Dependencies (UD) project, and our annotations elaborate on the general and German UD version 2 guidelines. In this document, we detail how to preprocess and tokenize Bavarian data, provide an overview of the POS tags and dependencies we use, explain annotation decisions that would also apply to closely related languages like German, and lastly we introduce and motivate decisions that are specific to Bavarian grammar.

Via

Access Paper or Ask Questions