Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Raj Dabre

Nilekani Centre at AI4Bharat, Indian Institute of Technology Madras, India, National Institute of Information and Communications Technology, Kyoto, Japan, Indian Institute of Technology Bombay, India

PrahokBART: A Pre-trained Sequence-to-Sequence Model for Khmer Natural Language Generation

Dec 15, 2025

Hour Kaing, Raj Dabre, Haiyue Song, Van-Hien Tran, Hideki Tanaka, Masao Utiyama

Abstract:This work introduces {\it PrahokBART}, a compact pre-trained sequence-to-sequence model trained from scratch for Khmer using carefully curated Khmer and English corpora. We focus on improving the pre-training corpus quality and addressing the linguistic issues of Khmer, which are ignored in existing multilingual models, by incorporating linguistic components such as word segmentation and normalization. We evaluate PrahokBART on three generative tasks: machine translation, text summarization, and headline generation, where our results demonstrate that it outperforms mBART50, a strong multilingual pre-trained model. Additionally, our analysis provides insights into the impact of each linguistic module and evaluates how effectively our model handles space during text generation, which is crucial for the naturalness of texts in Khmer.

* Published at COLING 2025, 14 pages

Via

Access Paper or Ask Questions

The Reasoning Lingua Franca: A Double-Edged Sword for Multilingual AI

Oct 23, 2025

Alan Saji, Raj Dabre, Anoop Kunchukuttan, Ratish Puduppully

Abstract:Large Reasoning Models (LRMs) achieve strong performance on mathematical, scientific, and other question-answering tasks, but their multilingual reasoning abilities remain underexplored. When presented with non-English questions, LRMs often default to reasoning in English, raising concerns about interpretability and the handling of linguistic and cultural nuances. We systematically compare an LRM's reasoning in English versus the language of the question. Our evaluation spans two tasks: MGSM and GPQA Diamond. Beyond measuring answer accuracy, we also analyze cognitive attributes in the reasoning traces. We find that English reasoning traces exhibit a substantially higher presence of these cognitive behaviors, and that reasoning in English generally yields higher final-answer accuracy, with the performance gap increasing as tasks become more complex. However, this English-centric strategy is susceptible to a key failure mode - getting "Lost in Translation," where translation steps lead to errors that would have been avoided by question's language reasoning.

* 14 pages, 13 figures, 5 tables

Via

Access Paper or Ask Questions

Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts

Jun 04, 2025

Sidharth Pulipaka, Sparsh Jain, Ashwin Sankar, Raj Dabre

Figure 1 for Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts

Figure 2 for Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts

Figure 3 for Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts

Figure 4 for Mark My Words: A Robust Multilingual Model for Punctuation in Text and Speech Transcripts

Abstract:Punctuation plays a vital role in structuring meaning, yet current models often struggle to restore it accurately in transcripts of spontaneous speech, especially in the presence of disfluencies such as false starts and backtracking. These limitations hinder the performance of downstream tasks like translation, text to speech, summarization, etc. where sentence boundaries are critical for preserving quality. In this work, we introduce Cadence, a generalist punctuation restoration model adapted from a pretrained large language model. Cadence is designed to handle both clean written text and highly spontaneous spoken transcripts. It surpasses the previous state of the art in performance while expanding support from 14 to all 22 Indian languages and English. We conduct a comprehensive analysis of model behavior across punctuation types and language families, identifying persistent challenges under domain shift and with rare punctuation marks. Our findings demonstrate the efficacy of utilizing pretrained language models for multilingual punctuation restoration and highlight Cadence practical value for low resource NLP pipelines at scale.

* Work in Progress

Via

Access Paper or Ask Questions

Limited-Resource Adapters Are Regularizers, Not Linguists

May 30, 2025

Marcell Fekete, Nathaniel R. Robinson, Ernests Lavrinovics, E. Djeride Jean-Baptiste, Raj Dabre, Johannes Bjerva, Heather Lent

Abstract:Cross-lingual transfer from related high-resource languages is a well-established strategy to enhance low-resource language technologies. Prior work has shown that adapters show promise for, e.g., improving low-resource machine translation (MT). In this work, we investigate an adapter souping method combined with cross-attention fine-tuning of a pre-trained MT model to leverage language transfer for three low-resource Creole languages, which exhibit relatedness to different language groups across distinct linguistic dimensions. Our approach improves performance substantially over baselines. However, we find that linguistic relatedness -- or even a lack thereof -- does not covary meaningfully with adapter performance. Surprisingly, our cross-attention fine-tuning approach appears equally effective with randomly initialized adapters, implying that the benefit of adapters in this setting lies in parameter regularization, and not in meaningful information transfer. We provide analysis supporting this regularization hypothesis. Our findings underscore the reality that neural language processing involves many success factors, and that not all neural methods leverage linguistic knowledge in intuitive ways.

Via

Access Paper or Ask Questions

CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

May 30, 2025

Emilio Villa-Cueva, Sholpan Bolatzhanova, Diana Turmakhan, Kareem Elzeky, Henok Biadglign Ademtew, Alham Fikri Aji, Israel Abebe Azime, Jinheon Baek, Frederico Belcavello, Fermin Cristobal(+24 more)

Figure 1 for CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

Figure 2 for CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

Figure 3 for CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

Figure 4 for CaMMT: Benchmarking Culturally Aware Multimodal Machine Translation

Abstract:Cultural content poses challenges for machine translation systems due to the differences in conceptualizations between cultures, where language alone may fail to convey sufficient context to capture region-specific meanings. In this work, we investigate whether images can act as cultural context in multimodal translation. We introduce CaMMT, a human-curated benchmark of over 5,800 triples of images along with parallel captions in English and regional languages. Using this dataset, we evaluate five Vision Language Models (VLMs) in text-only and text+image settings. Through automatic and human evaluations, we find that visual context generally improves translation quality, especially in handling Culturally-Specific Items (CSIs), disambiguation, and correct gender usage. By releasing CaMMT, we aim to support broader efforts in building and evaluating multimodal translation systems that are better aligned with cultural nuance and regional variation.

Via

Access Paper or Ask Questions

TikZero: Zero-Shot Text-Guided Graphics Program Synthesis

Mar 14, 2025

Jonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka, Masao Utiyama, Raj Dabre, Steffen Eger, Simone Paolo Ponzetto

Figure 1 for TikZero: Zero-Shot Text-Guided Graphics Program Synthesis

Figure 2 for TikZero: Zero-Shot Text-Guided Graphics Program Synthesis

Figure 3 for TikZero: Zero-Shot Text-Guided Graphics Program Synthesis

Figure 4 for TikZero: Zero-Shot Text-Guided Graphics Program Synthesis

Abstract:With the rise of generative AI, synthesizing figures from text captions becomes a compelling application. However, achieving high geometric precision and editability requires representing figures as graphics programs in languages like TikZ, and aligned training data (i.e., graphics programs with captions) remains scarce. Meanwhile, large amounts of unaligned graphics programs and captioned raster images are more readily available. We reconcile these disparate data sources by presenting TikZero, which decouples graphics program generation from text understanding by using image representations as an intermediary bridge. It enables independent training on graphics programs and captioned images and allows for zero-shot text-guided graphics program synthesis during inference. We show that our method substantially outperforms baselines that can only operate with caption-aligned graphics programs. Furthermore, when leveraging caption-aligned graphics programs as a complementary training signal, TikZero matches or exceeds the performance of much larger models, including commercial systems like GPT-4o. Our code, datasets, and select models are publicly available.

* Project page: https://github.com/potamides/DeTikZify

Via

Access Paper or Ask Questions

IteRABRe: Iterative Recovery-Aided Block Reduction

Mar 08, 2025

Haryo Akbarianto Wibowo, Haiyue Song, Hideki Tanaka, Masao Utiyama, Alham Fikri Aji, Raj Dabre

Figure 1 for IteRABRe: Iterative Recovery-Aided Block Reduction

Figure 2 for IteRABRe: Iterative Recovery-Aided Block Reduction

Figure 3 for IteRABRe: Iterative Recovery-Aided Block Reduction

Figure 4 for IteRABRe: Iterative Recovery-Aided Block Reduction

Abstract:Large Language Models (LLMs) have grown increasingly expensive to deploy, driving the need for effective model compression techniques. While block pruning offers a straightforward approach to reducing model size, existing methods often struggle to maintain performance or require substantial computational resources for recovery. We present IteRABRe, a simple yet effective iterative pruning method that achieves superior compression results while requiring minimal computational resources. Using only 2.5M tokens for recovery, our method outperforms baseline approaches by ~3% on average when compressing the Llama3.1-8B and Qwen2.5-7B models. IteRABRe demonstrates particular strength in the preservation of linguistic capabilities, showing an improvement 5% over the baselines in language-related tasks. Our analysis reveals distinct pruning characteristics between these models, while also demonstrating preservation of multilingual capabilities.

* 8 pages

Via

Access Paper or Ask Questions

RomanLens: Latent Romanization and its role in Multilinguality in LLMs

Feb 11, 2025

Alan Saji, Jaavid Aktar Husain, Thanmay Jayakumar, Raj Dabre, Anoop Kunchukuttan, Mitesh M. Khapra, Ratish Puduppully

Abstract:Large Language Models (LLMs) exhibit remarkable multilingual generalization despite being predominantly trained on English-centric corpora. A fundamental question arises: how do LLMs achieve such robust multilingual capabilities? For non-Latin script languages, we investigate the role of romanization - the representation of non-Latin scripts using Latin characters - as a bridge in multilingual processing. Using mechanistic interpretability techniques, we analyze next-token generation and find that intermediate layers frequently represent target words in romanized form before transitioning to native script, a phenomenon we term Latent Romanization. Further, through activation patching experiments, we demonstrate that LLMs encode semantic concepts similarly across native and romanized scripts, suggesting a shared underlying representation. Additionally in translation towards non Latin languages, our findings reveal that when the target language is in romanized form, its representations emerge earlier in the model's layers compared to native script. These insights contribute to a deeper understanding of multilingual representation in LLMs and highlight the implicit role of romanization in facilitating language transfer. Our work provides new directions for potentially improving multilingual language modeling and interpretability.

* 18 pages, 18 figures

Via

Access Paper or Ask Questions

Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages

Dec 14, 2024

Poulami Ghosh, Raj Dabre, Pushpak Bhattacharyya

Figure 1 for Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages

Figure 2 for Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages

Figure 3 for Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages

Figure 4 for Are Language Models Agnostic to Linguistically Grounded Perturbations? A Case Study of Indic Languages

Abstract:Pre-trained language models (PLMs) are known to be susceptible to perturbations to the input text, but existing works do not explicitly focus on linguistically grounded attacks, which are subtle and more prevalent in nature. In this paper, we study whether PLMs are agnostic to linguistically grounded attacks or not. To this end, we offer the first study addressing this, investigating different Indic languages and various downstream tasks. Our findings reveal that although PLMs are susceptible to linguistic perturbations, when compared to non-linguistic attacks, PLMs exhibit a slightly lower susceptibility to linguistic attacks. This highlights that even constrained attacks are effective. Moreover, we investigate the implications of these outcomes across a range of languages, encompassing diverse language families and different scripts.

* Work in Progress

Via

Access Paper or Ask Questions

Pralekha: An Indic Document Alignment Evaluation Benchmark

Nov 28, 2024

Sanjay Suryanarayanan, Haiyue Song, Mohammed Safi Ur Rahman Khan, Anoop Kunchukuttan, Mitesh M. Khapra, Raj Dabre

Figure 1 for Pralekha: An Indic Document Alignment Evaluation Benchmark

Figure 2 for Pralekha: An Indic Document Alignment Evaluation Benchmark

Figure 3 for Pralekha: An Indic Document Alignment Evaluation Benchmark

Figure 4 for Pralekha: An Indic Document Alignment Evaluation Benchmark

Abstract:Mining parallel document pairs poses a significant challenge because existing sentence embedding models often have limited context windows, preventing them from effectively capturing document-level information. Another overlooked issue is the lack of concrete evaluation benchmarks comprising high-quality parallel document pairs for assessing document-level mining approaches, particularly for Indic languages. In this study, we introduce Pralekha, a large-scale benchmark for document-level alignment evaluation. Pralekha includes over 2 million documents, with a 1:2 ratio of unaligned to aligned pairs, covering 11 Indic languages and English. Using Pralekha, we evaluate various document-level mining approaches across three dimensions: the embedding models, the granularity levels, and the alignment algorithm. To address the challenge of aligning documents using sentence and chunk-level alignments, we propose a novel scoring method, Document Alignment Coefficient (DAC). DAC demonstrates substantial improvements over baseline pooling approaches, particularly in noisy scenarios, achieving average gains of 20-30% in precision and 15-20% in F1 score. These results highlight DAC's effectiveness in parallel document mining for Indic languages.

* Work in Progress

Via

Access Paper or Ask Questions