Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Forget What You Know about LLMs Evaluations - LLMs are Like a Chameleon

Feb 11, 2025

Nurit Cohen-Inger, Yehonatan Elisha, Bracha Shapira, Lior Rokach, Seffi Cohen

Figure 1 for Forget What You Know about LLMs Evaluations - LLMs are Like a Chameleon

Figure 2 for Forget What You Know about LLMs Evaluations - LLMs are Like a Chameleon

Figure 3 for Forget What You Know about LLMs Evaluations - LLMs are Like a Chameleon

Figure 4 for Forget What You Know about LLMs Evaluations - LLMs are Like a Chameleon

Share this with someone who'll enjoy it:

Abstract:Large language models (LLMs) often appear to excel on public benchmarks, but these high scores may mask an overreliance on dataset-specific surface cues rather than true language understanding. We introduce the Chameleon Benchmark Overfit Detector (C-BOD), a meta-evaluation framework that systematically distorts benchmark prompts via a parametric transformation and detects overfitting of LLMs. By rephrasing inputs while preserving their semantic content and labels, C-BOD exposes whether a model's performance is driven by memorized patterns. Evaluated on the MMLU benchmark using 26 leading LLMs, our method reveals an average performance degradation of 2.15% under modest perturbations, with 20 out of 26 models exhibiting statistically significant differences. Notably, models with higher baseline accuracy exhibit larger performance differences under perturbation, and larger LLMs tend to be more sensitive to rephrasings indicating that both cases may overrely on fixed prompt patterns. In contrast, the Llama family and models with lower baseline accuracy show insignificant degradation, suggesting reduced dependency on superficial cues. Moreover, C-BOD's dataset- and model-agnostic design allows easy integration into training pipelines to promote more robust language understanding. Our findings challenge the community to look beyond leaderboard scores and prioritize resilience and generalization in LLM evaluation.

View paper on

Share this with someone who'll enjoy it:

Title:Forget What You Know about LLMs Evaluations - LLMs are Like a Chameleon

Paper and Code