Picture for Nicola Cancedda

Nicola Cancedda

AI Research Agents for Machine Learning: Search, Exploration, and Generalization in MLE-bench

Add code
Jul 03, 2025
Viaarxiv icon

Don't Make It Up: Preserving Ignorance Awareness in LLM Fine-Tuning

Add code
Jun 17, 2025
Viaarxiv icon

HalluLens: LLM Hallucination Benchmark

Add code
Apr 24, 2025
Viaarxiv icon

Calibrating Verbal Uncertainty as a Linear Feature to Reduce Hallucinations

Add code
Mar 18, 2025
Viaarxiv icon

LUNAR: LLM Unlearning via Neural Activation Redirection

Add code
Feb 11, 2025
Viaarxiv icon

Robust LLM safeguarding via refusal feature adversarial training

Add code
Sep 30, 2024
Figure 1 for Robust LLM safeguarding via refusal feature adversarial training
Figure 2 for Robust LLM safeguarding via refusal feature adversarial training
Figure 3 for Robust LLM safeguarding via refusal feature adversarial training
Figure 4 for Robust LLM safeguarding via refusal feature adversarial training
Viaarxiv icon

Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources

Add code
Sep 12, 2024
Figure 1 for Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources
Figure 2 for Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources
Figure 3 for Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources
Figure 4 for Source2Synth: Synthetic Data Generation and Curation Grounded in Real Data Sources
Viaarxiv icon

PISTOL: Dataset Compilation Pipeline for Structural Unlearning of LLMs

Add code
Jun 24, 2024
Figure 1 for PISTOL: Dataset Compilation Pipeline for Structural Unlearning of LLMs
Figure 2 for PISTOL: Dataset Compilation Pipeline for Structural Unlearning of LLMs
Figure 3 for PISTOL: Dataset Compilation Pipeline for Structural Unlearning of LLMs
Figure 4 for PISTOL: Dataset Compilation Pipeline for Structural Unlearning of LLMs
Viaarxiv icon

Know When To Stop: A Study of Semantic Drift in Text Generation

Add code
Apr 08, 2024
Viaarxiv icon

Spectral Filters, Dark Signals, and Attention Sinks

Add code
Feb 14, 2024
Viaarxiv icon