Picture for Xavier Fernandes

Xavier Fernandes

Steering Language Model Refusal with Sparse Autoencoders

Add code
Nov 18, 2024
Viaarxiv icon

From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond

Add code
Nov 06, 2024
Figure 1 for From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond
Figure 2 for From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond
Figure 3 for From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond
Figure 4 for From Medprompt to o1: Exploration of Run-Time Strategies for Medical Challenge Problems and Beyond
Viaarxiv icon

TrojanPuzzle: Covertly Poisoning Code-Suggestion Models

Add code
Jan 06, 2023
Figure 1 for TrojanPuzzle: Covertly Poisoning Code-Suggestion Models
Figure 2 for TrojanPuzzle: Covertly Poisoning Code-Suggestion Models
Figure 3 for TrojanPuzzle: Covertly Poisoning Code-Suggestion Models
Figure 4 for TrojanPuzzle: Covertly Poisoning Code-Suggestion Models
Viaarxiv icon