Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:A Bi-objective Perspective on Controllable Language Models: Reward Dropout Improves Off-policy Control Performance

Oct 06, 2023

Changhun Lee, Chiehyeon Lim

Figure 1 for A Bi-objective Perspective on Controllable Language Models: Reward Dropout Improves Off-policy Control Performance

Figure 2 for A Bi-objective Perspective on Controllable Language Models: Reward Dropout Improves Off-policy Control Performance

Figure 3 for A Bi-objective Perspective on Controllable Language Models: Reward Dropout Improves Off-policy Control Performance

Figure 4 for A Bi-objective Perspective on Controllable Language Models: Reward Dropout Improves Off-policy Control Performance

Share this with someone who'll enjoy it:

Abstract:We study the theoretical aspects of CLMs (Controllable Language Models) from a bi-objective optimization perspective. Specifically, we consider the CLMs as an off-policy RL problem that requires simultaneously maximizing the reward and likelihood objectives. Our main contribution consists of three parts. First, we establish the theoretical foundations of CLM by presenting reward upper bound and Pareto improvement/optimality conditions. Second, we analyze conditions that improve and violate Pareto optimality itself, respectively. Finally, we propose Reward Dropout, a simple yet powerful method to guarantee policy improvement based on a Pareto improvement condition. Our theoretical outcomes are supported by not only deductive proofs but also empirical results. The performance of Reward Dropout was evaluated on five CLM benchmark datasets, and it turns out that the Reward Dropout significantly improves the performance of CLMs.

* 25 pages, 14 figures, conference

View paper on

Share this with someone who'll enjoy it:

Title:A Bi-objective Perspective on Controllable Language Models: Reward Dropout Improves Off-policy Control Performance

Paper and Code