Picture for Tengyu Xu

Tengyu Xu

Think Smarter not Harder: Adaptive Reasoning with Inference Aware Optimization

Add code
Jan 31, 2025
Viaarxiv icon

HyperZero: A Customized End-to-End Auto-Tuning System for Recommendation with Hourly Feedback

Add code
Jan 30, 2025
Figure 1 for HyperZero: A Customized End-to-End Auto-Tuning System for Recommendation with Hourly Feedback
Figure 2 for HyperZero: A Customized End-to-End Auto-Tuning System for Recommendation with Hourly Feedback
Figure 3 for HyperZero: A Customized End-to-End Auto-Tuning System for Recommendation with Hourly Feedback
Figure 4 for HyperZero: A Customized End-to-End Auto-Tuning System for Recommendation with Hourly Feedback
Viaarxiv icon

Step-KTO: Optimizing Mathematical Reasoning through Stepwise Binary Feedback

Add code
Jan 18, 2025
Viaarxiv icon

Multi-IF: Benchmarking LLMs on Multi-Turn and Multilingual Instructions Following

Add code
Oct 21, 2024
Viaarxiv icon

The Perfect Blend: Redefining RLHF with Mixture of Judges

Add code
Sep 30, 2024
Figure 1 for The Perfect Blend: Redefining RLHF with Mixture of Judges
Figure 2 for The Perfect Blend: Redefining RLHF with Mixture of Judges
Figure 3 for The Perfect Blend: Redefining RLHF with Mixture of Judges
Figure 4 for The Perfect Blend: Redefining RLHF with Mixture of Judges
Viaarxiv icon

Provably Efficient Offline Reinforcement Learning with Trajectory-Wise Reward

Add code
Jun 13, 2022
Viaarxiv icon

Model-Based Offline Meta-Reinforcement Learning with Regularization

Add code
Feb 07, 2022
Figure 1 for Model-Based Offline Meta-Reinforcement Learning with Regularization
Figure 2 for Model-Based Offline Meta-Reinforcement Learning with Regularization
Figure 3 for Model-Based Offline Meta-Reinforcement Learning with Regularization
Figure 4 for Model-Based Offline Meta-Reinforcement Learning with Regularization
Viaarxiv icon

Faster Algorithm and Sharper Analysis for Constrained Markov Decision Process

Add code
Oct 20, 2021
Figure 1 for Faster Algorithm and Sharper Analysis for Constrained Markov Decision Process
Viaarxiv icon

PER-ETD: A Polynomially Efficient Emphatic Temporal Difference Learning Method

Add code
Oct 13, 2021
Figure 1 for PER-ETD: A Polynomially Efficient Emphatic Temporal Difference Learning Method
Figure 2 for PER-ETD: A Polynomially Efficient Emphatic Temporal Difference Learning Method
Figure 3 for PER-ETD: A Polynomially Efficient Emphatic Temporal Difference Learning Method
Figure 4 for PER-ETD: A Polynomially Efficient Emphatic Temporal Difference Learning Method
Viaarxiv icon

A Unified Off-Policy Evaluation Approach for General Value Function

Add code
Jul 06, 2021
Figure 1 for A Unified Off-Policy Evaluation Approach for General Value Function
Figure 2 for A Unified Off-Policy Evaluation Approach for General Value Function
Viaarxiv icon