Picture for Xiangru Peng

Xiangru Peng

Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Add code
Jun 26, 2024
Figure 1 for Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Figure 2 for Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Figure 3 for Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Figure 4 for Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Viaarxiv icon