tzwilliam0/maxmin-dpo-init-kl-coef-0.5-fix-reward-norm-dongnan Reinforcement Learning • Updated 15 days ago • 7
tzwilliam0/maxmin-dpo-init-kl-coef-0.1-fix-reward-norm-dongnan Reinforcement Learning • Updated 15 days ago • 7
tzwilliam0/maxmin-dpo-init-kl-coef-0.5-fix-lora-dongnan Reinforcement Learning • Updated 21 days ago • 61
tzwilliam0/maxmin-dpo-init-kl-coef-0.1-fix-lora-dongnan Reinforcement Learning • Updated 21 days ago • 65