README.md · RLHFlow/LLaMA3-SFT at d770fe53be6d89592387556cf05a36757f2bc7fc

metadata

library_name: transformers
tags: []

This is the SFT checkpoint used for the project RLHFlow/Online-RLHF

Technical Report: RLHF Workflow: From Reward Modeling to Online RLHF
Authors: Hanze Dong*, Wei Xiong*, Bo Pang*, Haoxiang Wang*, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, Tong Zhang
Code: https://github.com/RLHFlow/Online-RLHF

The model is trained from meta-llama/Meta-Llama-3-8B on a mixture of diverse open-source high-quality data for 1 epoch with detailed parameters in the report. It has not been trained by RLHF and can serve as a good starting point for the RLHF research.