RLHFlow
/

LLaMA3-SFT

Text Generation

text-generation-inference

Inference Endpoints

Model card Files Files and versions Community

LLaMA3-SFT / README.md

Haoxiang-Wang's picture

Update README.md

0a31e34 verified 14 days ago

|

history blame contribute delete

781 Bytes

	---
	library_name: transformers
	tags: []
	---

	This is the SFT checkpoint used for the project [RLHFlow/Online-RLHF](https://github.com/RLHFlow/Online-RLHF)

	* Paper: [RLHF Workflow: From Reward Modeling to Online RLHF](https://arxiv.org/pdf/2405.07863) (Published in TMLR, 2024)
	* Authors: Hanze Dong, Wei Xiong, Bo Pang, Haoxiang Wang, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, Tong Zhang
	* Code: https://github.com/RLHFlow/Online-RLHF

	The model is trained from [meta-llama/Meta-Llama-3-8B](https://huggingface.co/meta-llama/Meta-Llama-3-8B) on a mixture of diverse open-source high-quality data for 1 epoch with detailed parameters in the report. It has not been trained by RLHF and can serve as a good starting point for the RLHF research.