RLHFlow
/

LLaMA3-SFT

Text Generation

text-generation-inference

Inference Endpoints

Model card Files Files and versions Community

Haoxiang-Wang commited on Oct 14

Commit

d770fe5

•

1 Parent(s): ead1cff

Update README.md

Files changed (1) hide show

README.md +5 -1

README.md CHANGED Viewed

@@ -3,7 +3,11 @@ library_name: transformers
 tags: []
 ---
-This is the SFT checkpoint used for the project [Online-RLHF](https://github.com/RLHFlow/Online-RLHF). Also check our [technical report here](https://arxiv.org/pdf/2405.07863).
 The model is trained from [meta-llama/Meta-Llama-3-8B](https://huggingface.co/meta-llama/Meta-Llama-3-8B) on a mixture of diverse open-source high-quality data for 1 epoch with detailed parameters in the report. It has not been trained by RLHF and can serve as a good starting point for the RLHF research.

 tags: []
 ---
+This is the SFT checkpoint used for the project [RLHFlow/Online-RLHF](https://github.com/RLHFlow/Online-RLHF)
+* **Technical Report**: [RLHF Workflow: From Reward Modeling to Online RLHF](https://arxiv.org/pdf/2405.07863)
+* **Authors**: Hanze Dong*, Wei Xiong*, Bo Pang*, Haoxiang Wang*, Han Zhao, Yingbo Zhou, Nan Jiang, Doyen Sahoo, Caiming Xiong, Tong Zhang
+* **Code**: https://github.com/RLHFlow/Online-RLHF
 The model is trained from [meta-llama/Meta-Llama-3-8B](https://huggingface.co/meta-llama/Meta-Llama-3-8B) on a mixture of diverse open-source high-quality data for 1 epoch with detailed parameters in the report. It has not been trained by RLHF and can serve as a good starting point for the RLHF research.