Trained on a different random sampling of the same datasets used by loyal-piano-m7, then with cDPO on a blend of RLHF datasets.

Several intermediate checkpoints (of cDPO training) are on branches.

Uses the Alpaca prompt format.

Downloads last month: 755

Safetensors

Model size

7.24B params

Tensor type

BF16

Inference Examples

Text Generation

This model does not have enough activity to be deployed to Inference API (serverless) yet. Increase its social visibility and check back later, or deploy to Inference Endpoints (dedicated) instead.

Model tree for chargoddard/servile-harpsichord-cdpo

Merges

1 model

chargoddard
/

servile-harpsichord-cdpo

Model tree for chargoddard/servile-harpsichord-cdpo

Datasets used to train chargoddard/servile-harpsichord-cdpo

Spaces using chargoddard/servile-harpsichord-cdpo 5