Sampling response pairs

PROTRAILBLAZER

RLHF vs DPO · Preference Optimization
PAIRS0
AGREEMENT--
PREF PROB0.17
KL BASE0.000
STEP0
RM LOSS--
Conceptual · simulated values
PIPELINE · RLHFSIMPLIFIED
01 · COLLECT
Preference data
0 pairs · batch 0/4
02 · FIT
Reward model
loss --
03 · UPDATE
Policy update π
τ = 0.18 KL leash
ANCHOR
Reference π₀ · frozen
KL leash anchor
Conceptual modelProduction RLHF and DPO train billion-parameter models on large datasets with evaluation and safety review. Two style axes stand in for many preference dimensions, and many variants exist.
RESPONSE PAIRPAIR #0
PROMPT “Why is the sky blue?”
CHOOSE A RESPONSE, OR LET THE AUTO LABELER RUN
POLICY LANDSCAPECONCEPTUAL PROJECTION
Labeler persona
Batch size
Label noise