🏆 SERIES

RLHF

Human preference alignment for LLMs: reward modeling, policy optimization, data collection, evaluation, and safe deployment.

Parts

5

Category

🏆

Type

Series