[Explainer] What is RLHF, and how does it actually turn a raw language model into something like ChatGPT?

Started by Inference Python, Today at 02:07 PM

Previous topic - Next topic

0 Members and 1 Guest are viewing this topic.

Topic: [Explainer] What is RLHF, and how does it actually turn a raw language model into something like ChatGPT?   Views(Read 57 times)
Active members in this topic:
Inference Python(1)

Inference Python

A raw, freshly pretrained large language model is only trained to predict the next word in a huge dataset of internet text, which means it has no built in sense of what humans actually consider helpful, safe or truthful, since plenty of that training text is exactly the opposite. Reinforcement Learning from Human Feedback, or RLHF, is the technique used to bridge that gap, training a model to produce outputs humans actually prefer rather than just outputs that are statistically plausible continuations of text

The process works by having human evaluators compare pairs of model outputs and mark which one they prefer, or rank several outputs against each other. That preference data is used to train a separate reward model, essentially a system that learns to predict how a human would rate any given output. The original language model is then further trained using reinforcement learning, adjusting its behavior to produce outputs the reward model predicts humans will rate highly, effectively letting human judgment substitute for a hand coded reward function that would be nearly impossible to write out explicitly for something as open ended as good conversation

RLHF has become the central technique for aligning modern chatbots and virtual assistants with human expectations, and it is generally the most resource intensive stage of the whole alignment process, often preceded by a lighter weight step called supervised fine-tuning. It is not without real limitations though, since it depends entirely on the quality and consistency of human feedback, can be gamed if a model learns to satisfy the reward model rather than genuinely being helpful, and researchers are already questioning whether it will scale cleanly to future systems with capabilities well beyond what human evaluators can reliably judge

Save money on everyday spending Free cashback on thousands of retailers
View offer