Skip to content

AI & Intelligence

From labels to judgment: how supervised learning grew into RLHF

Supervised learning has always been a conversation between humans and machines. What has changed is how much the machine is asking.

Gopal Bhat, Chief Strategy Officer1 min read

As AI agents are built for specific industries, human input during data preparation remains the mechanism that keeps a model aligned with human values and preferences.

Where it started: labels

The first generation of supervised learning ran on binary and categorical labels. Is this image a cat or a dog? Is this email spam or not? A human decided, the machine learned the boundary, and the quality of the decision set the ceiling on the model.

Richer annotation for computer vision

Computer vision raised the bar. Object detection, the automotive industry’s core problem, needed bounding boxes. Understanding a scene needed pixel-level segmentation masks and written scene descriptions. Annotation became skilled work, and data preparation became a discipline in its own right.

Language models and human feedback

Large language models and generative AI changed the question again. For tasks like translation there is no single right label, only better and worse answers. Reinforcement learning from human feedback asks people to compare and rank outputs, and the model learns a preference rather than a category. The human is no longer labelling data; the human is supervising judgment.

What stays constant

The tools have changed at every stage. The role of the person has not. As AI agents are developed for specific industries, human input during training-data preparation remains essential to aligning those systems with human values and preferences. The organisations that treat that input as expert work, and organise it accordingly, will build the models that people trust.

In this insight

  • Supervised learning
  • Data labelling
  • Computer vision
  • Large language models
  • RLHF

Visual explainer

The idea as one route

  1. Business problem

    Binary and categorical labels

    Cat or dog, spam or not

  2. Technology

    Bounding boxes and segmentation

    Object detection and scene understanding

  3. Branta thinking

    Human feedback (RLHF)

    Ranking outputs, teaching preference

  4. Measured outcome

    Aligned AI

    Models that reflect human values

Have a business problem? Let's figure out the technology.

Tell us what is not working. We will work backwards from the outcome you need.