Back to Discover
Curiosity

Is Modern AI Actually Unsupervised?

A clear separation between the textbook meaning of 'unsupervised learning' and the hybrid training pipelines behind modern AI, supported by concrete examples of each stage.

Before you enter

A complete interactive classroom, not just a preview.

Start when you are ready to enter this Stage's 9 scenes and explore, respond, and learn as you go.

9
Scenes
18 min
Estimated
Content language: en-US
Start this Stage
Sign-in may be required to play
What happens inside
  1. 01The Word on Everyone's Lipsslide
    Question

    Open the investigation with the exact question the learner came to answer, framed by a media narrative that keeps calling modern AI 'unsupervised.'

    • Headlines and research papers often describe today's AI as 'unsupervised learning.'
    • But the training pipelines for frontier models visibly involve curated data, human raters, and written rules.
    • We will test whether the label matches the reality.
  2. 02Commit to an Initial Verdictquiz
    Prediction

    Before any evidence is shown, the learner locks in a single prediction: is modern AI truly unsupervised, partly supervised, or a misnomer entirely.

    • One committed answer, not a survey of options.
    • The answer will be revisited in the resolution scene.
  3. 03What 'Unsupervised' Is Supposed to Meanslide
    Evidence

    State the textbook definition of unsupervised learning — raw input, no labels, the model discovers structure on its own — so the learner has a yardstick to measure reality against.

    • Unsupervised learning works on unlabeled data.
    • The objective is internal: clustering, density, or self-reconstruction.
    • No human answers are provided during training.
  4. 04What Actually Builds a Modern Modelslide
    Evidence

    Lay out the real pipeline: web-scale scraping, filtering, instruction datasets, human demonstrations, preference rankings, and safety rules. Make the human involvement visible.

    • Pretraining on raw text is the only stage that looks unsupervised.
    • Instruction tuning uses millions of human-written examples with labels.
    • RLHF and constitutional methods add human or rule-based feedback signals.
    • Safety filters and system prompts hard-code additional constraints after training.
  5. 05Sort the Pipelineinteractive
    Evidence

    Drag-and-drop each training stage into the correct column: Purely Unsupervised, Supervised, or Feedback-Driven. The widget reveals how few stages are actually unsupervised.

    • Pretraining on raw text is the closest fit to unsupervised.
    • Instruction tuning is supervised on labeled prompt-response pairs.
    • Preference training uses ranked human judgments as a learning signal.
    • Safety rules are constraints, not a learning stage — but they shape behavior.
  6. 06Why the Label Stuck Anywayslide
    Explanation

    Explain the linguistic and historical reasons 'unsupervised' became a marketing-friendly shorthand for self-supervised pretraining, even when the rest of the pipeline is anything but.

    • 'Self-supervised' is the technically correct term for the pretraining stage.
    • Pretraining is the largest and most visible stage, so it dominated the narrative.
    • The later human-shaped stages are often described as 'alignment,' hiding them from the headline.
    • The result: a partly accurate label turned into a category error.
  7. 07Apply the Lens to a New Systeminteractive
    Transfer

    Given a short description of a different AI product, the learner tags each training signal as unsupervised, supervised, or feedback-driven — testing whether the distinction holds beyond language models.

    • Image classifiers trained on labeled photo datasets are supervised, not unsupervised.
    • Recommendation systems trained on click logs are closer to self-supervised on behavior data.
    • Robotics systems trained from human teleoperation are supervised at the demonstration level.
    • The 'unsupervised' label only fits when no human signal is used at all.
  8. 08Where the Line Gets Blurryslide
    Boundary

    Show the edge cases where the categories start to leak — self-supervised objectives, weakly supervised web data, and synthetic labels generated by other models — so the learner doesn't overgeneralize the rule.

    • Self-supervised learning is technically supervised by the data itself, but uses no human labels.
    • Curating web data implicitly encodes human judgments about what counts as 'clean.'
    • Synthetic labels from other models are labels, just not human ones — they still count as supervision.
    • The cleanest definition: who or what provided the learning signal.
  9. 09The Verdictslide
    Resolution

    Return to the driving question and resolve it directly, restating where the initial intuition was right, where it was wrong, and the precise wording that replaces the misleading shorthand.

    • Modern AI is not unsupervised in the textbook sense.
    • It is a pipeline: one large self-supervised stage plus several smaller human-shaped stages.
    • 'Trained mostly on raw text with human feedback at the end' is more accurate than 'unsupervised.'
    • Tighter vocabulary leads to sharper questions about data, labor, and control.
Discussion

Discussion threads for a Stage aren't available yet.

Where this leads
Explore more

More in Technology & Computing

See all