Back to Discover
Lesson

Hidden Dimensions in Data

How unsupervised learning reveals latent structure — clusters, low-dimensional manifolds, and anomalies — hidden inside high-dimensional data.

Before you enter

A complete interactive classroom, not just a preview.

Start when you are ready to enter this Stage's 13 scenes and explore, respond, and learn as you go.

13
Scenes
26 min
Estimated
Content language: en-US
Start this Stage
Sign-in may be required to play
What happens inside
  1. 01A Question Hidden in Plain Sightslide
    Orientation

    Frame the opening question: can a machine see patterns we cannot? Set up why high-dimensional data is opaque to humans and why machines have an edge.

    • Real data often has hundreds of features
    • Humans see in 2D/3D — machines do not
    • Unsupervised learning explores without labels
  2. 02The Curse of Dimensionalityslide
    Model building

    Explain why high-dimensional space behaves counter-intuitively: distances concentrate, neighborhoods become empty, and intuition breaks.

    • Volume explodes as dimensions grow
    • All points become roughly equidistant
    • Nearest-neighbor loses meaning
  3. 03Predict the Hidden Shapeinteractive
    PredictionPredict

    Learners look at a scatter of points that look random and predict how many underlying variables generated them before revealing the answer.

    • Observe the point cloud
    • Predict the number of latent variables
    • Compare guess with the true dimensionality
  4. 04PCA: Finding the Lying Axesslide
    Model building

    Introduce Principal Component Analysis as the linear way to find the axes along which data varies most, projecting into a viewable space.

    • Eigenvectors of the covariance matrix
    • Linear projection preserving global variance
    • Loses local and nonlinear structure
  5. 05PCA vs t-SNE: Same Data, Different Storyinteractive
    Misconception repairObserve

    Learners toggle between PCA and t-SNE projections of the same dataset and notice how each surfaces different structure.

    • Switch projection method
    • Observe global variance view vs local cluster view
    • Note that neither is 'wrong'
  6. 06Autoencoders: Learning the Curveslide
    Model building

    Show how a neural network compresses data to a small latent code and reconstructs it, learning the curved manifold automatically.

    • Encoder → bottleneck → decoder
    • Bottleneck forces the manifold
    • Nonlinear version of PCA
  7. 07Which Tool Wins?quiz
    Assessment

    Three short questions: pick the right reduction method for a stated goal, predict when an autoencoder will beat PCA, and explain why t-SNE distorts global distances.

    • Match tool to goal
    • Reason about nonlinear structure
    • Interpret t-SNE's local-only guarantee
  8. 08Clustering: Naming the Islandsslide
    Model building

    Survey three clustering families: k-means (centroids), hierarchical (nested merges), DBSCAN (dense regions).

    • k-means: assume round blobs
    • Hierarchical: produce a dendrogram
    • DBSCAN: find dense islands
  9. 09Cluster Thisinteractive
    PracticeChoose

    Learners choose an algorithm and parameters on a 2D toy dataset and watch the resulting clusters form.

    • Switch algorithm
    • Adjust parameters
    • See how shape assumptions affect results
  10. 10Are Clusters Real?slide
    Misconception repair

    Address the trap of treating clusters as truth: they reflect patterns in the chosen features, not objective categories.

    • Clusters are patterns, not facts
    • Different features give different clusters
    • Always sanity-check with domain knowledge
  11. 11Anomalies: The Points That Don't Fitslide
    Model building

    Introduce anomaly detection: statistical (z-score), density-based (LOF), and isolation (Isolation Forest). Explain why simple distance from the mean fails in high dimensions.

    • Z-score breaks in high dimensions
    • LOF compares local density
    • Isolation Forest uses random cuts
  12. 12Spot the Odd One Outinteractive
    ApplicationApply

    Learners try to flag anomalies on a shaped dataset, then compare their picks with an Isolation Forest result and see where simple distance fails.

    • Click suspicious points
    • Reveal the model's picks
    • Reflect on where distance-based intuition breaks
  13. 13What Machines Cannot Seeslide
    Synthesis

    Pull together the lesson: unsupervised learning reveals structure but cannot tell us whether that structure is meaningful, fair, or safe.

    • Hidden structure ≠ ground truth
    • Choice of method shapes the answer
    • Ethics of automated pattern discovery
Discussion

Discussion threads for a Stage aren't available yet.

Where this leads
Explore more

More in Technology & Computing

See all