How AI Image Generation Works
A clear mental picture of how diffusion-based AI image generators turn noise and text into coherent images, and why prompt design matters.
A complete interactive classroom, not just a preview.
Start when you are ready to enter this Stage's 11 scenes and explore, respond, and learn as you go.
How does a computer turn a text prompt like "a cat in a spacesuit" into an actual image?
- diffusion-basics
- Diffusion models learn by adding noise to images and then learning to reverse that process step by step.
- text-encoder
- A model like CLIP maps a text prompt into a numeric representation (embedding) that the image generator can condition on.
- denoising-loop
- The generator repeatedly denoises a noisy image, each step nudging it closer to something that matches the prompt.
- latent-space
- Modern generators like denoise work in a compressed "latent" representation of images, not on raw pixels, for efficiency.
- training-data
- Generators learn from millions of image–caption pairs scraped from the web, which shapes both their strengths and biases.
- prompt-craft
- The words, order, and style cues in a prompt steer the denoising process and strongly influence the output.
AI image generators search a database of existing images and pick one.
Show that the generator creates a brand-new image from learned patterns, never by retrieving a stored image.
The AI "understands" the prompt the way a person does.
Show the AI has no semantic understanding; it only matches statistical patterns between text embeddings and image features.
More words in a prompt always makes a better image.
Show that prompt clarity, order, and weighting matter more than sheer length; noisy prompts can hurt results.
- advanced math of diffusion equations
- GAN-based generators
- fine-tuning, LoRA, or training your own model
- video generation
- Learner explains in their own words the two main stages of diffusion (noising and denoising).
- Learner describes what the text encoder contributes to image generation.
- Learner rewrites a weak prompt into a stronger one and predicts how the output will change.
- Apply prompt-design intuition to a generator they have not used before.
Curious adults with no AI or coding background. Basic comfort with the idea that computers can be trained on data.
- 01What Does "AI Image Generation" Mean?slideOrientation
Set the big-picture question: how does text turn into a picture? Preview the journey from noise to image.
- From prompt to picture
- Start with static, end with art
- What we will explore
- 02It Does Not Copy PhotosinteractiveMisconception repairObserve
Learners see pairs of generated outputs from the same starting noise but different prompts, building intuition that images are synthesized, not retrieved.
- Watch images emerge from static
- Compare two prompts side-by-side
- Notice: same noise, different pictures
- 03A Picture Is Just NumbersslideModel building
Explain that an image is a grid of pixel values, and "noise" is just random pixel values with no structure.
- Pixels = numbers
- Structure = signal
- No structure = noise
- 04Guess the PromptinteractivePredictionPredict
Learners watch a denoising animation and predict which prompt produced it, sharpening their model of how prompts steer generation.
- Watch structure appear over steps
- Choose the matching prompt
- Explain your reasoning
- 05Map of Images (Latent Space)interactiveModel buildingConstruct
An interactive 2D map where learners place image concepts as points, then watch how prompts pull the generation toward a region.
- Similar images cluster together
- Prompts act like coordinates
- Steering by moving through the map
- 06The Denoising LoopinteractivePracticeApply
Step through the denoising process manually: each click removes one layer of noise. Learners see how structure forms gradually.
- One step = a little less noise
- Structure appears over many steps
- Early steps set the layout
- 07Seeds: Why Two Runs DifferslideModel building
Explain that the random starting noise (the seed) sets the path, so the same prompt produces different images each time.
- Seed = starting noise
- Same seed + same prompt = same image
- Different seed = new variation
- 08More Steps = Better?interactiveMisconception repairChoose
Learners adjust step count and compare quality, discovering that low steps are too noisy while excessive steps can wash out detail.
- Try very few steps
- Try very many steps
- Find the sweet spot
- 09Prompt Tuner GameinteractiveApplicationApply
Action game: the player adjusts prompt sliders (subject, style, mood) to make the generation match a target image as closely as possible within a limited "budget" of changes.
- Tune prompt features
- Match the target
- Score by similarity
- 10Build Your Own PipelineinteractiveSynthesisConstruct
Drag-and-drop diagram: learners arrange the stages (noise → prompt embedding → denoising loop → final image) in the correct order.
- Order the stages
- Spot the feedback loop
- See the full pipeline
- 11Putting It All TogetherslideSynthesis
Recap the full pipeline and connect each step back to a misconception the learner has now repaired.
- Noise → prompt guidance → denoise → image
- We repaired three myths
- You can reason about any new tool
Discussion threads for a Stage aren't available yet.