The Prisoner's Dilemma: Why Self-Interest Can Backfire
The Prisoner's Dilemma shows that individually optimal choices can produce collectively worse results, and that the structure of repeated play reshapes which strategy wins.
A complete interactive classroom, not just a preview.
Start when you are ready to enter this Stage's 7 scenes and explore, respond, and learn as you go.
If two rational people each pursue their own best interest, why do they often end up with a worse outcome than if they had cooperated?
Two suspects are questioned separately. Each can either stay silent or betray the other. Their sentences depend on a payoff matrix most people get wrong on first try.
Intuition says 'don't snitch' is the safe move, yet game theory predicts the opposite. Something subtle is hiding in the numbers.
A simulated payoff matrix the learner can manipulate: choose Cooperate or Defect against different AI strategies and watch cumulative scores diverge across rounds.
When both players chase their own best outcome, both end up worse than if they had trusted each other — and that paradox is the engine of cooperation, defection, and real-world negotiations.
A learner will probably guess that 'Cooperate' is the safe, decent choice and 'Defect' is the selfish one, expecting cooperation to score higher overall.
- Iterated tournament results beyond a brief mention of Axelrod
- Evolutionary stable strategy formal proofs
- Zero-determinant strategies and their controversies
- Applications to arms races, trade, or climate negotiations beyond passing references
- 01Two Suspects, Separate RoomsslideQuestion
Set the scene with a short narrative framing and the canonical payoff matrix so the learner knows exactly what each player faces.
- Each prisoner can Cooperate (stay silent) or Defect (confess).
- Payoffs depend on what the OTHER prisoner does, not just your own choice.
- The numbers contain a trap that intuition usually misses.
- 02Make Your First MoveinteractivePrediction
Learner commits to one choice against an opponent that will mirror them, setting up a comparison that the later simulation will overturn.
- Choose Cooperate or Defect with no further information about the opponent.
- See the immediate payoff and reflect on the reasoning behind the choice.
- Lock in a prediction before seeing how repeated play changes things.
- 03Run the Repeated MatchslideEvidence
Visualize cumulative scores across 20+ rounds for four strategies: Always Cooperate, Always Defect, Tit-for-Tat, and Random.
- Always Defect wins against Always Cooperate but collapses against Tit-for-Tat.
- Tit-for-Tat starts with one defection and then sustains mutual cooperation.
- Cumulative totals make the structural advantage of reciprocity visible.
- No permanent 'winner' — payoffs shift with opponent's strategy.
- 04Face Off Against the AlgorithmsinteractiveExplanation
Manipulable sandbox: pick a strategy, pick an opponent strategy, run the match, and watch the round-by-round payoff table update.
- Switch opponent between Always Cooperate, Always Defect, Tit-for-Tat, Grim Trigger, and Random.
- Inspect each round's payoff to see why each strategy thrives or collapses.
- Observe how Tit-for-Tat rewards cooperation and punishes defection.
- Build intuition for why 'best move' depends on horizon and opponent.
- 05When Does Cooperation Survive?slideBoundary
Examine the boundary conditions: Tit-for-Tat needs a long horizon, a chance to retaliate, and a way to recover from accidental defection.
- Short or unknown horizons push play toward Defect.
- Noise and miscommunication can lock Tit-for-Tat into mutual defection.
- Generous Tit-for-Tat and Pavlov improve robustness under noise.
- The dilemma softens but rarely disappears in real social systems.
- 06Apply the LogicquizTransfer
One transfer question: identify which real-world situation matches the Prisoner's Dilemma structure and which move risks mutual punishment.
- Recognize the temptation-to-defect payoff pattern in a new setting.
- Commit to a single best answer before seeing the rationale.
- 07Why Self-Interest Can BackfireslideResolution
Resolve the driving question by naming the Nash equilibrium of the one-shot game and explaining why repetition changes the answer.
- In a single round, Defect is dominant — and that dominance produces mutual harm.
- Repeated play lets players reward and punish, shifting the equilibrium toward cooperation.
- The Prisoner's Dilemma reveals the gap between individual rationality and collective outcomes.
Discussion threads for a Stage aren't available yet.