Back to recaps

Grant Sanderson on the Locked-Box Thought Experiment for AI Writing

  • Agents
  • AI For Science
  • Frontier Models And Capabilities
  • Jobs, GDP, And Economic Growth

Watch the recap

Autoregression is actually a really weird way to produce stuff, if you think about it.

A thought experiment:

Imagine you have to write an essay under the following constraints.

You're locked in a box. Into that box, someone passes you a slip of paper with the essay so far. You're asked to provide the single next word. After providing that word, your memory is wiped.

You're then passed a new slip of paper with the current essay, including your most recently provided word. This repeats for thousands of turns until the essay is complete.

You cannot keep a separate plan, choose the points you wish to make, or circle back to the beginning. Just one blind word choice after the next, based on the existing string of words.

The essay may read somewhat coherently and be grammatically correct. But it would nonetheless be a very shitty essay, lacking insight, surprise, and interesting nonlinear connections.

Now imagine everyone judging your ability to write based on this final essay.

In his conversation with Dwarkesh Patel, Grant Sanderson of 3Blue1Brown uses this thought experiment to describe autoregressive generation. His point: it's no surprise that large language models (LLMs), under the current training paradigm, are bad at writing.

Ideas

  • Dwarkesh Patel and Dwarkesh PodcastGrant Sanderson

    Explanation Is Not Necessarily the Human Remainder

    Idea

    AI may become exceptionally good at explaining and distilling mathematical ideas, so explanation is not necessarily the human role left after theorem proving.

    Insight and explanation
  • Dwarkesh Patel and Dwarkesh PodcastGrant Sanderson

    Social Curation Shapes What Matters

    Idea

    Even if AI can solve and explain mathematical ideas, people rely on trusted human curators to decide what is worth attention, pursuit, and trust.

    Curation and attention
  • Dwarkesh Patel and Dwarkesh PodcastGrant Sanderson

    Embodied Understanding Is More Than Pattern Recognition

    Idea

    People may partly understand another person's emotion through embodied facial feedback, underscoring the complexity of human understanding.

    Botox and embodied understanding
  • Dwarkesh Patel and Dwarkesh PodcastGrant Sanderson

    Writing Requires Reader Modeling, Not Just Distillation

    Idea

    Writing is not merely clear distillation; it involves original insight and navigating a reader's evolving mental state.

    Writing as way-finding
  • Dwarkesh Patel and Dwarkesh PodcastGrant Sanderson

    The Locked-Box Problem of Local Prediction

    Idea

    Autoregressive generation may privilege locally probable continuations while valuable cross-field connections can be unlikely.

    Locked-box question
  • Dwarkesh Patel and Dwarkesh PodcastGrant Sanderson

    Better Environments May Reward Rare Connections

    Idea

    Rare and valuable connections may become trainable through better data, environments, and incentives without changing the basic generation paradigm.

    Environment and incentives counterpoint
  • Dwarkesh Patel and Dwarkesh PodcastGrant Sanderson

    Digital Research Scales Through Parallel, Diverse Contexts

    Idea

    AI's advantage may come from many parallel agents with fresh contexts, deliberate disagreement, and systematically different heuristics.

    Parallel diverse agents

Tags

  • Frontier Models
  • Agent Orchestration
  • Research Labor Productivity
  • Labor Automation