Back to recaps

Maybe intelligence ain’t all that — and why Steve Hou disagrees

Cartoon cover showing a 200 IQ brain waiting for reality
Image: Clifford Sosin

Recap

Intro

In an X article, Clifford Sosin argues that machine intelligence is becoming abundant but real progress still depends on new observations and experiments. Steve Hou's response agrees on the need for contact with reality but argues that agents and laboratory automation will accelerate that evidence loop.

Superintelligence arrives without an explosion

Sosin opens with a blunt claim: superintelligence has arrived and feels incremental. His test is informal. Spend an hour with Fable 5, then compare it with ten very smart friends.

The machine writes excellent code and handles many tasks better than people. It has not produced the discontinuity promised by intelligence-takeoff forecasts. No warp drives. No cascade of discoveries people cannot follow.

“Superintelligence” remains undefined here. METR’s task-horizon work measures rising agent capability on software tasks, not general superiority over humans. Economic diffusion also remains incomplete. U.S. Census data put business AI use around 17–20% from December 2025 through May 2026. Those facts fit an incremental arrival. They do not explain it.

Reasoning fills the gaps between facts

In his definition, intelligence extrapolates between a thin set of known facts. Language models learned those connections from human text. Sosin expects this kind of reasoning to become cheap and abundant.

It works best where answers follow stable rules and can be checked. He puts coding, math, law, and administrative work in that category. A model holding the governing rules and cases can compare candidate answers quickly.

The boundaries are less clean than the article suggests. Software benchmarks support strong performance on structured tasks. Controlled medical studies also show large gains: one randomized trial found much better diagnostic-reasoning scores for AI-trained physicians with LLM access. A separate public-user study found that high standalone model accuracy did not improve users’ diagnoses or disposition choices. Checkable tasks still depend on tools, interfaces, user skill, and the cost of errors.

Complex systems have no reasoning shortcut

Most consequential systems do not behave like a legal code or a software test. Simple local rules can produce storms, turbulent fluids, and group behaviour that cannot be inferred cleanly from a few facts. The system has to run.

Weather supplies Sosin’s main example. Small errors in the measured state of the atmosphere grow over time. ECMWF reports that initial-condition errors and model approximations both weaken forecasts. More computation does not make either one vanish.

His two-week claim is too broad. ECMWF has found probabilistic forecast skill beyond two weeks for some variables and for spatial or temporal averages. Precise local forecasts still decay much faster. Chaos limits deterministic detail while leaving room for useful probabilities.

Sosin locates the constraint at contact with reality. Reasoning connects facts already in hand. Observation supplies facts the model does not have.

Reality sets the experimental clock

A model might design a better jet-turbine blade than an engineering team. The blade still has to be built and stressed until it fails or survives. At the edge of knowledge, a plausible design is not a result.

Physical testing works this way. A NASA-indexed turbine-blade endurance program simulated a full service life to test whether the design target held. Digital models help choose designs and failure modes, while sensor data and endurance tests settle the claim.

Intelligence can shorten this cycle. Autonomous materials labs combine computation, robotics, accumulated data, and active learning to choose and execute experiments. They increase the number and value of reality checks. They do not remove them.

An industrial revolution without an incomprehensible takeoff

Sosin remains optimistic. Anything people can solve at a desk becomes cheaper. That could lift living standards on the scale of mechanized labour.

Early evidence supports material gains. An AI assistant raised output by 14% on average in a study of 5,179 customer-support agents, with larger gains for less experienced workers. Stanford’s 2025 AI Index reports a more than 280-fold drop in the inference cost of GPT-3.5-level performance between late 2022 and late 2024.

The evidence does not establish a literal 200 IQ or an Industrial Revolution-sized outcome. Sosin’s forecast is narrower: cheap reasoning can create enormous value while physical feedback prevents an incomprehensible explosion.

Agents are already connecting intelligence to reality

Hou starts from the same requirement for data, tools, and iteration. He rejects the ceiling Sosin draws from it. Agentic AI is being built to browse, operate computers, call analytical tools, inspect results, and revise its work.

Software agents already act inside controlled environments. The OpenAI Agents SDK lets models inspect files, run commands, edit code, and use tools. Physical systems are narrower but real. The Robin research system linked literature search and data analysis to propose experiments, interpret results, and update hypotheses in biology. AutoLabs generated and checked instructions for a liquid-handling robot.

These systems support Hou’s direction of travel. They also preserve Sosin’s constraint. Current lab agents rely on specialized instruments, structured environments, and frequent human oversight. Intelligence reaches reality through machinery that remains expensive and narrow.

Hou expects that machinery to improve. Better agents can choose experiments. Robots and drones can run them. Each cycle can produce new data for the next decision.

Adoption and cost delay visible takeoff

Hou gives two reasons for the missing economic acceleration. The first is deployment. Models change quickly. Firms must redesign workflows, train people, set evaluation rules, and decide where errors are acceptable.

Broad adoption remains uneven. Census found 19.8% of U.S. employer businesses using AI in the period ending May 3, 2026, compared with 37% of firms with at least 250 employees. Stanford’s 2026 AI Index reports much higher organizational adoption in a different survey population, while agent deployment stayed in the single digits across nearly all business functions.

Inference keeps getting cheaper, but token prices are only part of deployment cost. Integration, governance, latency, supervision, and error recovery remain. Existing gains in customer support and other bounded workflows show billable value before those costs reach zero.

Capability is still too unreliable and scarce

Hou’s second reason is model quality. Frontier systems fail in predictable ways and still need close supervision outside their strongest domains. Most measured long-horizon performance remains concentrated in software.

METR reports longer software task horizons, but defines them at a chosen probability of success and warns that measurements above 16 human-hours are unreliable with its current suite. “Multi-hour task” describes how long a human expert would need for the benchmark task. It does not mean broad, dependable autonomy for the same duration.

Usage is moving beyond engineering. OpenAI reports agent adoption across legal, finance, and recruiting, though its internal usage data do not independently measure work quality. Medicine shows the same split: models can excel in controlled evaluations and still fail when ordinary users must apply their output.

Hou therefore treats today’s economy as an early deployment snapshot. Capability, reliability, cost, and adoption are all moving. He wants to let them run longer before calling intelligence overrated.