September 2026

These are hypotheses, not doctrine. I update them as evidence changes, and previous versions remain available in the site’s history.

Intelligence is scoped task-solving capacity

I use intelligence to mean the capacity of a specified system to solve tasks under specified conditions. It is an operational performance property, not a claim about the mechanism responsible for that performance.

An intelligence claim is incomplete unless it identifies the system, task or task distribution, operating conditions, and evaluation. Evidence of higher performance within that scope does not establish a general scalar ordering across systems.

Intelligence does not by itself imply understanding, grounding, reasoning by a particular mechanism, agency, autonomy, learning, reflection, meaning, selfhood, sentience, or consciousness.

The current AI cycle is a search problem

The durable question is not where a model can be deployed. It is which implementations can perform valuable jobs under real conditions. Current products are competing hypotheses in that search.

The job should be defined one level above the obvious AI task. In software engineering, for example, writing code is a task; safely delivering a useful software change is the job.

LLMs generate proposals; evidence drives selection

LLMs make it cheap to generate plausible candidate explanations, plans, hypotheses, interpretations, and actions. They do not produce an unbiased sample of every possibility, and plausibility is not evidence.

The useful division of labor is:

LLM = proposal generator
tools and environment = evidence generators
humans and objective tests = evaluators

The value lies in the complete loop: generate alternatives, derive their consequences, seek discriminating evidence, reject inconsistent candidates, test the survivors, and update.

Measure the whole job and follow the bottleneck

Making one task dramatically cheaper does not necessarily improve the whole job in proportion. Requirements, context, review, integration, testing, coordination, or judgment may become the new constraint.

An AI intervention should therefore be followed by measurement of the whole outcome, identification of the new bottleneck, redesign, and another cycle of evaluation.

Prefer strong reality signals

The most informative environments have valuable outcomes, digital evidence, results that can be checked, short feedback loops, and mistakes that are detectable and usually reversible. Closed loops allow weak hypotheses to fail quickly and useful ones to improve.

A recurring thesis

Plausible candidates are becoming abundant. Good experiments and discriminating evidence remain scarce.

The durable engineering work is in building systems that turn cheap variation into useful outcomes through evidence, evaluation, feedback, and selection.