Slow Fast

Gardeners of Intelligence:
Where the Next Model Grows

Three people stand in a brick archway behind a barred iron gate at night, the one on the right holding a large sack of garden fertiliser against their chest.
The growers, the crop and the fertiliser are all on their side of the gate. The business is on the other side of it.Lock, Stock and Two Smoking Barrels (1998), written and directed by Guy Ritchie.

1“You don’t look like your average horti-fucking-culturalist!”

Winston, Lock, Stock and Two Smoking Barrels, written and directed by Guy Ritchie, 1998.

The line comes from a scene about fertiliser and a money counter, and the film around it separates two accomplishments that the AI industry keeps treating as one: knowing how to grow something valuable, and controlling the business that forms around it. Winston and his friends grow the crop. Rory Breaker, the man who pays them, controls the trade, and by the end of the film the crop and the money have left in someone else’s van.1

For three years the AI industry followed that division. A hospital, a law firm, a bank or a marketplace held material worth learning from, a laboratory held the means of learning from it, and the owner could license the material or leave it unused. That division depended on the means of learning being scarce. Once an organisation with a valuable source can obtain a capable foundation model, adapt it and use models to help with the adaptation, it has a third choice, which is to produce the capability itself, and specialist suppliers are beginning to sell parts of that production to anyone. The general model can travel to the scarce source. A refusal to license then has two readings: the owner may misjudge its asset, or it may have a better use for it, since the source that makes an owner an attractive supplier also makes it a credible competitor.

Whether the next improvement in general models strengthens that owner or makes its asset dispensable depends on three effects, each on a different margin. A better learner can extract more from the source, which raises the source’s contribution. Accessible foundations and development assistance lower the cost of turning the source into a product or an internal capability. And the same progress equips competitors to reach the result without the source, which increases the supply of substitutes. The three can move in opposite directions, the owner’s position depends on their net effect against alternatives that also improve, and none of them can be read off a benchmark score.

The unit of analysis is therefore a learning regime: the activity that produces experience, the observations it makes available, the method that learns from them, and the costs and arrangements that keep it running. A regime can be a curated corpus, an executable task generator, a simulator, a scientific experiment or an instrumented setting for human work.

Two bets follow. The technical bet is that for work involving judgement, observing people who have relevant freedom of action teaches a model something transferable that observing a prescribed workflow does not. Watching people also changes what they do, as the Hawthorne studies at Western Electric made famous in the 1920s, so observation is one of the variables in that bet and not a neutral record of it.2 The strategic bet is that as foundations and development assistance improve, some owners of scarce sources will find it more valuable to produce capability than to license their sources. They are independent of each other: an owner can become a producer with a constrained expert workflow, a measurement source or a verified simulator even if freedom adds nothing to learning, and freedom can improve learning without giving anyone a profitable production option.

2A rigged table and a bad hand leave the same debt

Eddy sits down at Hatchet Harry’s three-card brag with the hundred thousand pounds he and three friends have pooled and stands up owing five hundred thousand, with a week to pay. The game was rigged, and nothing in the debt says so.

A supplier’s plant floods on a Tuesday. By Friday there is an agreement: a new price, a delivery date, a penalty if the date slips. It is signed and filed, and a year later it is in a training set.

Two negotiations can produce that agreement. In one the buyer correctly diagnosed the supplier’s operating constraint and priced it; in the other the buyer made an unnecessary concession and was lucky. The signed agreement is enough to learn the contractual terms and does not distinguish the two policies, so it is of little use for learning to diagnose the next interruption. The distinction lies in what the buyer knew at each point, the proposals made and withdrawn, the revisions, and whether the delivery date was met. An invoice has that property too: it records that something was agreed and nothing about why.

one agreement, two policies: a correct diagnosis or a lucky concession The information available Proposed actions Revisions The signed agreement Subsequent delivery Enough for extracting contractual terms What separates the two policies
Two negotiations produce one signed agreement after the flood, and the record of what the buyer knew, proposed and revised, and whether the date was met, is what separates the two policies.Drawn from the sequence described in this essay.

The problem is consequential ambiguity, a record that leaves open a distinction the target task needs. Text is not limited to outcomes, since corpora contain arguments, revisions, failed experiments and reports of interventions, so the question is which evidence survives in a given record and whether a learner can use it. The economical observation is the one that resolves the needed distinction at the lowest cost. Sometimes that is a detailed chronology, and often it is something cheaper: an expert’s explanation, a targeted question, a natural experiment. The criterion is what becomes learnable, and the volume of recorded activity does not measure it.

An invoice records that something was agreed and nothing about why.

There is a harder limit underneath. Suppose two worlds produce identical historical records but call for different interventions, as the bad hand and the rigged table do. No amount of computation over those records can tell the worlds apart; only discriminating evidence, such as a randomised intervention, or an additional justified assumption can.3 The limit bears differently on training and on deployment. Rich training records can teach a policy or an inference procedure that transfers to new cases, but they cannot tell a model about a hidden state at deployment when the model receives no signal of that state.

3The leaderboard hides the bill

Suppose a correct decision is worth $100 and a wrong one costs $900. A model that costs $1 an attempt and is right 90 per cent of the time has an expected net value of minus $1 a decision, and a model that costs $10 an attempt and is right 99 per cent of the time has an expected value of $80. Under these payoffs the model that is ten times dearer per attempt is the only one worth using. The example proves no general relationship between accuracy and market value, and it shows why the price of an attempt cannot settle the comparison.

The $1 model, 90 per cent accurate The $10 model, 99 per cent accurate +$90 −$90 −$1 −$1 +$99 −$9 −$10 $80 Gain Loss Cost Net Gain Loss Cost Net 0.9 × $100 0.1 × $900 per attempt 0.99 × $100 0.01 × $900 per attempt
The $10 model costs ten times as much per attempt and is the only one of the two that makes money.Drawn from the arithmetic in this essay.

The price hides at least four separate quantities: how well a system performs under specified conditions, what it costs to reach adequate performance when those conditions change, the range of tasks on which it reaches the required performance within a declared budget, and the continuing work of observing, checking and correcting its behaviour. These are adaptedness, adaptability, capability and regulation burden.

A system can improve on one while becoming more expensive on another, and enterprise software has long kept the fourth off its price list, in a repair process around the product. Two models with similar accuracy can need very different amounts of human review once a reliability target is fixed.4 A system is useful when its responses match the disturbances that matter, which gives no reason to maximise how many behaviours it can produce; Ross Ashby stated this in 1956 as the law of requisite variety. A model that can generate a thousand versions of the wrong answer has variety in abundance and a modest claim to usefulness.5

Capability is therefore better described as a frontier than as a score, along eight properties that interact and are not scores to multiply: breadth, the range of tasks served; depth, how hard a task can be done well within a budget; transfer, how far competence carries to unfamiliar cases; adaptability, the effort needed to reach a new behaviour; steerability, how reliably the model does what the objective and its authority require; dependability, meaning robustness, calibration and the severity and detectability of failures; learning efficiency, the full cost of acquiring the competence; and execution efficiency, the cost of using it at the required quality. The evaluation has to match the value claimed: a gain in one professional domain is a narrower result than general intelligence, and it does not have to be general to be worth its cost.

The bill also runs over time. Producing experience means building environments, paying for participants’ time, observing, verifying, curating, running experiments that fail, training and maintaining, and using the result means inference and continuing correction. Volume changes the answer, because an investment that is uneconomic for a small deployment can be cheap per use at scale, and counting the expected cost of serving a model changes the size it should be trained to, as Sardana and colleagues showed in 2024.6

Component benchmarks share the weakness. TypeSafe launched Jev in September as a decision model that answers typed questions with choices and probabilities instead of text, and it reports that on selected workflows Jev is twenty to two hundred times faster and forty to four hundred times cheaper than a language model, ratios drawn from company-defined evaluations. In an independent edge-network experiment, Jev’s faster decisions produced 459 correct, on-time completions out of 1,080, against DeepSeek’s 463.7 The difference is too small to rank the models, and it shows what the customer pays for: completed work under latency, cost and failure constraints, to which a component’s speed is one input.

Three accounts have to be kept apart. The customer’s benefit, the producer’s revenue and the wider social effect can move differently: a user’s saved hour is not revenue to whoever trained the model, and a payment to a contributor is the buyer’s cost and the contributor’s income, while the contributor’s time is a real cost either way. David Teece showed in 1986 that inventing something and keeping the returns from it are different problems, decided by different assets.8 What a unit of valued behaviour costs to produce, with every cost counted, and who keeps the difference are properties of the regime that produced the model, and no leaderboard reports either.

4Progress does not have to stop for scarcity to move

Large-scale language modelling made broadly useful behaviour learnable without programming each behaviour. The transformer made sequence training tractable at scale, scaling studies made the relationship between model size, data, compute and loss measurable within their tested ranges, Chinchilla showed that splitting a compute budget well between parameters and data mattered more than parameter count, and instruction and preference training changed what a model does without making it bigger.9 Design moved into architectures, objectives, data preparation and the allocation of compute, where it is harder to see. None of this depends on scaling being finished. The advantage from a particular improvement can decline while the technology keeps advancing, because rivals can obtain it, and a published method can remain costly to implement, so similar scores do not make producers interchangeable. The useful question is narrower: which investment, given the alternatives available, produces the required behaviour at the lowest full cost.

There are three routes to better learning, and a new source has to beat the other two. New evidence can distinguish possibilities that existing records leave unresolved. A different representation can make an existing relationship easier for a bounded learner to acquire. Selection, weighting, sequencing or additional computation can extract more from information already held. The routes overlap, and only the first requires a new source.

The third route is the main competitor to any expensive new source, and it has grown stronger. In 2023 DoReMi’s re-weighting of an existing corpus reached baseline accuracy in about 2.6 times fewer steps. In May, OP-Mix adapted the mixture to the model as training progressed and, in one three-stage comparison at 7 billion parameters, matched the tuned recipe’s accuracy, 60.0 against 60.2 per cent, with 95 per cent fewer training FLOPs, the cost of selection included. Other work this year found more learning in archives already held by allocating repetition better or re-presenting the material faithfully.10 These methods help owners and rivals alike, and their gains sit on different denominators (search spend, loss, data efficiency, task scores) that cannot be added into one speed-up. They also fail in instructive places: CausalMix’s gains on development tasks did not become superiority on unseen tasks, so a newer method can still lose the comparison that matters.11

The second route has evidence too. Generated rationales have made behaviour easier to learn without adding a fact about the world, although a useful explanation need not be a faithful record of how an answer was produced.12 The location of a gain matters as much as its size. The prominent result on step-by-step verification came from a reward model selecting among generated solutions, which is a result about selection and not about the unaided generator, and DeepSeek’s R1 improved its policy with longer reasoning, which puts inference spend inside the comparison.13 Novelty, secrecy and apparent richness describe an input. A source earns strategic attention through its marginal contribution against an improving alternative, and that contribution has to be measured.

When a source does contribute, the contribution takes one or more of four technical forms. It can supply missing coverage, meaning tasks, states or exceptions absent from the learner’s effective experience; missing relationships, meaning dependencies that the record obscures or the representation makes hard to acquire; missing evaluation, the ability to tell a successful result from a persuasive one; or missing adaptation opportunities, practice at responding when goals, rules or conditions change. Each has to be shown against the cheaper routes, and a claim about any source needs four links: a distinction a model can learn, an effect on the model, an outcome someone values, and a reason an equivalent result is not cheaper elsewhere.

5Asking is cheap, and the permission to ask is scarce

Manufactured experience has a history. AlphaGo Zero learned Go without human games from the rules, a win signal and self-play; DeepMind’s AdA adapted quickly within an authored task space using a curriculum and memory; Absolute Zero had pretrained models propose tasks for themselves and learn from code execution without an external post-training dataset.14 In September a small model went further, starting from random initialisation, generating executable byte programs as its curriculum and transferring to prediction on natural data, although natural validation data still guided the choice of model and settings.15 The result is far from general intelligence and cannot invent private facts, and it weakens the claim that useful general structure can only be learned from human records.

Building the environments is becoming a learned task as well. SPADE trains a designer that uses documents, execution and memory to keep challenges useful, and RACES found that composing verifiable environments raised a 14-billion-parameter model’s mean score from 48.8 to 51.3 compared with training on the components, with instances and training steps matched.16 Results like these weaken any asset claim that rests on hand-authored tasks being scarce. The designers still depend on checks, inputs and selection that someone specified, and those remain costs of production.

The limit lies in what each generator can establish. A code executor establishes consequences under the rules somebody implemented, and a hand-built simulator establishes consequences under the dynamics somebody assumed. A learned simulator estimates the dynamics from evidence, and only an experiment supplies new evidence about the external process; visual realism or long trajectories do not turn a simulator into one. Transfer is a separate achievement from manufacture: detectors trained on randomised synthetic images located real objects, game agents overfitted thousands of generated levels, and models trained on human demonstrations failed on unfamiliar websites.17 A workable division of labour uses external activity to discover missing distinctions, simulation to multiply practice around them and independent outcomes to test what transfers, and it counts the cost of keeping the simulator calibrated.

Let me take it one step further. A stored record answers the questions its production happened to preserve, while an environment allows a new question: change an action, vary a condition, observe the consequence. Choosing informative experiments is an established discipline, and models are making it cheaper.18 The scarce input then shifts from the archive to the ability to arrange an informative encounter with a process that matters, which requires permission to intervene, access to experts, reliable measurement and the patience to observe delayed outcomes. A record can often be copied after it exists, while a new cooperative episode has to happen before anyone can copy it. Proposing the question becomes cheaper with each model generation, and the permission, the access, the measurement and the valid interpretation of the answer often belong to other people. An organisation that contributes no unique text can therefore hold the scarce input, provided it can tell a successful candidate from a plausible one.

6Schrödinger’s workflow

A system that lets participants choose between two supplied explanations cannot record the third explanation they were never allowed to propose. A workflow recorded after its consequential choices were designed out teaches the workflow and not the judgement it replaced. Structure and constraint are different properties: a record can be precisely structured without prescribing each action, and a fully instrumented workflow can still exclude the alternatives one hoped to learn about.

The technical bet is specific. For tasks that need contextual judgement or revision under changing constraints, observing the process should improve transfer more when participants can exercise relevant discretion than when their strategy is prescribed, provided the distinctions they reveal are learnable and the outcomes can be assessed. The prediction concerns an interaction and makes no claim that maximum freedom is best, since more freedom also makes exploration, interpretation and credit assignment harder. It can be tested by crossing two conditions, prescribed or broader choice during the work and terminal or process records for training, and asking whether the benefit of process records is larger under broader choice, measured on tasks outside the environment that produced the records.19 Constraint has evidence on its side as well: HLER reported fewer critical research failures with a constrained architecture around an unchanged model, and because its gates can reject work, fewer failures did not mean more research completed.20

Observation has to keep its evidential categories apart. An action, a state change, a stated intention and an inferred motivation are different records, and explanatory text is useful without being a transparent trace of the process, for models as for people, since how far a model’s stated reasoning drives its answer varies with the model and the task.21 Feedback has limits as well: execution can verify a program against an incomplete test, human preferences carry judgement and can stay divided, and a learned reward can be exploited at its weaknesses.22 An independent evaluator is valuable when it measures the intended consequence, and a second model with a different name is not an independent evaluator.

Structured outputs have a related limit. Jev answers inside a set of options the customer supplies, which prevents invalid answers and still permits the wrong valid one. Five properties therefore need separate tests: a valid output type, a correct meaning, calibrated uncertainty, resistance to adversarial instructions, and an authorised, successful execution. TypeSafe’s documentation lists Jev’s failures on arithmetic, multi-hop questions, long context and prompt injection, and structured outputs do not remove substantive error.23

Recording people also changes them. In the illumination experiments at Western Electric’s Hawthorne plant between 1924 and 1927, output was reported to rise whenever the lighting changed, brighter or dimmer, and Henry Landsberger later read the relay-assembly results as showing that the workers responded to the attention of the study more than to the conditions being varied. The original data are weaker than the story, and the lesson survived in research design.24 Interfaces, payment, observation and the choice of participants all alter behaviour, so a record of people working with their discretion intact is an aim to be tested and no guarantee of representative behaviour.

There is still a bounded reason to care about actual human behaviour, since agents trained to coordinate with copies of themselves coordinate poorly with people.25 The most direct trained-agent evidence supports curated process supervision: training a software agent on curated records of how fixes were reached raised it from 39.6 to 42.4 per cent on SWE-bench Verified in a size-matched comparison, and to 50.4 per cent with the full curated set, a gain that also reflects recovered coverage of hard issues.26 It gives no support to maximal logging or unrestricted activity. What a model can learn from a record is bounded by the choices the people in it were allowed to make and by how well the record captures their consequences.

7The terms of access change what there is to learn

Offer a purchasing team a fee to record how it handles the next supply interruption, and the team changes. Some of the best negotiators decline, those who agree may negotiate differently while observed, and the resulting record reflects the offer as well as the negotiation. The terms of access therefore enter the production of information through a chain: the arrangement changes participation and conduct, participation and conduct change what can be observed, training extracts a difference, and the difference changes an outcome someone values. Each link can fail. More contributors need not produce better learning, and a reluctant population need not hold anything a willing substitute cannot supply.

A set of supplied options prevents invalid answers and still permits the wrong valid one.

Refusal can be rational, because information is non-rival in use and competitive in consequence: a record can be learned from without being used up, while the competence learned from it reduces its producer’s advantage. Jones and Tonetti formalised that tension in 2020.27 Both kinds of offer have costs in conduct. A generous arrangement recruits expertise and also attracts behaviour optimised for the payment, and a restrictive one preserves trust and can prevent the learning that would have paid for it.

The source owner also has to be identified. An employee, an employer, a customer, an expert and a rights holder can each control part of an activity, so the owner is whoever has the effective authority and practical capacity to supply a specified use, sometimes as a coalition. Holding the file establishes neither. Privacy protection and competitive protection address different losses: a model can learn a professional skill without revealing any confidential record, so a promise to learn only general patterns does not answer a contributor whose concern is competition, unless the resulting capability complements the contributor’s work, which has to be shown in the setting.

Suppliers set terms from the other side. Jev’s weights are shared across accounts, customers adapt it through the state, criteria and code they supply, and TypeSafe’s agreement permits training on customer data only with prior consent while reserving telemetry and voluntary feedback as separate channels.28 The default lets an enterprise supply useful private context while keeping its learning assets under its control, and it means TypeSafe’s durable advantage has to come from its technology, its operating cost, its distribution or a demonstrated renewal process, because no automatic loop converts customer traffic into better weights. Thomson Reuters says it does not use customer information to train its model either.29

Documented acquisitions establish demand for particular inputs and the obstacles to trading them, without measuring what any dataset taught any model. Reddit gave Google negotiated access to its content under a partnership announced in February 2024, and a federal judge in California set out in June 2025 how Anthropic had acquired its books, distinguishing training on them from building a library of pirated copies.30 Treating such records as prices confuses a legal outcome with a learning contribution.

A credible bargain can change what experience comes to exist, through compensation, useful services, limits on downstream use or a share of the return, and its credibility matters after the learning has happened, when the contributor’s bargaining position has changed. Rival laboratories can improve their offers too, so access on better terms is a candidate advantage and not a durable one by default. Withholding is one rational response to exposure and partial sharing is another: in Dubus and Legros’s model, sharing can create value by revealing synergies while strengthening a prospective rival.31 The analysis has to include the complete bargain, with shared development, selective licensing and acquisition, or it presents owners with a choice between surrender and isolation that they do not face.

8The growers can work for themselves

Winston’s crop and money leave in Dog’s van, and the fence who moves the stolen crop arranges to sell it to Rory Breaker, the man it was grown for. The growers had the garden and none of the trade, which is the position source owners held before they could produce capability themselves.

Thomson Reuters has taken the other position. It discussed its Thomson model publicly in July 2026, deployed it for professionals in August and published the technical report later that month. Its model starts from Qwen foundations and adds proprietary content, synthetic material, general replay, expert evaluation and an existing product channel. According to the report, the final large-model training run cost under $450,000 in GPU time and total development about $40 million, so most of the cost of becoming a producer lay in building the capacity to run and improve the process, not in the final run.32 The company also describes more than an archive: lawyers write rubrics and preference judgements, failures are diagnosed and fed back, and the starting foundation has been replaced more than once. The production asset is that organised learning capacity, the ability to state what a useful professional answer requires and to turn that judgement into training and a service, and no single ablation shows which ingredient is indispensable.

Final large-model training run, GPU cost below $450,000 Total development approximately $40 million
Thomson as its developers report it, with the final large-model training run below $450,000 in GPU cost against total development of about $40 million; these are developer estimates and not audited totals.Drawn from the figures the developers report, cited in the notes.

The model improved, and specific regressions bound the result. On the report’s broad aggregate Thomson rose from the Qwen model’s 73.0 to 78.5, against 79.5 for Opus 4.8, and its mean on seven public legal benchmarks rose from 63.2 to 69.6, with six of the seven improving; it was slightly less factual in matched legal research, 0.83 against 0.85, and fell on a general terminal benchmark from 54.0 to 48.3.33 Training also stood in for serving effort. Across four selected task families a single Thomson pass scored 71.9 against 71.2 for the base model making four calls, at about a quarter of the decode cost, so an up-front learning investment replaced expenditure that would otherwise recur at every use. Whether that repays the development cost depends on volume, refresh costs, margins and the alternatives available during the payback period.

The institution competes as a complete system. In a blind study of 3,035 tasks judged by 35 attorney-editors, the Thomson system was preferred in aggregate over each of five OpenAI and Anthropic systems, winning 54 to 64 per cent of legal queries against losses of 26 to 33 per cent and rating higher on completeness and usefulness against all five, with legal-specific tools and Reuters news search against their web search; Harvey and Legora were not tested. Holding the model fixed and removing the legal tools and instructions lowered legal preference, with the full system winning 57 per cent and losing 24 per cent of those comparisons, so the controlled contribution is the bundle and not an isolated archive.34 A rival has to replace the useful result of that combination, and matching its foundation alone is an incomplete substitute test.

Distribution supplies the commercial link. CoCounsel had reached one million professionals in 107 countries and territories by February 2026, before Thomson launched, and the model powers Tabular Analysis inside CoCounsel Legal.35 Distribution matters twice: it is a route for selling differentiated quality, and it supplies the volume over which a lower serving cost can repay development. The company keeps a multi-model strategy, has released a small research model, sells no standalone Thomson model, and in September announced a partnership with OpenAI covering its HighQ platform and a preview of a further CoCounsel experience, so producing, partnering and granting access sit inside one strategy whose strongest feature is the ability to replace the public foundation underneath while keeping the learning and distribution system on top.

The case shows competitive conversion: an information owner has exercised the production option, delivered measured gains within the evaluated scope and connected the capability to an existing commercial channel. Persistence is the open part, since the combination has to sustain customer preference, margins and further improvement through the payback period, and customer reach alone is no automatic training loop, particularly when customer data are excluded from training.

TypeSafe occupies a different position in the production economy the growers are entering. Jev is a component: the customer supplies state and criteria, Jev returns choices, scores or probabilities, the customer’s software executes, and a general model can supply the reasoning Jev does not attempt.36 Thomson brings intelligence to a source its owner holds, while TypeSafe sells a tool with which many owners can exploit theirs. A successful supplier is therefore not an example of a source owner choosing to produce, and passing most of a checklist does not establish a durable advantage for either company. William Stanley Jevons observed in 1865 that making a resource cheaper to use tends to increase how much of it is used, and better general models may increase the use of both Thomson and Jev while reducing the scarcity of what each sells.37 The comparison has to be between the contributions each company controls, whatever labels are attached to them.

General model progress assists both routes and their substitutes Thomson Reuters An information owner produces a professional model and product Sources + expert judgement Domain training + evaluation CoCounsel professional product Professional outcomes TypeSafe / Jev A specialist supplies a decision component to source owners Customer state + criteria Jev shared weights (inference) Customer code + optional general reasoning Outcomes + customer-held feedback Documented roles, not measured effects. Customer requests do not imply Jev weight updates.
Two routes into the same production economy: Thomson Reuters produces a professional model from the sources it holds, and TypeSafe supplies a decision component to many source owners. These are documented roles and not measured effects, and customer requests do not update Jev’s weights.Drawn from the roles the two companies document, cited in the notes.

Both positions can gain from better general models, and they retain returns differently. For a source owner, accessible production capability shortens the distance from possessing a valuable input to owning the differentiated service, and distribution then supports both repeated sales and the recovery of fixed learning costs, while a component supplier has to show why its contribution stays preferable as alternative components improve. The line between them does not follow company labels either. In August, Harvey’s Tenet research preview described post-training a Kimi K3 foundation with Fireworks on synthetic, public legal and expert data, with an agenda that includes helping law firms develop and own specialised models.38 Organisations outside the foundation laboratories are acquiring the means of model production, and the open question is which combination of evidence, expertise, technology and distribution keeps the strongest economics.

Any owner that can obtain general capability and combine it with a source that is hard to substitute has the production option, and exercising it requires adaptation expertise, evaluation, deployment, financing and a route to revenue, the complementary assets in Teece’s account.39 The owner need not pretrain a frontier model or sell general models. It can specialise a foundation, replace purchased capability, sell a domain product or keep a model for internal use, and each is a different business with different costs.

The decision rests on three measurements that should be kept apart: the source’s additional learned performance against the strongest alternative without it, which shows whether general progress makes the source more useful or replaces it; the owner’s net value from keeping the rights, which reflects its ability to exploit the asset; and a comparison of that value with the best available arrangement that grants the rights, which decides ownership.

Crossing the first measurement with the third gives four cases, and two of them correct common errors. An owner can want control for a transitional product opportunity while its learning advantage erodes, so a preference for ownership proves no unique technical value; a valuable source can rationally remain available to other producers, so licensing proves no lack of value. The prediction is a threshold. As foundations and development assistance improve, owners with substantial residual learning value, a development bottleneck that assistance can reduce and a practical route to deployment will find that producing under retained control yields more than their best external bargain, and a measurable subset will change strategy. An owner with distribution already in place can cross the threshold sooner, because it has a way to deploy and sell the gain. The option also works before entry, by raising the opportunity cost of granting access and strengthening the owner’s hand in negotiating it. That is a narrower claim than a withdrawal of valuable data from the market.

The price at which an owner grants the rights follows from that comparison. The minimum payment is the owner’s value with the rights retained less its value after granting them, and it rises only when the first rises faster than the second. A buyer can still meet it, a narrower grant can preserve the uses that matter, and a partnership can enlarge the total surplus.40 The assets that contribute most to future learning also create the strongest internal production opportunities, which is why the learning contribution and the outside option have to be measured independently.

The threshold moves because models supply parts of model production. OpenAI’s September account reports 3.1 agent-workdays per researcher workday, with more than half of the successful four-to-eight-hour tasks still needing human intervention; Anthropic’s index has AI leading 26 per cent of weighted research and development work, which is different from doing that work autonomously; and AlphaEvolve’s 23 per cent kernel speed-up reduced Gemini training time by 1 per cent, because the denominator is the entire training run.41 METR has begun measuring the experiment budget needed to find a validated training improvement, which points to the right unit: validated progress per complete expenditure, failed trials and human intervention included. A randomised study of experienced developers using early-2025 tools found them 19 per cent slower, a counterexample to universal productivity gains and not a verdict on current tools.42

For a clean comparison, the starting foundation, the development assistance and the original source should be varied separately. Holding the foundation and the source fixed while changing the assistance identifies a different channel from improving the foundation, and assistance that supplies new supervision is an additional information input as well as saved labour.43 The laboratories receive assistance earlier and run it on better infrastructure, and assistance also makes substitute data, simulators and verification cheaper, so the prediction concerns relative full costs and does not imply decentralisation. The growers can work for themselves where the assistance they can buy closes their development gap faster than alternatives erase their source’s contribution, and those conditions should be measured before entry is observed, since successful firms cannot supply, after the fact, the definition of the conditions that predicted their success.

9The robbers want the van, not the garden

Dog’s gang robs the growers with the help of Plank, one of their customers, and Eddy and his friends take the haul from Dog’s gang when it returns next door. None of them needs to know how to grow anything.

A rival needs behaviour good enough for the customer, and several routes reach it without the owner’s source: public data, a different architecture, a specialist model, verified synthetic practice, more inference, and permitted distillation from the owner’s outputs.44 MirrorCode rebuilt the behaviour of target programs from an executable reference and its tests, without access to the source code during the task. Success was incomplete and the running reference was itself a valuable input, but the study shows that control of original code does not always control the cheapest route to adequate behaviour.45

Tom bought the two antique shotguns for the robbery from a fence, neither man knowing what they were, and they were worth hundreds of thousands of pounds to a collector and nothing extra to a robber. A proprietary archive can be in that position, irreplaceable as a historical object and replaceable as a means of producing behaviour. Economic equivalence is narrower than general intelligence: a competitor serving one valuable task family does not need the original model’s breadth, while matching a benchmark average can miss the rare failures that decide professional value. Target outcomes, acceptable failure severity, latency and assistance budget have to be stated before equivalence can be judged.

None of this depends on incumbent laboratories being unable to adopt new architectures, or on challengers replacing language modelling. Methods, architectures and sources are different routes to a gain, and they differ in how far the gain travels: a large technical improvement can pass readily to rivals, while a modest one can last when its production is hard to reproduce. The test is whether the source keeps its value against the strongest learner and substitute available to either side.

Every comparison here is dynamic, because the alternative advances alongside the asset. The September releases from OpenAI and Anthropic reported further gains in agent performance and cost per task, so the substitute set has to include the current generation.46

Assets fall into two classes as general capability improves. Some fill a capability gap that general progress can close, such as a specialised corpus, an annotation procedure or a workflow teaching a skill the next foundation will supply adequately. Absolute performance using such an asset can rise while its incremental contribution falls, though the investment can still pay if the return arrives within its useful life. Others control a structurally better way to produce or validate the next useful increment: measurements or interventions that existing evidence cannot supply, a lower full cost of generating and verifying experience, or participation and complementary assets that rivals cannot reproduce on equal terms. For those, better general models can raise the usefulness of the source and lower the cost of exploiting it at once. An asset can move from one class to the other as its target is solved or its production improves, and persistence requires the system to keep offering a better attainable result or cost, without requiring imitation to be impossible.

Which constraint binds changes the priorities. Where broad transferable competence is the binding constraint, differences in general capability dominate, and the leading research priorities are relevant coverage and valid feedback, followed by economical selection, learner-appropriate difficulty and useful process structure. Where domain evidence, dependable execution or cost binds, specialist capability can carry the decisive value while the general frontier keeps advancing, through dependable completion, specialist reasoning with evidence, steerability, consistent action over time, adaptation to changed conditions and efficiency at the required quality. These are different constraints and not historical stages, so a specialist opportunity need not wait for general capability to converge. The requirement then depends on the workload: formal mathematics needs correct reasoning and economical search, empirical science needs external validity, professional analysis needs specialist reasoning with evidence, and operational agents need correct completion under constraints.

Private return also has a time limit. A regime can create substantial value and still be a poor investment if the advantage is copied before its cost is recovered. A temporary edge can still pay if its returns arrive before substitution; a durable advantage keeps a useful performance or cost difference through the payback period; and a reinforcing moat needs one more mechanism, in which the owner keeps enough of the benefit to improve the next round of production and that improvement stays ahead as rivals advance. Four measurements separate them: payback time, time until a competitor reaches equivalence at a stated quality and cost, validated learning yield per full expenditure, and the share of created value the owner retains.

A corpus can stay useful for years, a live environment can produce increasingly redundant episodes, and a loop compounds only when earlier learning makes later learning cheaper to produce or extract.47 AgentCL illustrates the difference, since external memory improved reuse within its task stream without improving the held-out coding score, so stored activity, useful local memory and transferable competence are separate assets.48 Concentration in general foundations and durable specialisation among source owners can coexist, and a winner-take-most market is a stronger claim that needs evidence about substitution, access, bargaining and reinforcement, which a learning curve or a feedback loop does not supply.

The value of an archive and the value of access can therefore diverge. As models improve at exploiting existing evidence, some archives lose value while the capacity to obtain the next discriminating observation gains value, provided the target keeps changing and the yield repays the cost of observation. On a stable task, the robbers take everything of value in the van. On a moving one, the value stays with the capacity to grow the next crop.

10Tom is still on the bridge

The film ends with Tom hanging over the rail of a bridge, the two shotguns on the ledge beyond it and his phone in his mouth, while his friends ring to tell him what the guns are worth; the frame freezes before he answers. Both bets are in a similar position, with their outcomes open and the ways they could lose stated in advance.

The technical bet fails if process records help but broader choice adds nothing economically meaningful under a precise test, if the gain disappears outside the environment that produced it, or if synthetic practice and targeted questions win at full cost. The strategic bet fails if stronger foundations eliminate the source’s incremental contribution or make an adequate substitute cheaper over the claimed horizon, and it also fails if the value of retained control rises while the best available bargain rises faster. It narrows to a bargaining result if rights become more valuable while supply remains available at better prices, and the claim of a durable private advantage fails, while the learning result survives, if cheap imitation arrives before payback. These are distinct losses, and the general point that learning has an economic organisation cannot be used to claim every outcome as success.

Five stronger claims would be convenient, and none of them survives the evidence: that better models necessarily strengthen owners, that every valuable AI supplier owns a learning source, that observed entry proves superior retained returns, that structured outputs eliminate substantive error, and that synthetic methods remove all information scarcity. The three effects hold as a conditional account: current research strengthens both the pressure from substitutes and the range of production inputs available to owners, without establishing either concentration or decentralisation.

Each bet implies its test. The technical bet needs the factorial comparison of freedom and observation, with transfer measured on independently authored tasks and deployment inputs held fixed. Answering the investment question requires each alternative, from curation and retrospective annotation to targeted queries, outcome-only learning, verified synthetic experience, calibrated simulation and additional compute, to compete at matched total cost, and the strategic bet needs adaptation repeated across successively stronger foundations, with the source’s contribution, the owner’s cost and the best offer estimated separately, and a competent rival team given a realistic budget to reach the result another way.49 Payback and catch-up belong on one horizon, and repeating the comparison on more than one capable foundation separates a contribution that survives a change of learner from one tied to a particular architecture. Controlled experiments are one source of evidence among three: an independently evaluated working system settles an existence claim that argument alone leaves open, and the market shows who enters, which rights they keep, how offers change and how long catching up takes.

A remaining valuable asset can therefore be described before the experiments are run. Technically, it makes a consequential contribution against improving alternatives, offers a usable route from source to capability, supports valid discrimination between better and worse results, and holds up across the range of cases claimed for it. Economically, it connects to valued outcomes, stays favourable on a complete account as production methods improve, has a feasible path to deployment, and beats the relevant substitutes. Strategically, its controller can capture a return from it, through access, cooperation, deployment or distribution, general progress does not promptly absorb its superiority, it persists through durable evidence or renewed advantage, and, where entry is claimed, producing under retained control beats the best available licence. The twelve conditions apply together to a named asset, workload and horizon, and they are not points to add into a score. Uniqueness, richness, secrecy and human provenance are not among them.

The technical bet waits for an experiment. The strategic bet is already being run: Thomson shows that source ownership, obtainable model capability and a route to customers can be combined into a productive business, and the open question is whether such a combination holds its advantage through successive model generations. As general models improve, they make the means of producing intelligence more accessible, and where valuable learning experience stays harder to substitute, its owners can acquire those means, turn the experience into differentiated capability and keep the resulting business instead of selling its essential input. That is a stronger production option, neither an inevitable withdrawal of data nor a guaranteed moat, and the contest is whether obtainable production capability advances faster than economical substitution for the next valuable learning opportunity. Once a capable model can be brought to the garden, the gardener has another business to consider.

Notes

  1. 1Lock, Stock and Two Smoking Barrels, written and directed by Guy Ritchie (1998). The line is Winston’s, and the exchange it belongs to is about fertiliser and a money counter. The plot as the essay uses it: Eddy loses the £100,000 he and three friends pooled and ends up owing £500,000 in Hatchet Harry’s rigged game of three-card brag, with a week to pay; Winston and his friends grow the crop and Rory Breaker is their paymaster; Dog’s gang robs them with the help of Plank, one of their customers, and Eddy’s four take the haul off Dog’s gang next door; Nick the Greek, the fence, arranges to sell the stolen crop to Rory and has already sold Tom the two antique shotguns that Harry wanted, with neither man aware of what they were; the film ends with Tom hanging over the rail of a bridge, the guns on the ledge and his phone in his mouth. Take it as the distinction the essay rests on: the grower and the business are different accomplishments, and the film is what happens to growers who have only the first.
  2. 2The studies ran at Western Electric’s Hawthorne Works near Chicago from 1924 to 1932: the illumination experiments with the National Research Council from 1924 to 1927, then the Relay Assembly Test Room from 1927 to 1932 and later phases under Elton Mayo’s Harvard team, with the canonical account in F. J. Roethlisberger and W. J. Dickson, Management and the Worker (Harvard University Press, 1939). The label is usually credited to Henry A. Landsberger, Hawthorne Revisited (Cornell University, 1958). Go there for what was run at the plant, and for the book the account comes from.
  3. 3Take a fair binary condition U drawn afresh each episode, with the historical action always equal to U; in one world the outcome is the action, in the other it is U, so both worlds write the same record, and an intervention chosen before U is observed pays three quarters for action one in the first world and one half for action zero in the second. No computation over the record tells the worlds apart; a randomised intervention does. Judea Pearl, Causal Inference in Statistics: An Overview (Statistics Surveys, 2009), for seeing against doing with the assumptions written out. Take it as the counterexample to treating richer logs as identified causality.
  4. 4Veronica Chatrath and colleagues, READY or Not: Reliable Enterprise Agent Deployment (September 2026): under an explicit reliability target, similar autonomous accuracy can imply different amounts of human review, although the review success is assumed and not observed in deployment. Go there for the review arithmetic, and take the assumption along with it.
  5. 5W. Ross Ashby, An Introduction to Cybernetics (1956), chapter 11. Requisite variety concerns effective responses to the disturbances that matter and offers no reason to maximise arbitrary variety; nothing here converts Ashby’s measure into a price. Falk Lieder and Thomas Griffiths, Resource-Rational Analysis (Behavioral and Brain Sciences, 2020), for the bridge between outcome quality and computational spend, which is also not a price. Go there for the law as Ashby stated it.
  6. 6Nikhil Sardana, Jacob Portes, Sasha Doubov and Jonathan Frankle, Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws (ICML 2024). Under their cost assumptions, expected inference demand changes the preferred model size and training duration. Go there for the curves that move as expected inference demand rises.
  7. 7TypeSafe AI, Introducing System One Models & Jev (15 September 2026), for the product and its speed and cost ratios, which come from selected workflows and company-defined evaluation against model-consensus references and are not estimates of typical savings. Delong Li and colleagues, Fast Intent-Driven Service Orchestration with Jev for 6G Edge Networks (September 2026), for the 459 against 463 correct, on-time completions out of 1,080, a small descriptive difference and not a ranking. Go there for the launch claims and the first independent test side by side.
  8. 8David Teece, Profiting from Technological Innovation (Research Policy, 1986). Go there for the complementary assets that decide who profits from an invention, and for the imitators who did.
  9. 9Ashish Vaswani and colleagues, Attention Is All You Need (2017). Jared Kaplan and colleagues, Scaling Laws for Neural Language Models (2020). Jordan Hoffmann and colleagues, Training Compute-Optimal Large Language Models (2022): Chinchilla, at 70 billion parameters, outperformed the 280-billion Gopher on the same training compute. Long Ouyang and colleagues, Training Language Models to Follow Instructions with Human Feedback (2022): raters preferred the 1.3-billion-parameter InstructGPT to the 175-billion GPT-3. Go there for the loss curves, the 70 against the 280, and the rater preferences.
  10. 10Sang Michael Xie and colleagues, DoReMi: Optimizing Data Mixtures Speeds Up Language Model Pretraining (2023), for the 2.6 times fewer steps and 6.5 points on five tasks, with proxy training at 8 per cent of the large-model cost. Hu, Gandhi, Cho, Linzen and Sharma, Always Learning, Always Mixing (OP-Mix, May 2026): 60.0 against 60.2 per cent mean accuracy in one three-stage 7-billion-parameter comparison, with 95 per cent fewer training FLOPs including selection, a scoped alternative to the compared tuning recipe and not a general reduction in training cost. Sedova and colleagues, Scaling Laws for Mixture Pretraining Under Data Constraints (May 2026), and Yu and Xiong, Generating Pretraining Tokens from Organic Data for Data-Bound Scaling (SynPro, May 2026, consulted in its September revision), for more learning from an archive already held. Go there for the step counts and the FLOP savings, and keep each on its own denominator.
  11. 11Tang and colleagues, CausalMix: Data Mixture as Causal Inference for Language Model Training (July 2026): development gains that do not become superiority on the unseen-task aggregate. Take it as the warning that a current method can still fail the transfer comparison.
  12. 12Eric Zelikman, Yuhuai Wu, Jesse Mu and Noah Goodman, STaR: Bootstrapping Reasoning With Reasoning (2022). Cheng-Yu Hsieh and colleagues, Distilling Step-by-Step! (Findings of ACL 2023). A useful explanation is not necessarily a faithful record of the process that produced the answer. Go there for the bootstrapping loop, and for the smaller model that beat larger ones on less data.
  13. 13Hunter Lightman and colleagues, Let’s Verify Step by Step (ICLR 2024, cited in its 2023 first version): the prominent result concerns reward-model selection among generated mathematical solutions. DeepSeek-AI, DeepSeek-R1 (technical report of 22 January 2025, cited in that version): policy improvement and distillation, with longer reasoning that puts inference expenditure inside the comparison. Go there for the selection setup, and for the response lengths that grow through training in the R1 figures.
  14. 14David Silver and colleagues, Mastering the Game of Go without Human Knowledge (Nature, 2017). Jakob Bauer and colleagues, Human-Timescale Adaptation in an Open-Ended Task Space (ICML 2023). Andrew Zhao and colleagues, Absolute Zero: Reinforced Self-play Reasoning with Zero Data (2025): pretrained models generate their own tasks and learn through code-execution feedback with no externally supplied post-training task dataset, which is not learning without prior knowledge, a task language or a checking mechanism. David Silver and Richard Sutton, Welcome to the Era of Experience (2025), for the agenda stated from inside the lab that built Zero. Go there for what each system was given, and for the checker that Absolute Zero could not do without.
  15. 15Cowsik, Dolev, Li, De Luca, Cohen, Goodman and Levine, Self-Play Pretraining with Zero Data (24 September 2026): random initialisation, executable byte-program curricula and transfer to natural-domain prediction, with natural validation data guiding the choice of model and hyperparameters; a very recent small-model result, with no independent replication assumed. Take it as the limit case, and read the validation caveat before the headline.
  16. 16Liu and colleagues, SPADE: Self-Play in Adaptive Synthetic Executable Environments (August 2026). Xiang and colleagues, Verifiable Environments Are LEGO Bricks: Recursive Composition for Reasoning Generalization (RACES, June 2026): the mean rises from 48.8 to 51.3 for the DeepSeek-distilled 14-billion-parameter model and from 60.1 to 61.1 for Qwen3-14B, against individual environments drawn from the same pool with instances and training steps matched. SCOPE (May 2026) extends co-evolving practice to document-grounded, open-ended tasks. Go there for environment construction becoming a learned skill, and for the checks it still needs.
  17. 17Josh Tobin and colleagues, Domain Randomization for Transferring Deep Neural Networks from Simulation to the Real World (IROS 2017). Karl Cobbe and colleagues, Quantifying Generalization in Reinforcement Learning (ICML 2019). Xing Han Lu, Zdeněk Kasner and Siva Reddy, WebLINX: Real-World Website Navigation with Multi-Turn Dialogue (ICML 2024), where models trained on human demonstrations fail on websites they had not seen. Go there for the transfer that worked, the overfitting that persisted, and the websites that broke the demonstrations.
  18. 18Adam Foster, Desi Ivanova, Ilyas Malik and Tom Rainforth, Deep Adaptive Design: Amortizing Sequential Bayesian Experimental Design (ICML 2021). Go there for how to choose the informative experiment efficiently, and take from it that choosing is the part a model makes cheap.
  19. 19Nested views of the same authorised episodes, from terminal artifacts to process records linked to consequences; a factorial crossing prescribed or broader choice with terminal or process records, with the interaction as the endpoint; transfer evaluated on independently authored tasks with deployment inputs frozen. Take it as the design of the comparison this sentence asks for.
  20. 20Chen Zhu, Xiaolu Wang and Weilong Zhang, (Human) Attention Is (Still) All You Need (HLER, June 2026): fewer critical research failures with a constrained architecture around the same model, whose gates can reject unsound outputs. Take it as evidence for deliberate control of consequential actions, which leaves the interaction between human choice and observation to its own test.
  21. 21Tamera Lanham and colleagues, Measuring Faithfulness in Chain-of-Thought Reasoning (2023). Go there for the faithfulness figures by model and task.
  22. 22Paul Christiano and colleagues, Deep Reinforcement Learning from Human Preferences (NeurIPS 2017). Leo Gao, John Schulman and Jacob Hilton, Scaling Laws for Reward Model Overoptimization (ICML 2023). Go there for the preferences that carried judgement, and for the curve where the proxy and the reference part company.
  23. 23TypeSafe AI, Jev 1.13 jaggedness (live page, consulted in its 17 September 2026 revision), for the arithmetic, multi-hop, context and prompt-injection failures. Go there for a vendor recording its model’s failures in public.
  24. 24The legend has since been re-examined. H. M. Parsons, What Happened at Hawthorne? (Science, 1974), attributed the relay-room gains to feedback and pay; Stephen R. G. Jones, Was There a Hawthorne Effect? (American Journal of Sociology, 1992), found little or no evidence of the effect in the relay data; and Steven D. Levitt and John A. List, Was There Really a Hawthorne Effect at the Hawthorne Plant? An Analysis of the Original Illumination Experiments (American Economic Journal: Applied Economics, 2011), recovered the illumination data, found that the dramatic patterns in the usual accounts were not there, and found subtler signs consistent with observation effects. Go there for how a story about lights became a rule of research design, and read the re-analyses before the legend.
  25. 25Micah Carroll and colleagues, On the Utility of Learning about Humans for Human-AI Coordination (NeurIPS 2019). The finding motivates calibration to people; it does not establish broad human trajectories as the best input for general intelligence. Go there for the self-play agents that could not cook with a human.
  26. 26Ma and colleagues, From Patches to Trajectories: Privileged Process Supervision for Software-Engineering Agents (May 2026): 39.6 to 42.4 per cent on SWE-bench Verified in a size-matched comparison for the Qwen2.5-Coder-32B student, and 50.4 per cent with the full curated set, a larger gain that also includes recovered coverage of difficult issues. OpenAI, Improving Instruction Hierarchy in Frontier LLMs (March 2026), where a control capability improved from 0.76 to 0.91 while GPQA Diamond held at 0.83 and some preference measures declined, is the companion result that a control property can move without the reasoning benchmark moving. Go there for the size-matched numbers, and for the trade-offs the second report records against itself.
  27. 27Charles Jones and Christopher Tonetti, Nonrivalry and the Economics of Data (American Economic Review, 2020). Daron Acemoglu, Ali Makhdoumi, Azarakhsh Malekian and Asu Ozdaglar, Too Much Data (American Economic Journal: Microeconomics, 2022), for the opposite distortion, where one person’s disclosure reveals their neighbours. Go there for the hoarding model, and for the sharing one.
  28. 28TypeSafe AI, Master customer agreement and privacy policy (live pages, consulted September 2026), and Jev models for the shared weights. Go there for the consent clause and the reserved channels, and take the terms as snapshots that can change.
  29. 29Thomson Reuters, launch announcement (24 August 2026), How we built Thomson, product page and FAQ (accessed 25 September 2026), Thomson 1.0 Small, and the OpenAI partnership announcement (September 2026). Go there for what the company says it does not do with customer data.
  30. 30Reddit, Expanding Our Partnership with Google (22 February 2024). Bartz v. Anthropic PBC, Order on Fair Use, No. 3:24-cv-05417-WHA (N.D. Cal., 23 June 2025), which granted summary judgment that the training use was fair use while refusing to treat the pirated central-library copies as training copies. Both are historical records: neither is used here to state a rule of copyright law or the present state of the litigation, and neither acquisition nor adoption identifies a dataset’s causal contribution, establishes informed consent or shows net social benefit. Go there for the two documents as filed, and take from them intent and obstacles, not a price per token.
  31. 31Antoine Dubus and Patrick Legros, Make or Buy Decisions and Data Sharing (Journal of Industrial Economics, 2026): in a model of data sharing and mergers, partial sharing can create value by revealing potential synergies even while strengthening a prospective rival. Go there for the case where sharing with the rival was the rational move.
  32. 32Chen and colleagues, Thomson: Continual Learning of Frontier Models for SovereignAI (technical report, 2026): Qwen starting checkpoints; the final large-model run below $450,000 in GPU cost and total development of approximately $40 million, developer estimates and not audited totals. Go there for the two cost figures side by side.
  33. 33Chen and colleagues, Thomson: Continual Learning of Frontier Models for SovereignAI (technical report, 2026), again, for the matched tables: the broad aggregate from 73.0 for Qwen to 78.5, against 79.5 for Opus 4.8 (Table 1); seven public legal benchmarks from a mean of 63.2 to 69.6, six of them improving (Table 13); legal deep-research factuality 0.83 against 0.85 for Qwen and completeness 0.87 against 0.82; Terminal-Bench 2.1 from 54.0 to 48.3; and a single Thomson pass at 71.9 against 71.2 for Qwen with four-call Fusion on four selected families, at roughly a quarter of the decode cost. These are distinct model, system and cost comparisons and not additive effects, and common-harness model comparisons are kept apart from complete-system comparisons with different information access. Go there for the matched tables, and read the two kinds of comparison separately.
  34. 34The complete-system study is section 4.4 of the technical report as published by Thomson Reuters: 3,035 tasks, 35 attorney-editors, win rates of 54 to 64 per cent on legal queries against losses of 26 to 33 per cent and higher completeness and usefulness ratings against all five systems, with Thomson using legal-specific tools and Reuters news search and the other systems web search. The domain-access ablation on legal queries is Figure 32 of the same report, where the full system wins 57 per cent and loses 24 per cent against the same model with its legal tools and instructions removed. Go there for the judged comparison, and for the one place the model is held fixed.
  35. 35Thomson Reuters, One Million Professionals Turn to CoCounsel as Thomson Reuters Scales AI for Regulated Industries (February 2026), and the launch of the next generation of CoCounsel Legal (August 2026). These are CoCounsel adoption figures, not a count of Thomson-model users or an estimate of its revenue. Take them as the channel, and not yet as the returns.
  36. 36TypeSafe AI, Jev models, for the shared weights and the absence of customer fine-tuning: every customer queries one set of weights, and a customer request does not update it. Go there for what the component does and for what it leaves untouched.
  37. 37William Stanley Jevons, The Coal Question (1865), for the paradox. Launch coverage reports that Jev is named after it (Ecosistema Startup, September 2026). Take it as the reason cheaper decisions raise the number of decisions, and the bill with them.
  38. 38Harvey, Tenet research preview (August 2026): post-training a Kimi K3 foundation with Fireworks on synthetic, public legal and human expert data, with an agenda that includes enabling law firms to develop and own specialised models. Take it as evidence that the line between application company and model producer does not follow company labels.
  39. 39David Teece, Profiting from Technological Innovation (Research Policy, 1986), again, and this time for the complementary assets themselves: the manufacturing, distribution and service an innovator needs before it can keep the returns, which here are adaptation expertise, evaluation, deployment, financing and a route to revenue. Go there for the assets, and for the innovators who had the invention and none of them.
  40. 40For a rights grant g and an accessible production package q, the minimum compensating payment is the owner’s best expected net value with the rights retained less its best expected net value after the grant, excluding the payment, and between two packages it changes by the change in the first less the change in the second. A negative reservation payment means the grant pays before compensation; a reservation-price increase, an ownership preference and reduced supply are three distinct empirical outcomes. Take it as the arithmetic that keeps “they will not sell” from being inferred from “it is valuable”.
  41. 41OpenAI, Research acceleration: The view inside OpenAI (6 September 2026). Anthropic, Measurements for understanding the pace of AI development inside frontier labs (17 September 2026). METR, Expenditure Horizon (21 July 2026). Alexander Novikov and colleagues, AlphaEvolve (2025): a 23 per cent kernel speed-up, a 1 per cent reduction in Gemini training time. Corporate measurement reports are not controlled estimates of whole-project acceleration. Go there for the workday ratio and the intervention rate, and for the 1 per cent that is small until you read its denominator.
  42. 42Joel Becker, Nate Rush, Beth Barnes and David Rein, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity (2025): a 19 per cent increase in completion time in its randomised setting. Go there for the 19 per cent, and hold it next to the 1 per cent in the note before.
  43. 43Separate the starting foundation, the development assistance and the original source, and measure complete cost and elapsed time to a declared, independently validated target, unsuccessful attempts included; record any teacher knowledge the assistant introduces, since a fixed archive is not a fixed information budget. Take it as the design that keeps a better assistant from being mistaken for a better foundation.
  44. 44DeepSeek-AI, DeepSeek-R1 (22 January 2025), the distillation section. Go there for the distilled models and their scores, and read them as the reason source exclusivity is no defence.
  45. 45Tom Adamczewski and colleagues, MirrorCode: AI can rebuild entire programs from behavior alone (full paper, July 2026): reconstruction from an executable reference and tests without access to the source during the task; prior exposure in pretraining is not fully ruled out, and success is incomplete. Go there for what the reference supplied, and read it as the reason owning the code is no defence of the behaviour.
  46. 46OpenAI, Introducing GPT-6 Sol and Luna, and Anthropic, Claude Opus 5.5 (September 2026): vendor evaluations of agent performance and cost per task, which differ from token-price reductions and are not matched estimates of any customer’s savings. Take them as the reason the substitute set has to include the current generation.
  47. 47Andrei Hagiu and Julian Wright, Data-Enabled Learning, Network Effects, and Competitive Advantage (RAND Journal of Economics, 2023): the form of learning and of competition decides what advantage results, and the model’s assumptions are not evidence about concentration in AI. The value of a regime over its horizon is its initial cost set against the discounted stream of valued outcomes less renewal and deployment costs, with each cost counted once. Go there for the conditions under which data-enabled learning does and does not lock a market, and for the accounting boundary.
  48. 48Shu and colleagues, AgentCL: Toward Rigorous Evaluation of Continual Learning in Language Agents (June 2026). Go there for the memory that helped inside the stream and not outside it.
  49. 49The rest of the programme: a mechanism study and an investment study run side by side at matched total cost; the bargain varied where randomisation is honest; adaptation repeated across successively stronger foundations; and a competing team given a realistic budget and permission to use public outputs. Yue and colleagues, Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? (2025), for why finite sampling cannot establish a universal boundary of what a base model could ever produce. Take it as the rest of the programme these tests belong to.