1“You don’t look like your average horti-fucking-culturalist!”
Winston, Lock, Stock and Two Smoking Barrels, written and directed by Guy Ritchie, 1998.
The line comes from a scene about fertiliser and a money counter, and the film around it separates two accomplishments that the AI industry keeps treating as one: knowing how to grow something valuable, and controlling the business that forms around it. Winston and his friends grow the crop. Rory Breaker, the man who pays them, controls the trade, and by the end of the film the crop and the money have left in someone else’s van.1
For three years the AI industry followed that division. A hospital, a law firm, a bank or a marketplace held material worth learning from, a laboratory held the means of learning from it, and the owner could license the material or leave it unused. That division depended on the means of learning being scarce. Once an organisation with a valuable source can obtain a capable foundation model, adapt it and use models to help with the adaptation, it has a third choice, which is to produce the capability itself, and specialist suppliers are beginning to sell parts of that production to anyone. The general model can travel to the scarce source. A refusal to license then has two readings: the owner may misjudge its asset, or it may have a better use for it, since the source that makes an owner an attractive supplier also makes it a credible competitor.
Whether the next improvement in general models strengthens that owner or makes its asset dispensable depends on three effects, each on a different margin. A better learner can extract more from the source, which raises the source’s contribution. Accessible foundations and development assistance lower the cost of turning the source into a product or an internal capability. And the same progress equips competitors to reach the result without the source, which increases the supply of substitutes. The three can move in opposite directions, the owner’s position depends on their net effect against alternatives that also improve, and none of them can be read off a benchmark score.
The unit of analysis is therefore a learning regime: the activity that produces experience, the observations it makes available, the method that learns from them, and the costs and arrangements that keep it running. A regime can be a curated corpus, an executable task generator, a simulator, a scientific experiment or an instrumented setting for human work.
Two bets follow. The technical bet is that for work involving judgement, observing people who have relevant freedom of action teaches a model something transferable that observing a prescribed workflow does not. Watching people also changes what they do, as the Hawthorne studies at Western Electric made famous in the 1920s, so observation is one of the variables in that bet and not a neutral record of it.2 The strategic bet is that as foundations and development assistance improve, some owners of scarce sources will find it more valuable to produce capability than to license their sources. They are independent of each other: an owner can become a producer with a constrained expert workflow, a measurement source or a verified simulator even if freedom adds nothing to learning, and freedom can improve learning without giving anyone a profitable production option.
2A rigged table and a bad hand leave the same debt
Eddy sits down at Hatchet Harry’s three-card brag with the hundred thousand pounds he and three friends have pooled and stands up owing five hundred thousand, with a week to pay. The game was rigged, and nothing in the debt says so.
A supplier’s plant floods on a Tuesday. By Friday there is an agreement: a new price, a delivery date, a penalty if the date slips. It is signed and filed, and a year later it is in a training set.
Two negotiations can produce that agreement. In one the buyer correctly diagnosed the supplier’s operating constraint and priced it; in the other the buyer made an unnecessary concession and was lucky. The signed agreement is enough to learn the contractual terms and does not distinguish the two policies, so it is of little use for learning to diagnose the next interruption. The distinction lies in what the buyer knew at each point, the proposals made and withdrawn, the revisions, and whether the delivery date was met. An invoice has that property too: it records that something was agreed and nothing about why.
The problem is consequential ambiguity, a record that leaves open a distinction the target task needs. Text is not limited to outcomes, since corpora contain arguments, revisions, failed experiments and reports of interventions, so the question is which evidence survives in a given record and whether a learner can use it. The economical observation is the one that resolves the needed distinction at the lowest cost. Sometimes that is a detailed chronology, and often it is something cheaper: an expert’s explanation, a targeted question, a natural experiment. The criterion is what becomes learnable, and the volume of recorded activity does not measure it.
An invoice records that something was agreed and nothing about why.
There is a harder limit underneath. Suppose two worlds produce identical historical records but call for different interventions, as the bad hand and the rigged table do. No amount of computation over those records can tell the worlds apart; only discriminating evidence, such as a randomised intervention, or an additional justified assumption can.3 The limit bears differently on training and on deployment. Rich training records can teach a policy or an inference procedure that transfers to new cases, but they cannot tell a model about a hidden state at deployment when the model receives no signal of that state.
3The leaderboard hides the bill
Suppose a correct decision is worth $100 and a wrong one costs $900. A model that costs $1 an attempt and is right 90 per cent of the time has an expected net value of minus $1 a decision, and a model that costs $10 an attempt and is right 99 per cent of the time has an expected value of $80. Under these payoffs the model that is ten times dearer per attempt is the only one worth using. The example proves no general relationship between accuracy and market value, and it shows why the price of an attempt cannot settle the comparison.
The price hides at least four separate quantities: how well a system performs under specified conditions, what it costs to reach adequate performance when those conditions change, the range of tasks on which it reaches the required performance within a declared budget, and the continuing work of observing, checking and correcting its behaviour. These are adaptedness, adaptability, capability and regulation burden.
A system can improve on one while becoming more expensive on another, and enterprise software has long kept the fourth off its price list, in a repair process around the product. Two models with similar accuracy can need very different amounts of human review once a reliability target is fixed.4 A system is useful when its responses match the disturbances that matter, which gives no reason to maximise how many behaviours it can produce; Ross Ashby stated this in 1956 as the law of requisite variety. A model that can generate a thousand versions of the wrong answer has variety in abundance and a modest claim to usefulness.5
Capability is therefore better described as a frontier than as a score, along eight properties that interact and are not scores to multiply: breadth, the range of tasks served; depth, how hard a task can be done well within a budget; transfer, how far competence carries to unfamiliar cases; adaptability, the effort needed to reach a new behaviour; steerability, how reliably the model does what the objective and its authority require; dependability, meaning robustness, calibration and the severity and detectability of failures; learning efficiency, the full cost of acquiring the competence; and execution efficiency, the cost of using it at the required quality. The evaluation has to match the value claimed: a gain in one professional domain is a narrower result than general intelligence, and it does not have to be general to be worth its cost.
The bill also runs over time. Producing experience means building environments, paying for participants’ time, observing, verifying, curating, running experiments that fail, training and maintaining, and using the result means inference and continuing correction. Volume changes the answer, because an investment that is uneconomic for a small deployment can be cheap per use at scale, and counting the expected cost of serving a model changes the size it should be trained to, as Sardana and colleagues showed in 2024.6
Component benchmarks share the weakness. TypeSafe launched Jev in September as a decision model that answers typed questions with choices and probabilities instead of text, and it reports that on selected workflows Jev is twenty to two hundred times faster and forty to four hundred times cheaper than a language model, ratios drawn from company-defined evaluations. In an independent edge-network experiment, Jev’s faster decisions produced 459 correct, on-time completions out of 1,080, against DeepSeek’s 463.7 The difference is too small to rank the models, and it shows what the customer pays for: completed work under latency, cost and failure constraints, to which a component’s speed is one input.
Three accounts have to be kept apart. The customer’s benefit, the producer’s revenue and the wider social effect can move differently: a user’s saved hour is not revenue to whoever trained the model, and a payment to a contributor is the buyer’s cost and the contributor’s income, while the contributor’s time is a real cost either way. David Teece showed in 1986 that inventing something and keeping the returns from it are different problems, decided by different assets.8 What a unit of valued behaviour costs to produce, with every cost counted, and who keeps the difference are properties of the regime that produced the model, and no leaderboard reports either.
4Progress does not have to stop for scarcity to move
Large-scale language modelling made broadly useful behaviour learnable without programming each behaviour. The transformer made sequence training tractable at scale, scaling studies made the relationship between model size, data, compute and loss measurable within their tested ranges, Chinchilla showed that splitting a compute budget well between parameters and data mattered more than parameter count, and instruction and preference training changed what a model does without making it bigger.9 Design moved into architectures, objectives, data preparation and the allocation of compute, where it is harder to see. None of this depends on scaling being finished. The advantage from a particular improvement can decline while the technology keeps advancing, because rivals can obtain it, and a published method can remain costly to implement, so similar scores do not make producers interchangeable. The useful question is narrower: which investment, given the alternatives available, produces the required behaviour at the lowest full cost.
There are three routes to better learning, and a new source has to beat the other two. New evidence can distinguish possibilities that existing records leave unresolved. A different representation can make an existing relationship easier for a bounded learner to acquire. Selection, weighting, sequencing or additional computation can extract more from information already held. The routes overlap, and only the first requires a new source.
The third route is the main competitor to any expensive new source, and it has grown stronger. In 2023 DoReMi’s re-weighting of an existing corpus reached baseline accuracy in about 2.6 times fewer steps. In May, OP-Mix adapted the mixture to the model as training progressed and, in one three-stage comparison at 7 billion parameters, matched the tuned recipe’s accuracy, 60.0 against 60.2 per cent, with 95 per cent fewer training FLOPs, the cost of selection included. Other work this year found more learning in archives already held by allocating repetition better or re-presenting the material faithfully.10 These methods help owners and rivals alike, and their gains sit on different denominators (search spend, loss, data efficiency, task scores) that cannot be added into one speed-up. They also fail in instructive places: CausalMix’s gains on development tasks did not become superiority on unseen tasks, so a newer method can still lose the comparison that matters.11
The second route has evidence too. Generated rationales have made behaviour easier to learn without adding a fact about the world, although a useful explanation need not be a faithful record of how an answer was produced.12 The location of a gain matters as much as its size. The prominent result on step-by-step verification came from a reward model selecting among generated solutions, which is a result about selection and not about the unaided generator, and DeepSeek’s R1 improved its policy with longer reasoning, which puts inference spend inside the comparison.13 Novelty, secrecy and apparent richness describe an input. A source earns strategic attention through its marginal contribution against an improving alternative, and that contribution has to be measured.
When a source does contribute, the contribution takes one or more of four technical forms. It can supply missing coverage, meaning tasks, states or exceptions absent from the learner’s effective experience; missing relationships, meaning dependencies that the record obscures or the representation makes hard to acquire; missing evaluation, the ability to tell a successful result from a persuasive one; or missing adaptation opportunities, practice at responding when goals, rules or conditions change. Each has to be shown against the cheaper routes, and a claim about any source needs four links: a distinction a model can learn, an effect on the model, an outcome someone values, and a reason an equivalent result is not cheaper elsewhere.
5Asking is cheap, and the permission to ask is scarce
Manufactured experience has a history. AlphaGo Zero learned Go without human games from the rules, a win signal and self-play; DeepMind’s AdA adapted quickly within an authored task space using a curriculum and memory; Absolute Zero had pretrained models propose tasks for themselves and learn from code execution without an external post-training dataset.14 In September a small model went further, starting from random initialisation, generating executable byte programs as its curriculum and transferring to prediction on natural data, although natural validation data still guided the choice of model and settings.15 The result is far from general intelligence and cannot invent private facts, and it weakens the claim that useful general structure can only be learned from human records.
Building the environments is becoming a learned task as well. SPADE trains a designer that uses documents, execution and memory to keep challenges useful, and RACES found that composing verifiable environments raised a 14-billion-parameter model’s mean score from 48.8 to 51.3 compared with training on the components, with instances and training steps matched.16 Results like these weaken any asset claim that rests on hand-authored tasks being scarce. The designers still depend on checks, inputs and selection that someone specified, and those remain costs of production.
The limit lies in what each generator can establish. A code executor establishes consequences under the rules somebody implemented, and a hand-built simulator establishes consequences under the dynamics somebody assumed. A learned simulator estimates the dynamics from evidence, and only an experiment supplies new evidence about the external process; visual realism or long trajectories do not turn a simulator into one. Transfer is a separate achievement from manufacture: detectors trained on randomised synthetic images located real objects, game agents overfitted thousands of generated levels, and models trained on human demonstrations failed on unfamiliar websites.17 A workable division of labour uses external activity to discover missing distinctions, simulation to multiply practice around them and independent outcomes to test what transfers, and it counts the cost of keeping the simulator calibrated.
Let me take it one step further. A stored record answers the questions its production happened to preserve, while an environment allows a new question: change an action, vary a condition, observe the consequence. Choosing informative experiments is an established discipline, and models are making it cheaper.18 The scarce input then shifts from the archive to the ability to arrange an informative encounter with a process that matters, which requires permission to intervene, access to experts, reliable measurement and the patience to observe delayed outcomes. A record can often be copied after it exists, while a new cooperative episode has to happen before anyone can copy it. Proposing the question becomes cheaper with each model generation, and the permission, the access, the measurement and the valid interpretation of the answer often belong to other people. An organisation that contributes no unique text can therefore hold the scarce input, provided it can tell a successful candidate from a plausible one.
6Schrödinger’s workflow
A system that lets participants choose between two supplied explanations cannot record the third explanation they were never allowed to propose. A workflow recorded after its consequential choices were designed out teaches the workflow and not the judgement it replaced. Structure and constraint are different properties: a record can be precisely structured without prescribing each action, and a fully instrumented workflow can still exclude the alternatives one hoped to learn about.
The technical bet is specific. For tasks that need contextual judgement or revision under changing constraints, observing the process should improve transfer more when participants can exercise relevant discretion than when their strategy is prescribed, provided the distinctions they reveal are learnable and the outcomes can be assessed. The prediction concerns an interaction and makes no claim that maximum freedom is best, since more freedom also makes exploration, interpretation and credit assignment harder. It can be tested by crossing two conditions, prescribed or broader choice during the work and terminal or process records for training, and asking whether the benefit of process records is larger under broader choice, measured on tasks outside the environment that produced the records.19 Constraint has evidence on its side as well: HLER reported fewer critical research failures with a constrained architecture around an unchanged model, and because its gates can reject work, fewer failures did not mean more research completed.20
Observation has to keep its evidential categories apart. An action, a state change, a stated intention and an inferred motivation are different records, and explanatory text is useful without being a transparent trace of the process, for models as for people, since how far a model’s stated reasoning drives its answer varies with the model and the task.21 Feedback has limits as well: execution can verify a program against an incomplete test, human preferences carry judgement and can stay divided, and a learned reward can be exploited at its weaknesses.22 An independent evaluator is valuable when it measures the intended consequence, and a second model with a different name is not an independent evaluator.
Structured outputs have a related limit. Jev answers inside a set of options the customer supplies, which prevents invalid answers and still permits the wrong valid one. Five properties therefore need separate tests: a valid output type, a correct meaning, calibrated uncertainty, resistance to adversarial instructions, and an authorised, successful execution. TypeSafe’s documentation lists Jev’s failures on arithmetic, multi-hop questions, long context and prompt injection, and structured outputs do not remove substantive error.23
Recording people also changes them. In the illumination experiments at Western Electric’s Hawthorne plant between 1924 and 1927, output was reported to rise whenever the lighting changed, brighter or dimmer, and Henry Landsberger later read the relay-assembly results as showing that the workers responded to the attention of the study more than to the conditions being varied. The original data are weaker than the story, and the lesson survived in research design.24 Interfaces, payment, observation and the choice of participants all alter behaviour, so a record of people working with their discretion intact is an aim to be tested and no guarantee of representative behaviour.
There is still a bounded reason to care about actual human behaviour, since agents trained to coordinate with copies of themselves coordinate poorly with people.25 The most direct trained-agent evidence supports curated process supervision: training a software agent on curated records of how fixes were reached raised it from 39.6 to 42.4 per cent on SWE-bench Verified in a size-matched comparison, and to 50.4 per cent with the full curated set, a gain that also reflects recovered coverage of hard issues.26 It gives no support to maximal logging or unrestricted activity. What a model can learn from a record is bounded by the choices the people in it were allowed to make and by how well the record captures their consequences.
7The terms of access change what there is to learn
Offer a purchasing team a fee to record how it handles the next supply interruption, and the team changes. Some of the best negotiators decline, those who agree may negotiate differently while observed, and the resulting record reflects the offer as well as the negotiation. The terms of access therefore enter the production of information through a chain: the arrangement changes participation and conduct, participation and conduct change what can be observed, training extracts a difference, and the difference changes an outcome someone values. Each link can fail. More contributors need not produce better learning, and a reluctant population need not hold anything a willing substitute cannot supply.
A set of supplied options prevents invalid answers and still permits the wrong valid one.
Refusal can be rational, because information is non-rival in use and competitive in consequence: a record can be learned from without being used up, while the competence learned from it reduces its producer’s advantage. Jones and Tonetti formalised that tension in 2020.27 Both kinds of offer have costs in conduct. A generous arrangement recruits expertise and also attracts behaviour optimised for the payment, and a restrictive one preserves trust and can prevent the learning that would have paid for it.
The source owner also has to be identified. An employee, an employer, a customer, an expert and a rights holder can each control part of an activity, so the owner is whoever has the effective authority and practical capacity to supply a specified use, sometimes as a coalition. Holding the file establishes neither. Privacy protection and competitive protection address different losses: a model can learn a professional skill without revealing any confidential record, so a promise to learn only general patterns does not answer a contributor whose concern is competition, unless the resulting capability complements the contributor’s work, which has to be shown in the setting.
Suppliers set terms from the other side. Jev’s weights are shared across accounts, customers adapt it through the state, criteria and code they supply, and TypeSafe’s agreement permits training on customer data only with prior consent while reserving telemetry and voluntary feedback as separate channels.28 The default lets an enterprise supply useful private context while keeping its learning assets under its control, and it means TypeSafe’s durable advantage has to come from its technology, its operating cost, its distribution or a demonstrated renewal process, because no automatic loop converts customer traffic into better weights. Thomson Reuters says it does not use customer information to train its model either.29
Documented acquisitions establish demand for particular inputs and the obstacles to trading them, without measuring what any dataset taught any model. Reddit gave Google negotiated access to its content under a partnership announced in February 2024, and a federal judge in California set out in June 2025 how Anthropic had acquired its books, distinguishing training on them from building a library of pirated copies.30 Treating such records as prices confuses a legal outcome with a learning contribution.
A credible bargain can change what experience comes to exist, through compensation, useful services, limits on downstream use or a share of the return, and its credibility matters after the learning has happened, when the contributor’s bargaining position has changed. Rival laboratories can improve their offers too, so access on better terms is a candidate advantage and not a durable one by default. Withholding is one rational response to exposure and partial sharing is another: in Dubus and Legros’s model, sharing can create value by revealing synergies while strengthening a prospective rival.31 The analysis has to include the complete bargain, with shared development, selective licensing and acquisition, or it presents owners with a choice between surrender and isolation that they do not face.
8The growers can work for themselves
Winston’s crop and money leave in Dog’s van, and the fence who moves the stolen crop arranges to sell it to Rory Breaker, the man it was grown for. The growers had the garden and none of the trade, which is the position source owners held before they could produce capability themselves.
Thomson Reuters has taken the other position. It discussed its Thomson model publicly in July 2026, deployed it for professionals in August and published the technical report later that month. Its model starts from Qwen foundations and adds proprietary content, synthetic material, general replay, expert evaluation and an existing product channel. According to the report, the final large-model training run cost under $450,000 in GPU time and total development about $40 million, so most of the cost of becoming a producer lay in building the capacity to run and improve the process, not in the final run.32 The company also describes more than an archive: lawyers write rubrics and preference judgements, failures are diagnosed and fed back, and the starting foundation has been replaced more than once. The production asset is that organised learning capacity, the ability to state what a useful professional answer requires and to turn that judgement into training and a service, and no single ablation shows which ingredient is indispensable.
The model improved, and specific regressions bound the result. On the report’s broad aggregate Thomson rose from the Qwen model’s 73.0 to 78.5, against 79.5 for Opus 4.8, and its mean on seven public legal benchmarks rose from 63.2 to 69.6, with six of the seven improving; it was slightly less factual in matched legal research, 0.83 against 0.85, and fell on a general terminal benchmark from 54.0 to 48.3.33 Training also stood in for serving effort. Across four selected task families a single Thomson pass scored 71.9 against 71.2 for the base model making four calls, at about a quarter of the decode cost, so an up-front learning investment replaced expenditure that would otherwise recur at every use. Whether that repays the development cost depends on volume, refresh costs, margins and the alternatives available during the payback period.
The institution competes as a complete system. In a blind study of 3,035 tasks judged by 35 attorney-editors, the Thomson system was preferred in aggregate over each of five OpenAI and Anthropic systems, winning 54 to 64 per cent of legal queries against losses of 26 to 33 per cent and rating higher on completeness and usefulness against all five, with legal-specific tools and Reuters news search against their web search; Harvey and Legora were not tested. Holding the model fixed and removing the legal tools and instructions lowered legal preference, with the full system winning 57 per cent and losing 24 per cent of those comparisons, so the controlled contribution is the bundle and not an isolated archive.34 A rival has to replace the useful result of that combination, and matching its foundation alone is an incomplete substitute test.
Distribution supplies the commercial link. CoCounsel had reached one million professionals in 107 countries and territories by February 2026, before Thomson launched, and the model powers Tabular Analysis inside CoCounsel Legal.35 Distribution matters twice: it is a route for selling differentiated quality, and it supplies the volume over which a lower serving cost can repay development. The company keeps a multi-model strategy, has released a small research model, sells no standalone Thomson model, and in September announced a partnership with OpenAI covering its HighQ platform and a preview of a further CoCounsel experience, so producing, partnering and granting access sit inside one strategy whose strongest feature is the ability to replace the public foundation underneath while keeping the learning and distribution system on top.
The case shows competitive conversion: an information owner has exercised the production option, delivered measured gains within the evaluated scope and connected the capability to an existing commercial channel. Persistence is the open part, since the combination has to sustain customer preference, margins and further improvement through the payback period, and customer reach alone is no automatic training loop, particularly when customer data are excluded from training.
TypeSafe occupies a different position in the production economy the growers are entering. Jev is a component: the customer supplies state and criteria, Jev returns choices, scores or probabilities, the customer’s software executes, and a general model can supply the reasoning Jev does not attempt.36 Thomson brings intelligence to a source its owner holds, while TypeSafe sells a tool with which many owners can exploit theirs. A successful supplier is therefore not an example of a source owner choosing to produce, and passing most of a checklist does not establish a durable advantage for either company. William Stanley Jevons observed in 1865 that making a resource cheaper to use tends to increase how much of it is used, and better general models may increase the use of both Thomson and Jev while reducing the scarcity of what each sells.37 The comparison has to be between the contributions each company controls, whatever labels are attached to them.
Both positions can gain from better general models, and they retain returns differently. For a source owner, accessible production capability shortens the distance from possessing a valuable input to owning the differentiated service, and distribution then supports both repeated sales and the recovery of fixed learning costs, while a component supplier has to show why its contribution stays preferable as alternative components improve. The line between them does not follow company labels either. In August, Harvey’s Tenet research preview described post-training a Kimi K3 foundation with Fireworks on synthetic, public legal and expert data, with an agenda that includes helping law firms develop and own specialised models.38 Organisations outside the foundation laboratories are acquiring the means of model production, and the open question is which combination of evidence, expertise, technology and distribution keeps the strongest economics.
Any owner that can obtain general capability and combine it with a source that is hard to substitute has the production option, and exercising it requires adaptation expertise, evaluation, deployment, financing and a route to revenue, the complementary assets in Teece’s account.39 The owner need not pretrain a frontier model or sell general models. It can specialise a foundation, replace purchased capability, sell a domain product or keep a model for internal use, and each is a different business with different costs.
The decision rests on three measurements that should be kept apart: the source’s additional learned performance against the strongest alternative without it, which shows whether general progress makes the source more useful or replaces it; the owner’s net value from keeping the rights, which reflects its ability to exploit the asset; and a comparison of that value with the best available arrangement that grants the rights, which decides ownership.
Crossing the first measurement with the third gives four cases, and two of them correct common errors. An owner can want control for a transitional product opportunity while its learning advantage erodes, so a preference for ownership proves no unique technical value; a valuable source can rationally remain available to other producers, so licensing proves no lack of value. The prediction is a threshold. As foundations and development assistance improve, owners with substantial residual learning value, a development bottleneck that assistance can reduce and a practical route to deployment will find that producing under retained control yields more than their best external bargain, and a measurable subset will change strategy. An owner with distribution already in place can cross the threshold sooner, because it has a way to deploy and sell the gain. The option also works before entry, by raising the opportunity cost of granting access and strengthening the owner’s hand in negotiating it. That is a narrower claim than a withdrawal of valuable data from the market.
The price at which an owner grants the rights follows from that comparison. The minimum payment is the owner’s value with the rights retained less its value after granting them, and it rises only when the first rises faster than the second. A buyer can still meet it, a narrower grant can preserve the uses that matter, and a partnership can enlarge the total surplus.40 The assets that contribute most to future learning also create the strongest internal production opportunities, which is why the learning contribution and the outside option have to be measured independently.
The threshold moves because models supply parts of model production. OpenAI’s September account reports 3.1 agent-workdays per researcher workday, with more than half of the successful four-to-eight-hour tasks still needing human intervention; Anthropic’s index has AI leading 26 per cent of weighted research and development work, which is different from doing that work autonomously; and AlphaEvolve’s 23 per cent kernel speed-up reduced Gemini training time by 1 per cent, because the denominator is the entire training run.41 METR has begun measuring the experiment budget needed to find a validated training improvement, which points to the right unit: validated progress per complete expenditure, failed trials and human intervention included. A randomised study of experienced developers using early-2025 tools found them 19 per cent slower, a counterexample to universal productivity gains and not a verdict on current tools.42
For a clean comparison, the starting foundation, the development assistance and the original source should be varied separately. Holding the foundation and the source fixed while changing the assistance identifies a different channel from improving the foundation, and assistance that supplies new supervision is an additional information input as well as saved labour.43 The laboratories receive assistance earlier and run it on better infrastructure, and assistance also makes substitute data, simulators and verification cheaper, so the prediction concerns relative full costs and does not imply decentralisation. The growers can work for themselves where the assistance they can buy closes their development gap faster than alternatives erase their source’s contribution, and those conditions should be measured before entry is observed, since successful firms cannot supply, after the fact, the definition of the conditions that predicted their success.
9The robbers want the van, not the garden
Dog’s gang robs the growers with the help of Plank, one of their customers, and Eddy and his friends take the haul from Dog’s gang when it returns next door. None of them needs to know how to grow anything.
A rival needs behaviour good enough for the customer, and several routes reach it without the owner’s source: public data, a different architecture, a specialist model, verified synthetic practice, more inference, and permitted distillation from the owner’s outputs.44 MirrorCode rebuilt the behaviour of target programs from an executable reference and its tests, without access to the source code during the task. Success was incomplete and the running reference was itself a valuable input, but the study shows that control of original code does not always control the cheapest route to adequate behaviour.45
Tom bought the two antique shotguns for the robbery from a fence, neither man knowing what they were, and they were worth hundreds of thousands of pounds to a collector and nothing extra to a robber. A proprietary archive can be in that position, irreplaceable as a historical object and replaceable as a means of producing behaviour. Economic equivalence is narrower than general intelligence: a competitor serving one valuable task family does not need the original model’s breadth, while matching a benchmark average can miss the rare failures that decide professional value. Target outcomes, acceptable failure severity, latency and assistance budget have to be stated before equivalence can be judged.
None of this depends on incumbent laboratories being unable to adopt new architectures, or on challengers replacing language modelling. Methods, architectures and sources are different routes to a gain, and they differ in how far the gain travels: a large technical improvement can pass readily to rivals, while a modest one can last when its production is hard to reproduce. The test is whether the source keeps its value against the strongest learner and substitute available to either side.
Every comparison here is dynamic, because the alternative advances alongside the asset. The September releases from OpenAI and Anthropic reported further gains in agent performance and cost per task, so the substitute set has to include the current generation.46
Assets fall into two classes as general capability improves. Some fill a capability gap that general progress can close, such as a specialised corpus, an annotation procedure or a workflow teaching a skill the next foundation will supply adequately. Absolute performance using such an asset can rise while its incremental contribution falls, though the investment can still pay if the return arrives within its useful life. Others control a structurally better way to produce or validate the next useful increment: measurements or interventions that existing evidence cannot supply, a lower full cost of generating and verifying experience, or participation and complementary assets that rivals cannot reproduce on equal terms. For those, better general models can raise the usefulness of the source and lower the cost of exploiting it at once. An asset can move from one class to the other as its target is solved or its production improves, and persistence requires the system to keep offering a better attainable result or cost, without requiring imitation to be impossible.
Which constraint binds changes the priorities. Where broad transferable competence is the binding constraint, differences in general capability dominate, and the leading research priorities are relevant coverage and valid feedback, followed by economical selection, learner-appropriate difficulty and useful process structure. Where domain evidence, dependable execution or cost binds, specialist capability can carry the decisive value while the general frontier keeps advancing, through dependable completion, specialist reasoning with evidence, steerability, consistent action over time, adaptation to changed conditions and efficiency at the required quality. These are different constraints and not historical stages, so a specialist opportunity need not wait for general capability to converge. The requirement then depends on the workload: formal mathematics needs correct reasoning and economical search, empirical science needs external validity, professional analysis needs specialist reasoning with evidence, and operational agents need correct completion under constraints.
Private return also has a time limit. A regime can create substantial value and still be a poor investment if the advantage is copied before its cost is recovered. A temporary edge can still pay if its returns arrive before substitution; a durable advantage keeps a useful performance or cost difference through the payback period; and a reinforcing moat needs one more mechanism, in which the owner keeps enough of the benefit to improve the next round of production and that improvement stays ahead as rivals advance. Four measurements separate them: payback time, time until a competitor reaches equivalence at a stated quality and cost, validated learning yield per full expenditure, and the share of created value the owner retains.
A corpus can stay useful for years, a live environment can produce increasingly redundant episodes, and a loop compounds only when earlier learning makes later learning cheaper to produce or extract.47 AgentCL illustrates the difference, since external memory improved reuse within its task stream without improving the held-out coding score, so stored activity, useful local memory and transferable competence are separate assets.48 Concentration in general foundations and durable specialisation among source owners can coexist, and a winner-take-most market is a stronger claim that needs evidence about substitution, access, bargaining and reinforcement, which a learning curve or a feedback loop does not supply.
The value of an archive and the value of access can therefore diverge. As models improve at exploiting existing evidence, some archives lose value while the capacity to obtain the next discriminating observation gains value, provided the target keeps changing and the yield repays the cost of observation. On a stable task, the robbers take everything of value in the van. On a moving one, the value stays with the capacity to grow the next crop.
10Tom is still on the bridge
The film ends with Tom hanging over the rail of a bridge, the two shotguns on the ledge beyond it and his phone in his mouth, while his friends ring to tell him what the guns are worth; the frame freezes before he answers. Both bets are in a similar position, with their outcomes open and the ways they could lose stated in advance.
The technical bet fails if process records help but broader choice adds nothing economically meaningful under a precise test, if the gain disappears outside the environment that produced it, or if synthetic practice and targeted questions win at full cost. The strategic bet fails if stronger foundations eliminate the source’s incremental contribution or make an adequate substitute cheaper over the claimed horizon, and it also fails if the value of retained control rises while the best available bargain rises faster. It narrows to a bargaining result if rights become more valuable while supply remains available at better prices, and the claim of a durable private advantage fails, while the learning result survives, if cheap imitation arrives before payback. These are distinct losses, and the general point that learning has an economic organisation cannot be used to claim every outcome as success.
Five stronger claims would be convenient, and none of them survives the evidence: that better models necessarily strengthen owners, that every valuable AI supplier owns a learning source, that observed entry proves superior retained returns, that structured outputs eliminate substantive error, and that synthetic methods remove all information scarcity. The three effects hold as a conditional account: current research strengthens both the pressure from substitutes and the range of production inputs available to owners, without establishing either concentration or decentralisation.
Each bet implies its test. The technical bet needs the factorial comparison of freedom and observation, with transfer measured on independently authored tasks and deployment inputs held fixed. Answering the investment question requires each alternative, from curation and retrospective annotation to targeted queries, outcome-only learning, verified synthetic experience, calibrated simulation and additional compute, to compete at matched total cost, and the strategic bet needs adaptation repeated across successively stronger foundations, with the source’s contribution, the owner’s cost and the best offer estimated separately, and a competent rival team given a realistic budget to reach the result another way.49 Payback and catch-up belong on one horizon, and repeating the comparison on more than one capable foundation separates a contribution that survives a change of learner from one tied to a particular architecture. Controlled experiments are one source of evidence among three: an independently evaluated working system settles an existence claim that argument alone leaves open, and the market shows who enters, which rights they keep, how offers change and how long catching up takes.
A remaining valuable asset can therefore be described before the experiments are run. Technically, it makes a consequential contribution against improving alternatives, offers a usable route from source to capability, supports valid discrimination between better and worse results, and holds up across the range of cases claimed for it. Economically, it connects to valued outcomes, stays favourable on a complete account as production methods improve, has a feasible path to deployment, and beats the relevant substitutes. Strategically, its controller can capture a return from it, through access, cooperation, deployment or distribution, general progress does not promptly absorb its superiority, it persists through durable evidence or renewed advantage, and, where entry is claimed, producing under retained control beats the best available licence. The twelve conditions apply together to a named asset, workload and horizon, and they are not points to add into a score. Uniqueness, richness, secrecy and human provenance are not among them.
The technical bet waits for an experiment. The strategic bet is already being run: Thomson shows that source ownership, obtainable model capability and a route to customers can be combined into a productive business, and the open question is whether such a combination holds its advantage through successive model generations. As general models improve, they make the means of producing intelligence more accessible, and where valuable learning experience stays harder to substitute, its owners can acquire those means, turn the experience into differentiated capability and keep the resulting business instead of selling its essential input. That is a stronger production option, neither an inevitable withdrawal of data nor a guaranteed moat, and the contest is whether obtainable production capability advances faster than economical substitution for the next valuable learning opportunity. Once a capable model can be brought to the garden, the gardener has another business to consider.