Slow Fast

Action Jack,
or the Problem With Large Action Models
in the Enterprise

A man in a suit addresses a start-up from the front of an open-plan office, arms spread, with the company name in large letters on the wall behind him.
Action Jack loves taking action, especially in contexts he does not entirely understand.Silicon Valley (HBO, 2014 to 2019), created by Mike Judge, John Altschuler and Dave Krinsky.

1A man of action arrives at a company he does not understand

In the third season of Silicon Valley the board of a start-up that has built a compression algorithm installs a chief executive called Jack Barker. He has run companies before, he speaks in diagrams, and within a week he has drawn the conjoined triangles of success on a whiteboard, found a way to sell the algorithm inside a metal box to enterprise buyers, and told the engineers that the product is whatever the customer will pay for this quarter. He is decisive, likeable and wrong, and the writers gave him the name the cast used: Action Jack.

In the summer of 2024 the next frontier for AI companies was the large action model. Adept, H and Orby were the names on the pitch decks, and between them and their imitators more than half a billion dollars had been raised in twelve months on an idea simple enough to fit on Jack’s whiteboard: train a model on enough recordings of people using software and it will learn to use software, clicking the buttons a person clicks, in every application a company owns, without anyone changing the applications.1 The prize was real. Companies were on course to spend a trillion dollars on software that year, four times what they had spent a decade earlier, while revenue per employee stayed flat and the headcount in IT and shared services grew, which is a polite way of saying that the last generation of software failed to deliver what it was bought for.2 We need less software and fewer people doing meaningless work, and a model that could operate the software we already have looked like the shortest road there.

I wrote then that it was much harder than it seemed, and two years later I would put it more strongly. The difficulty was never the model. It was the context the model was being sent to act in, and Action Jack is the patron saint of acting in a context you do not understand.

2The enterprise is not Phoenix, and it shoots back

Automating enterprise software sounds easier than building a self-driving car. Everything is already in a computer; the world is digital and controlled. It is not. A large company runs on mainframes, on PDFs of handwritten notes, on shared inboxes, on SharePoint folders nobody has opened since the merger, on fax, and on a spreadsheet from 1999 with a macro whose author has retired, all of it holding the everyday world together. The self-driving car took twenty years and more than a hundred and fifty billion dollars to reach the point where it could almost drive as well as a person, and in the month I wrote the post one of the best of them hit a telephone pole on a straight road in Phoenix.3 Actions are hard, and actions with consequences in the world are the hardest kind.

The enterprise is harder to drive in than a bad night in San Francisco, and it has a property no road has: it fights back. Put an agent on every employee’s browser to record what they do and you will meet the IT department, then the data protection officer, then the EU’s AI Act. Ask for a connection to the back-end systems and you will meet IT again. Let the model take a decision that touches the ledger and you will meet the chief financial officer, and the CFO does not laugh. A car in Phoenix has to cope with the road. An action model in a company has to cope with everyone whose job it is to stop things happening, and most of those people are right.

The fair concession, two years on, is that the cars got there. In 2026 the company whose car hit the pole carries paying passengers in several cities with nobody at the wheel, and the sneer in my 2024 post has aged worse than the argument. The argument holds because the road did not fight back and the enterprise still does.

3The data costs more than the model, and the humans cost more than the data

Suppose the war is won and the recordings are flowing. What is the model supposed to learn from: the process, the rules, the screenshots, the video? Reading a screen through a model is expensive, and interpreting what a person meant by a sequence of clicks is expensive work on top of it. A browser extension will do for consumers and will not survive a security review at a bank. And if the model is to be large, it needs a great deal of this material, which is owned by a great many vendors, an average of three hundred and seventy applications in a large company before the legacy systems are counted, each of them building its own assistant and none of them with a reason to hand over the record of how their product is used.4 Gathering that costs many times what it cost to scrape the consumer internet, before a single model is trained, and the rights question underneath it is where music was before streaming: everyone has a piece of the catalogue and nobody has a licence to the whole.

Then there are the humans, who are the hardest part. We are non-linear pattern machines carrying knowledge that is invisible to a recording. The consumer models reached their quality on human feedback and labelling done by large outsourced workforces, and that does not transfer to the enterprise, where the context and quality of the feedback are the point and the person who knows whether a purchase order was handled well is the buyer, not a contractor in another time zone. Legacy applications were not built to collect that judgement, so you build a layer to collect it, with the context of the record attached so that the feedback means something, and before long you have rebuilt most of the application from scratch in order to learn how to operate it.5

4Look, no hands

A convincing demo is easy. Adept’s early videos showed an agent working through Salesforce and buying a plane ticket, and they were as compelling as the first self-driving demonstrations were twenty years before. Three days before I published, Amazon hired Adept’s chief executive and most of its founders and took a licence to its work, which is what happens to a company with a co-author of the transformer paper and four hundred million dollars in funding when it tries to do actions top-down in the enterprise.6 H had raised the largest seed round in French history in May and lost three of its five founders by August. The device that had launched in January with a large action model on the box turned out, when people looked, to be running scripts.7 Two years on, none of the three companies on the pitch decks is what it set out to be. Adept never shipped the product; its investors were paid back and a rump of thirty people carried on under a new name for the same idea, and the founder Amazon hired left Amazon in February 2026 to start again. Orby was sold to a call-centre software company in August 2025 for an undisclosed sum and folded into a suite. H is the one still standing, with a new chief executive from Palantir, and what it sells now is models that read screens and click, which is the button-pusher under another name.8

Then the labs took the category over, and this is where the post has to be corrected in the open. In October 2024 Anthropic shipped a model that operates a computer from screenshots, OpenAI followed in January 2025, and on the standard test of using a real desktop the best systems went from finishing one task in eight to beating the human baseline inside two years.9 The button-pushing got good, and the button-pushing was never the difficulty. An agent clicking through a screen removes none of the cost of the screen, which still has to be designed, built and kept by the people who were doing that before; it adds a second user who stumbles when a button moves, and one that can be given orders by the page it is reading, a weakness OpenAI does not expect ever to be eliminated.10 It turns pixels back into the structure the application spent most of its code turning into pixels, and it pays for both translations, twelve minutes for a change of line spacing that takes a person under thirty seconds. The demo had no hands. The product has two hands and no idea what the buttons are for.

5Apple generated the world instead of driving in it

Apple is never first and is often right, and in June 2024, a few weeks before the post, it showed what it had been waiting for. Its assistant would not learn to operate applications by watching them. Every application on the phone would be made to expose its capabilities in a form the assistant could call, which Apple named App Intents, and the assistant would orchestrate what the applications had been made to offer.11 It is the difference between teaching a car to read the world and rebuilding the roads so that the car does not have to: Apple used its distribution and its control of the platform to generate the world, which is easier with today’s tools, better for the person holding the phone, and cheaper in compute than seeing, reasoning and acting, which is also why it runs on a battery.

Apple has since found out how hard even that is, and the personal assistant it demonstrated slipped by more than a year.12 The design was still the right one. Understanding a world that was not built for you is the expensive road; making the world expose what you need is the cheap one, and Apple could take it because it owns the world its assistant works in.

6Software is not bricks and asphalt

We cannot remake the physical world to suit a car. Software is not bricks and asphalt. Nothing about an enterprise application is fixed except the habits of the people who paid for it, and that is the way out of the problem: stop automating the software we have. Start from a contained problem, its data, its defined interfaces, what goes in and what must come out, and generate the application for that problem with the ability to act already inside it, so that there is no old screen for a model to learn and no button for it to find. The model that reads a screen is a bridge out of a system you are leaving; the model that writes the system is the destination.

In 2024 we called the unit of that a work block, and we said two things about it that I would still sign. Each block should be composable, so that a large piece of work is many small ones linked together rather than one large model doing everything, and each should be validated by a human who understands the outcome, so that trust in the work done by people and machines together is built in and not bolted on. The goal was never a single large action model. It was a platform for millions of small ones, and new systems of work built from the bottom up instead of a model trained to click through the old ones and reinforce every silo they made.13

The lesson of Apple and the lesson of the self-driving car are the same lesson: the intelligence is cheap and the world it acts in is the expensive part, which is why the world is what you design.14 Action Jack was not wrong to act. He was wrong to act in a context he had not built and did not understand, and the fix was never a smarter Jack. It was a company he could understand, made so that his actions had somewhere honest to land.15

Notes

  1. 1Adept raised a 350 million dollar Series B in March 2023, bringing its total to about 415 million at a valuation near a billion; H, in Paris, announced a 220 million dollar seed round in May 2024; Orby had raised about 55 million by 2025. Take it as the half billion of the opening.
  2. 2Gartner’s forecast for 2024 put worldwide enterprise spending on software at just over one trillion dollars, roughly four times the level of a decade earlier; the flat revenue per employee and the growth of IT and shared-service headcount were the 2024 post’s own observations from inside the trade. Take it as the size of the prize the action models were aimed at.
  3. 3In May 2024 Waymo issued a software recall after one of its vehicles struck a telephone pole in Phoenix on a straight road; the “twenty years and a hundred and fifty billion” is the industry’s own rough accounting of what autonomy has cost since the DARPA challenges. By 2026 Waymo carries paying passengers in several American cities with no safety driver, which is the concession made in section 2. Go there for the pole, and for how the sneer aged.
  4. 4Zylo’s SaaS management reports put the average number of SaaS applications in a large company at about 371 in 2023, not counting legacy systems. The assistant per application, and why none of the vendors has an interest in sharing, is in the author’s Why Enterprise Software Sucks (September 2023). Go there for the three hundred and seventy owners of the data an action model would need.
  5. 5The human feedback and labelling behind the consumer models is described in OpenAI’s InstructGPT paper, Ouyang and colleagues, Training language models to follow instructions with human feedback (2022). Take it as the method that does not transfer, because at work the person who can judge the result is the person doing the work.
  6. 6The Verge, 1 July 2024: Amazon hired Adept’s chief executive David Luan and its co-founders and licensed its technology, following the pattern Microsoft had set with Inflection in March. Go there for the fate of the best-funded action model, and as no comment on its people, whom Amazon hired for what they had built.
  7. 7H’s founder departures were reported in August 2024, three months after its seed round. The Rabbit R1, launched in January 2024 with a large action model as its selling point, was shown by reporting in June 2024 to be driving its supported services with pre-written browser scripts. Take it as the summer the category’s demos were taken apart.
  8. 8Semafor reported in August 2024 that Adept’s investors roughly recouped the 414 million raised, the remaining company continued with about thirty people under a new chief executive, and GeekWire reported David Luan’s departure from Amazon in February 2026. Uniphore announced its acquisition of Orby on 28 August 2025 for an undisclosed sum. H replaced its chief executive in June 2025 with the former head of Palantir France and now sells computer-use models under the name Holo. Go there for where the half billion went.
  9. 9Anthropic released computer use in public beta in October 2024 and OpenAI released Operator on 23 January 2025. OSWorld, the benchmark of 369 tasks on a real desktop, reported a human baseline of 72.36 per cent against 12.24 for the best early model, and by 2026 several systems report scores above the human figure. The same figures carry the argument in the author’s Back to the Future (of User Interfaces) (October 2023). Go there for how fast the button-pushers improved, which is conceded here.
  10. 10OpenAI, writing about its Atlas browser in December 2025, said it does not expect prompt injection ever to be eliminated. The twelve minutes against thirty seconds is from Abhyankar, Qi and Zhang, OSWorld-Human (2025). Go there for what the robot at the wheel costs, in minutes and in exposure.
  11. 11Apple announced Apple Intelligence and the expanded App Intents framework at its developers’ conference on 10 June 2024: applications expose actions and content that the system’s assistant can invoke. Take it as the platform owner’s version of generating the world.
  12. 12In March 2025 Apple said the more personal Siri it had demonstrated would take longer than expected, and the features moved into 2026. Take it as the concession that even the owner of the world found the world hard to wire, and note that the design did not change.
  13. 13The work block of 2024 became the workblock of Beyond Work, described in the author’s A Love Letter to Consulting and the Human Factor (February 2026). The platform is under construction, and no deployed outcome, measured saving or productivity result is claimed here. Take it as the maturity label on section 6.
  14. 14The same conclusion, reached from a different direction, is in the author’s The Psychology of Every Day Machines, the Inversion of Meaning and the Alien Mind Without a World (September 2026); the robot at the wheel is in Back to the Future (of User Interfaces). Take it as a pointer, not a prerequisite; this essay stands on its own.
  15. 15First published on LinkedIn on 3 July 2024, at about a thousand words, under the title Action Jack or the problem with Large Action Models in the enterprise. This edition keeps the post’s claims in the post’s order, the self-driving car, the warzone, the data and the humans, the demo, Apple and the generated world, corrects what the labs proved since, and adds the receipts. Take it as the lineage.