Slow Fast About

Back to the future (of user interfaces)
and why AI agents won’t solve
enterprise software’s problems

A passenger leans forward from the back seat of a taxi to argue with the driver, who is a smiling plastic mannequin bolted behind the wheel.
The car can drive itself, and somebody still built a driver to sit at the wheel.Total Recall (1990), directed by Paul Verhoeven.

1A stone needs no training

You want the inside of a nut. You pick up a stone and bring it down, the shell breaks, and you eat. One object stood between the wanting and the eating, a river had shaped it, and it came without a manual.

A tool translates intent into action, and the measure of a tool is how little is lost on the way. The stick, the axe and the lever improved the translation, with more reach or more force, and all of them kept the distance between what you meant and what happened to about the length of an arm. The part of a tool that faces the person is what we later learned to call its user interface, and for most of history it was a grip.

The computer stretched that distance. A machine that takes instructions and nothing else has to be addressed in its language, so for the first decades people learned the machine’s words and typed them, and the command line had a lopsided economy: it was cheap to build, because a new capability was a new word with a function behind it, and expensive to use, because the vocabulary lived in your head.

On 9 December 1968 Douglas Engelbart sat at a console on a stage in San Francisco with a small box on a cable under his right hand, and for ninety minutes about a thousand people watched a spot follow that hand across a display twenty-two feet high.1 They called the box a mouse, on account of the tail. Xerox built a desktop of windows and icons around it, Steve Jobs and Bill Gates both helped themselves, by Gates’s own account, and by the middle of the 1980s a person could point at what she wanted.2

The graphical interface reversed the economy of the command line. It was cheap to use, because what you could do was in front of you, and expensive to build, because someone had to draw whatever you might point at, decide where it went and keep it working from one release to the next. On those early screens a designer placed the pixels one at a time, which made the interface artisanal work in the way of a hand-stitched shoe, and the industry has never priced it that way. The cost of the interface had moved from the user’s memory to the maker’s payroll, and it has grown there ever since.

2The screen is still made by hand, and the screen is most of the software

Put two pieces of work software from different vendors side by side, an expense tool and a purchasing system, and you will have trouble telling them apart. The buttons, the tables, the date pickers and the grey sidebar come out of component libraries, so the pixels arrive ready-made, and the screen is still composed, wired and tested by hand. The libraries are why one application resembles the next, and they are why applications hold still. A screen built from components is fixed on the day it ships, and changing it takes a designer, an engineer and a release, or else a layer of configuration so convoluted that Salesforce certifies the people who operate it, in an exam that gives a fifth of its marks to objects, fields and page layouts.3 Adding a field to a form is a profession.

In 1992 Brad Myers and Mary Beth Rosson surveyed seventy-four software projects and found that on average forty-eight per cent of the code went to the user interface, along with half of the time spent writing it, and that the projects built on interface toolkits alone were nearer sixty.4 That was software for one kind of machine, used by someone at a desk. Work software in 2026 runs in a browser and on a phone, in a dozen languages, for people who cannot see the screen as well as for those who can, with form state held on the client and a suite of tests that clicks through all of it before a release, and nearly all of it is built on toolkits. I paid for an application like that for fifteen years, and my estimate is that about sixty per cent of the code in a modern SaaS product is interface, libraries included, that the product runs to more than a million lines, and that keeping those lines alive costs more over their life than writing them did.5

An industry that spends most of its money on the part with the buttons learns to ration buttons. Each new capability costs a screen, each screen costs a team for as long as the product lives, and the rational vendor ships fewer capabilities and more configuration, which goes a long way towards explaining why work software has changed so little in forty years.

Figma made drawing a screen cheaper than it had ever been, and that made matters worse, because a screen that is cheap to draw costs as much as before to build, test and keep, and cheap drawing produces more screens. The market has put a price on the drawing three times. In September 2022 Adobe, which had sold the first artisans their tools, offered twenty billion dollars for Figma; regulators in London and Brussels objected, and in December 2023 Adobe walked away and paid a billion dollars for the privilege. When Figma listed in New York in July 2025, its first day of trading valued it at more than twice what Adobe had offered. In September 2026 it is worth about twelve billion, with revenue growing by more than forty per cent a year, because investors have come to fear that the models will do the drawing.6 The market may be wrong about Figma, whose customers are still paying, and it is right about the drawing. The screen is most of the software, it is made by hand, and the price of the hand has started to fall.

3An agent is a robot at the wheel of a car that can drive itself

In Total Recall, made in 1990, the taxi comes with a driver called Johnny, a mannequin in a cap fixed into the front seat, who makes small talk and cannot understand where his passenger wants to go. The joke is that a car able to drive itself has been given a dummy to sit where the human used to sit.

The software industry built Johnny in earnest. For a decade UiPath and Blue Prism sold robotic process automation, which means scripts that push the buttons on interfaces drawn for people, and because the scripts had to be managed they were given interfaces too, drawn for the people who manage the robots that push the buttons, an Inception of user interfaces. In 2023 the language models arrived and the idea was funded again under a new name, the agent. Adept raised three hundred and fifty million dollars to build one that clicks and types its way through other companies’ software, and fifteen months later its chief executive and co-founders went to work for Amazon.7 Then the labs took the category over. In October 2024 Anthropic shipped a model that operates a computer by looking at screenshots and moving a cursor, OpenAI followed with Operator in January 2025, and on the standard test of operating a real desktop the best system went from finishing one task in eight to beating the human baseline inside two years.8

They push buttons well, and pushing the button was never the difficulty. An agent removes none of the cost of the interface, because the screens it clicks through still have to be designed, built, tested and maintained by the people who were doing that before, and it adds a second user to keep the screen working for, one that stumbles when a button moves and that can be given orders by the page it is reading, a weakness OpenAI does not expect ever to be eliminated.9

Let me take it one step further. The figures on the screen began as structure in a database. The application spends most of its code turning that structure into pixels a human eye can read, and the agent then spends money turning the pixels back into structure, decides, and steers a cursor to a button whose one job is to call a function that was there all along. Two translations cancel each other out and we pay for both. When researchers timed it in 2025, changing the line spacing of two paragraphs took an agent twelve minutes and a person under thirty seconds.10

It is a robot in the driver’s seat of a Tesla with its hands on the wheel. The car has a computer that can steer, the wheel is there for the human, and we have built a second machine to operate the first machine’s human interface. A trebuchet will crack a nut, and no army built one for that, because a stone is cheaper.

There is an honest use for the button-pusher, which is the thirty-year-old system with no other way in; robotic process automation earned its money there, and a model that reads screens is a better version of it. That is a bridge out of a system you are leaving, and the industry is selling it as the destination. The labs’ other answer, which lets the model skip the screen and call the function, is right for the machine and leaves the person where she was, in front of a screen someone drew in 2019 for a job that has since changed. The agent takes over the clicking and leaves the screen where it was, and the screen was the problem; while screens are made by hand, each new user of them, person or robot, is another bill.

4The interface should arrive with the problem and leave with it

A buyer has an invoice that does not match its purchase order. In the system she has, that is four screens, two of which she must remember the way to, and a field she may not edit without a ticket to an administrator with a certificate. In the system she should have, she says what is wrong, and what appears is the order, the invoice, the difference between them and the two actions open to her, each with its consequence stated before she chooses. She chooses, the choice is recorded, and the screen goes away, because it was made for that invoice and that buyer and there will be no second occasion for it.

An interface is a tool for a problem, so it should be shaped around the problem and the person who has it, present while she needs it and gone when she does not. The field’s name for this is generative UI, and in 2023 at Beyond Work we called ours PromptPlus.11 A model can draw a working screen in less time than it takes to file the ticket asking for one, and that removes the cost the industry was organised around. The command line was cheap to build and expensive to use, the drawn screen was cheap to use and expensive to build, and a generated screen is cheap on both sides, which no interface has been since the stone.

A generated screen is as trustworthy as what it renders. If the model that draws the table also decides what the number in the table means, the result is a persuasive picture of nothing, and at work somebody signs for the number. So the amount, the authority and the record of what was done have to live somewhere the screen cannot edit, exact and replayable, and above that layer the screen is free to be whatever helps this person judge this case.12

My bet is that an interface generated for the problem, over meaning kept exact underneath it, reaches a useful result at a lower total cost than an agent clicking through screens that were drawn for someone else, and that the drawn screen ends up where the command line did, kept by the specialists who like it. We are running that bet at Beyond Work, with a language called Sakio underneath it, and nothing here claims the result in advance.13

For fifty years we drew our way away from the stone, a tool that arrived with the problem, fitted the hand that held it, asked for no training and was put down once the nut was open. A robot sitting in front of the drawings does not bring that back, and a tool generated for the problem does. The future of the user interface is a stone.14

Notes

  1. 1Douglas Engelbart demonstrated the oN-Line System at the Fall Joint Computer Conference in San Francisco on 9 December 1968: ninety minutes, about a thousand people, a display twenty-two feet high on a projector lent by NASA, and the computer itself thirty miles away in Menlo Park. The session was later named the Mother of All Demos, and DARPA’s timeline entry has the essentials. Go there for the first public mouse, and for how much of the modern desktop was on stage in one afternoon.
  2. 2Andy Hertzfeld, A Rich Neighbor Named Xerox, Folklore.org, retold in Walter Isaacson’s Steve Jobs (2011). Accused by Jobs in 1983 of stealing the Macintosh’s interface, Gates replied that they both had a rich neighbour called Xerox, and that he had broken in for the television and found that Jobs had already taken it. Go there for the most candid history of the graphical interface, from one of the two men who took it.
  3. 3The Salesforce Certified Platform Administrator exam gives a fifth of its weight to a section called Object Manager and Lightning App Builder, which covers creating, deleting and customising fields and page layouts. Take it as the distance between configuring a screen and asking for one.
  4. 4Brad Myers and Mary Beth Rosson, Survey on User Interface Programming, CHI 1992. Seventy-four responses; on average 48 per cent of the code went to the interface, with 45 per cent of design time, 50 per cent of implementation time and 37 per cent of maintenance time; projects using toolkits alone came in around 60 per cent. Go there for the measurement that everyone’s estimate, mine included, leans on.
  5. 5The sixty per cent is my estimate, argued and not measured: the 1992 figure for toolkit-built projects is the floor, and the browser, the phone, localisation, accessibility, client-side state and interface testing have all arrived since. Robert Glass, Facts and Fallacies of Software Engineering (2002), puts maintenance at 40 to 80 per cent of what software costs over its life, 60 on average. Take it as an estimate with its reasoning showing, to be replaced by a measurement when someone has one.
  6. 6Adobe agreed to buy Figma for twenty billion dollars in September 2022 and abandoned the deal in December 2023 after objections from competition authorities, the British one among them, paying a termination fee of one billion dollars. Figma priced its initial public offering at 33 dollars a share on 30 July 2025 and closed its first day at 115.50, more than three times that price, which valued the company at between forty-five and fifty-six billion dollars depending on the count used. In September 2026 its market value is about twelve billion. Revenue in the first quarter of 2026 was up 46 per cent on the year before, and the explanation analysts give for the fall is fear of design tools built on the models, Google’s Stitch and Anthropic’s Claude Design among them. Take it as the market pricing the drawn screen twice, once for what it costs and once for its end.
  7. 7Adept raised a Series B of 350 million dollars in March 2023. In June 2024 Amazon hired its chief executive, David Luan, and four co-founders and took a licence to its technology, and the company carried on with the staff who remained. Take it as the fate of the best-funded button-pusher of 2023, and as no comment on its people, whom Amazon hired for what they had built.
  8. 8Anthropic released computer use in public beta in October 2024 and OpenAI released Operator on 23 January 2025; both read the screen and act through a cursor and a keyboard. OSWorld sets 369 tasks on a real desktop; its authors reported a human baseline of 72.36 per cent against 12.24 for the best early model, and by 2026 several systems report scores above the human figure, under test conditions that differ from one entry to the next. The other route, in which the model calls a function through a published interface and never sees the screen, is what the Model Context Protocol (Anthropic, November 2024) standardised. Go there for how fast the button-pushers improved, which is conceded here.
  9. 9Writing about its Atlas browser in December 2025, OpenAI said it does not expect prompt injection ever to be eliminated, any more than scams and social engineering have been; Britain’s National Cyber Security Centre had given a similar warning earlier that month. Take it as the new problem an interface acquires when its user obeys what is written on it.
  10. 10Reyna Abhyankar, Qi Qi and Yiying Zhang, OSWorld-Human: Benchmarking the Efficiency of Computer-Use Agents (2025). The twelve minutes against thirty seconds is their example; the leading agents took between 1.4 and 2.7 times the steps a person needs, and most of the waiting was the model planning and reflecting between clicks. The measurement dates from June 2025, and whatever the agents have gained since, the two translations remain. Go there for what the robot at the wheel costs in minutes.
  11. 11Yaniv Leviathan and colleagues at Google Research, Generative UI: LLMs are Effective UI Generators (November 2025). It was released in the Gemini app and in Search, and human raters strongly preferred the generated interfaces to ordinary model output once generation time was set aside. PromptPlus was Beyond Work’s name for the idea in 2023. Go there for the point at which the generated interface stopped being a demo.
  12. 12The author’s The psychology of every day machines, the inversion of meaning and the alien mind without a world (2026), section 6, on keeping the amount, the authority and the recorded effect exact while the rendering stays free. Go there for the long version of what has to sit underneath a generated screen.
  13. 13Sakio is in development. The bet is described as designed and under construction, and no deployed outcome, measured saving or productivity result is claimed. Take it as the maturity label on this section.
  14. 14First published on LinkedIn on 15 October 2023 and revised in September 2026. The argument is the one made in 2023. The receipts are new, because in the meantime the Adobe deal collapsed, Adept’s founders went to Amazon and the labs shipped the button-pushers, and the teaser that closed the original has been replaced by the argument it was teasing. Take it as the lineage label.