Published on

The Next Unhobbling Gain is Self-Creating Software

AI can increasingly do the software work required for its own adoption.

2758 words14 min read–––
Views
Two hands emerge from a sheet of paper, each drawing the other into existence.
M.C. Escher, Drawing Hands (1948). © The M.C. Escher Company.

By now it is clear that by simply scaling models and spending a lot of money, we can do work that would be impossible for a human to accomplish on a similar timescale or budget.

Why then has the diffusion of this powerful AI continued to be so small? Why does AI’s impact on the overall economy so far seem as if it is just another general purpose technology? I would argue that it’s because we haven’t yet achieved the unhobbling gain that will create an exponential in diffusion by making powerful AI intuitive and easy to use for all users.

Much of the work required to adopt AI is itself software work. Forward-deployed engineers come in and connect information and systems, create workflows, and build interfaces that a person can comfortably use. One bottleneck is the companies that exist to deploy AI, and the people they have to deploy it across a wide scope.

We could solve this problem of adoption sooner than people think, through the development of self-creating software: model harnesses that can build, maintain, and revise their own tools and interfaces to better suit the user’s needs. This requires greater model capabilities in software engineering and training UI taste in models, but seems to be the clear endgame.

This could make existing capabilities easier to use, while bringing models into more real tasks from which they can learn. Today, setting up this flywheel in a new field requires substantial human effort. Self-creating software could automate much of that work.

Self-creating software is, in effect, the automation of the software company and its deployment arm. Once models can perform this work cheaply enough, an application no longer needs a large market to justify its existence, and only needs to save one person more effort than it takes to create and maintain.

Coding shows what this could look like.

Intuitively, we all know that a chatbot is not the easiest interface for people to interact with. It’s why we created Claude Code and Codex for computer scientists and software engineers, and why new personal agents from Instinct to GrokBot and Meta’s Muse have emerged to try to target the general consumer, even if they still remain mostly adopted among San Francisco-adjacent circles.

Before Claude Code existed, we used to copy and paste a lot of files and ask the chatbot to write some Python code. With coding agents, the agent is going through the file system and has access to tools, getting much better traces of how the work gets done. It then becomes a lot easier for a user to entrust more complex work to it, so the data we can collect also changes because people are willing to delegate more work.

This is the software engineering flywheel: people keep using it, trusting it with more complex work, and providing traces that help models improve at real-world work. How do you get this flywheel in other fields? Today, this requires human effort and software teams to come in and do all of this work. But setting up the flywheel is fundamentally about creating software that makes it intuitive for people to use AI.

I extrapolate that existing frontier language model capabilities are severely underadopted because of the high friction to figure out how to use and integrate a new platform such as ChatGPT or Codex into a person’s daily work. I’d go so far as to argue that currently, AI diffusion across traditional software is counterintuitive, harmful, and alienates the end user; even people in tech don’t like Gemini AI overviews replacing Google’s Dictionary boxes that would often show up in search results.

Software has to be the pathway to widespread adoption. I think people sometimes believe that an arbitrarily powerful AI will show up in the workplace and immediately cause ripple effects around it as it displaces people who were previously working on certain tasks. But the real world is much more complicated – there are messy political and internal dynamics everywhere. Organizations and people also tend to resist change: we have a status quo bias, and our willingness to be flexible depends on how much value is created and how easy a tool is to adopt. Fortunately, software as a medium is something we’re all familiar with, because it exists to simplify friction when done right.

What does this endgame look like?

The beauty of this end stage is that it can look like whatever the end user wants it to look like. If a user prefers a style of interaction where everything is in messages such as Poke or Instinct, that is readily available to them. If the user prefers a style of interaction similar to a voice assistant, that is readily available to them. And for complex tasks where UIs are still preferred, that is readily available to them.

We have never created truly accessible personal software for people because it would simply require more effort, money, or technical skill than people could justify; it would effectively require a software team for every single individual. I’m sure you have myriad problems with the software you use in your day-to-day life. Tools like Arc Browser and Superhuman Mail both emerged because they identified a user experience that could be better, and found a loyal user base with whom this new experience resonated.

In the past, Geoffrey Litt has described how LLMs could make software more malleable (easier for its users to modify). I think the implications extend to AI adoption itself: models could increasingly perform the work required to make their capabilities useful to each person.

Self-creating software would effectively be whatever you want it to be. Say that a lawyer really doesn’t like the interface of the software that the firm uses to track billable hours; in this world, it’s incredibly easy for that lawyer to change their own personal user experience with the software, whether rearranging the components or having AI automate various parts of the information that they need to fill out. Initially, the lawyer wants to have entries grouped by client and matter rather than day by day. Later, they repeatedly correct how certain activities are categorized, and the system proposes a change that it checks against previously reviewed entries. Those successful changes persist. Over time, the lawyer spends less effort explaining preferences and correcting recurring mistakes.

This is software too cheap to meter. Already, we have seen a Cambrian explosion of software and personal projects that do everything from aggregating news and paper feeds for researchers to quality-of-life software that would have taken months and would have been impossible to justify. But this has largely remained constrained to SF-sphere, especially because creating good software still requires knowledge that remains in the hands of software engineers and designers. Net new software will be created for the express purpose of helping people with the work that they already do in their day-to-day lives. In this world, the work of software engineering is largely abstracted for the end user. A platform such as Codex may observe that the user spends a lot of time scrolling through their X bookmarks, and then propose and whip up a nice-looking tool that lets the user find exactly what they need on demand. The user will give some requests for additional features, and as the user finds more things that a specific piece of software can and should do, will request them.

How does this happen?

Today, a plethora of domain-specific harnesses exist to ease the diffusion of AI into the real world. You have Harvey and Legora for legal, or Sierra and Decagon for customer service. But these harnesses exist because the cost of diffusion is incredibly high, and a stellar company of a hundred or so scrappy engineers can out-execute a team of half a dozen or so people at Anthropic or OpenAI that make plugins for legal or customer service.

However, once capabilities are good enough, the harness simply becomes the domain-specific harnesses by building the tools and workflows that require dedicated product teams. There’s no reason why self-improving agents should draw the line at improving the harness and its underlying tools to purely elicit greater capabilities from the underlying model. They can also, with proper versioning and self-improvement, create better user experiences by modifying their own interfaces based on user preference and feedback. Why should you need to use the same harness and interface built for a billion other people when you have a working style that you are much better suited for? If you love using shortcuts, the harness can whip up shortcuts for you. If you’re a visual person, dashboards can turn into graphs and widgets.

Effectively, the cost to create what the end user desires falls to zero, leading to greater and greater adoption – friction costs become negligible compared to the gains that people receive in their productivity. Eventually a user might be leveraging dozens or even hundreds of custom-made harnesses for tasks that range from managing all of their context at work to one that simply handles billing disputes for them whenever they arise.

Creating the environment to end all RL environments

In many ways, this is the ultimate RL environment for the task of making something people want.

Once models are doing more and more useful work, their actions, mistakes, and user interventions become observable.

Consider this sidebar conversation about what continual learning could look like from Dwarkesh’s recent podcast episode about how close we are to RSI. In particular there are two quotes from Charlie O’Neill:

“A good example [of continual learning] is probably Composer. Harvey’s doing the same thing with legal agents. You have some sort of model, and you are getting very specific environments from the data that you have for that particular task, and things that users are complaining about, and all the feedback that you’re somehow extracting from your specific deployments.”

“I think the real world — and the reason people are thinking so much about continual learning — is not really a cumulative task. Imagine in a law firm, you have an agent acting as a legal associate. That’s a very non-stationary distribution. You have to be able to fit in your context all the relationships between all the important people at that company, which are also changing all the time. You have all these implicit ways about how things are done, where to find information, et cetera. That’s not as clean an example of a cumulative task as RSI is.”

I believe that systems capable of building useful applications could also construct training environments from their own operation – RL environment creation already largely leverages coding agents today. Models could capture the relevant inputs, preserve user signals, and propose criteria for judging whether a task was completed successfully.

Self-creating software could do two things at once:

  1. It helps us figure out lower-friction ways to do real-world tasks. In effect, we are creating the largest database of software and tooling that finds the right context to get work done. This is effectively what OpenAI envisioned with the chatbot marketplace back in 2024 or how users share skills and plugins they find effective today, except now it’s entire software ecosystems that may emerge. For what it's worth, I think the GPT Store was the right idea but long before it could be useful. A similar marketplace in 2027 might accelerate diffusion even further by having ready-to-install custom harnesses that immediately get up to speed and adapt to your work style.
  2. It lets us generate billions of environments where users continue to provide feedback, which is a scale at least two OOMs above what we have today. Today, dozens of AI data companies make billions in revenue putting AI talent into generating thousands of environments of okay quality (it is known that the RL environments today need much more quality control).

My point is that if “an inherent part of the learning there is interacting with the real world,” as Dwarkesh describes, the path to scale the data that is necessary for this will not come from an explosion of labor working on sandboxes and RL environments. It will instead come from lowering the adoption costs for models to start interacting with the real world, which is what self-creating software does.

What happens to the software companies?

Ultimately, it becomes harder and harder to justify certain software. For example, current subscription-based services that people pay to take notes on their tablet will likely disappear altogether.

Most software people currently use will be stripped down to its core functions, which I would argue is a data and security moat alongside permissioning and reliability. These still remain important, especially in enterprise settings that require the software they use to be compliant and secure.

There is a variety of software that will continue to benefit from network effects. Messaging platforms and social media (Facebook, Bloomberg Terminal, etc.) are intuitive examples; people use these platforms because other people use these platforms as well. But there is no reason why a user cannot take the data underlying all of these messaging tools and condense this into a simple app for themselves, even if they continue to rely on the underlying functions of the original apps.

There will also continue to be holdouts. Medical software may be doomed to continue running on software from the 1960s (MUMPS) for much longer. Sooner rather than later, external factors such as AI reshaping the context doctors ingest and automating administrative sprawl combined with continued capability gains in modernizing software will also mean that these systems too will become modernized. It’s also plausible that the sheer productivity gains from some countries, states, or firms adopting new software will cause cascading effects due to competition dynamics as well.

Are the labs doing this?

It’s very clear that both have their eyes on the ball.

Way back in March, we got the first rumors of OpenAI planning to create a sole super-app for ChatGPT and Codex, which was substantiated when ChatGPT and Codex combined in July. Many of the staff members that were on shuttered teams such as Sora, OpenAI’s short-form video generator, and Atlas, OpenAI’s web browser, have since shifted to Codex. With Codex Work, it seems pretty clear that a superapp is the endgame. Sam has also spoken many times about how he envisions more capable models onboarding its users to how to use it:

“If you ask ChatGPT, ‘Teach me how to use you’ it’s pretty good.” with Tyler Cowen in October 2025.

“Eventually, the models will get so good that they’ll help companies deploy themselves.” TBPN in February.

We've also known for a while that the endgame for Anthropic has been models that are really good at coding as an interface to interact with the world, hence the launch of Claude Code in May 2025.

Dario has described a similar change in the economics of software. At Davos this January, he said the following:

“There are still things for software engineers to do. It’s like, even if the software engineers are only doing 10% of it, they still have a job to do, or they can take a level up. That’s not going to last forever. The models are going to do more and more.”

“There’s an incredible amount of productivity here. Software is going to become cheap, maybe essentially free. The premise that you need to amortize the cost of building software across millions of users may start to be false… software may become very flexible and recyclable.”

To summarize, we need more experience with real tasks to help AI systems become much more effective at doing what people want. The cheat code to making the capabilities we already have useful for people is in the software effort it takes to deploy and diffuse AI. But if AI can build and maintain the software through which people use it, it can take on an increasing share of the work that’s required for its own adoption. All that is needed for this is continued improvement in building and serving software itself.

It is coming, and it is coming soon.