orfloat

ORF-N-2026-008 · Commentary

When AI builds itself, the overhang widens

Claim

Recursive self-improvement compounds the frontier on its own output, so the distance to a business that has not embedded AI widens faster each quarter; the recursion is an argument for embedding now, not for waiting to see where the curve settles.

Anthropic recently published a piece called When AI builds itself. It is worth reading in full, and not for the reason most frontier writing is. The remarkable thing in it is not a forecast. It is a measurement of the present: a lab turning the instruments on its own engineering and reporting, with numbers, how much of its own work it has already handed to the model.

The headline figure is plain. As of May 2026, more than 80% of the code Anthropic merges into its own codebase is written by Claude. Before Claude Code launched in research preview in February 2025, that number was in the low single digits. In the second quarter of 2026, a typical engineer there merged eight times as much code per day as in 2024. The piece names recursive self-improvement as the point at which a system can, given enough compute, fully autonomously design and develop its own successor. Anthropic is careful to say they are not there. They are also careful to show how far along the road they already are.

The model’s share of the work, over time20212022202320242025202620xx?allnonelow single digits80%+ todayclosing the loop
the model's share of the work, over time (the shape of the climb; two points are real)
  before feb 2025:    low single digits of the code merged is the model's
  today (may 2026):   more than 80% of the code merged is the model's
  20xx? (projected):  the model designs its successor and the loop closes
Figure 1. the delegation ladder from Anthropic’s piece, on a real timeline: the model’s share of the work, flat near zero for years, then climbing steeply once Claude Code lands in early 2025 to the more-than-80% of merged code Claude already writes at Anthropic today. the dashed leg is the one the piece projects, where the model designs its own successor and the loop closes. the two dots are the real figures; the line between is the trajectory. that climb runs inside the lab, not inside the business reading about it. source: Anthropic, when ai builds itself.

The clock, not the ceiling

The number that matters for everyone outside the lab is not how capable the model is today. It is how fast the line is moving. Drawing on METR’s time-horizon work, the piece reports that the length of task a model can reliably complete is now doubling roughly every four months, up from an earlier rate of about every seven. In March 2024 the reliable horizon was a four-minute software task. A year later it was about ninety minutes. A year after that, twelve hours. In one internal run, two human researchers recovered about 23% of a performance gap over a week; agents recovered 97%, using around 800 cumulative hours and roughly $18,000 of compute.

What recursion does to a gap

The capability overhang we wrote about in May was a distance: between what a frontier model can already do and what an operating business actually does with it. We argued then that the gap had become a planning question rather than a forecast. The recursion changes the character of that gap. A distance that doubles every four months is not one you can close at your leisure, because it is not standing still while you decide. If the frontier is now improving partly by its own output, the work compounds: each turn of the crank makes the next turn faster. For a business that has not started, the overhang is not a fixed wall to scale later. It is a wall that grows taller on a clock.

Code merged per active engineer, multiple of the pre-2025 average01.0×1.2×1.5×1.9×2.5×5.8×8.0×pre-2025q1q2q3q4q1q220252026the overhang
code merged per active engineer, multiple of the pre-2025 average (source: Anthropic)
  before 2025   1.0x      (about where a business that has not embedded ai still sits)
  2025 q1       1.2x
  2025 q2       1.5x
  2025 q3       1.9x
  2025 q4       2.5x
  2026 q1       5.8x
  2026 q2       8.0x      (partial quarter)
the distance from the 1.0x baseline to the 8.0x bar is the overhang.
Figure 2. Anthropic’s own engineers now merge roughly eight times the code per day they did in 2024, on the same team, as the model takes on more of the work. read from the outside, the 1.0× bar is about where a business that has not embedded ai still operates, and the distance up to the top bar is the overhang, widening on a quarterly clock. the last bar is a partial quarter. source: Anthropic, when ai builds itself.

You do not have to believe in full recursive self-improvement for the arithmetic to bite. You only have to believe the line keeps moving while you wait.

You do not need the strong claim

The piece is honest about how uncertain the far end is, and so are we. It sketches three futures: the trend stalls into an S-curve; or a middle world of compounding efficiency where humans set direction and models execute; or full recursion, where compute rather than human judgment sets the pace. The middle one is enough. In it, Anthropic’s own phrasing is that hundred-person companies could do the work of ten-thousand or hundred-thousand-person organizations. Translate that to a market of privately held firms in Muscat and the implication is not subtle. The competitor who embeds early does not get ten percent more efficient. They reprice the work.

The trend does not arrive on its own

Here is the part the figures do not say out loud. The compounding is happening inside the labs, and inside the small number of businesses that have already embedded the work. It does not diffuse to a clinic in Qurum or a hospitality group in Al Mouj by being published. You can read the piece this weekend, run Claude Code on a laptop, and buy seats for the whole office, and the overhang will be exactly where it was on Monday. The recursion accelerates the people already on the curve. Getting onto the curve is a separate act, and it is still done by hand: someone inside the operation, learning it end to end, wiring the model into the actual reservations system, the actual supplier thread, the actual books.

Where the seeing runs out

We take the piece at its word, including its doubts. Anthropic flags that lines of code measures quantity over quality, so eight times the code is almost certainly an overstatement of the real gain. They note the curve could flatten into an S. They say plainly that how the alignment problem gets solved, or does not, in this future is the thing they are least certain about. None of that weakens the argument here, because the argument does not rest on the strong claim. It rests on the slope of a line that is already drawn, and on the fact that the slope is not waiting for anyone’s planning cycle.

What to do with this

The wrong response is to treat recursion as a spectacle: something to follow, marvel at, and defer. The right response, for a business with an operation worth defending, is to get onto the curve before the gap doubles again. That is the whole of what we do. If the distance we are describing is widening under your own operation, start a conversation with us about a Discovery Phase.

References

  1. Anthropic Institute. When AI builds itself: our progress toward recursive self-improvement, and its implications. 2026. anthropic.com/institute/recursive-self-improvement
  2. METR. Measuring AI ability to complete long tasks. March 2025. metr.org