Back to all insights

AIBES Insights

AIBES 5-Point Friday #35

Two machine proofs landed in seven days, one of them a Millennium Prize problem, and the part that broke was not the math, it was deciding whose name goes on it.

Signal Worth Noticing

Verification scaled to machine speed. Attribution did not.

On September 4, Anthropic announced that Claude had produced the first complete computer-checked proof of Fermat's Last Theorem in Lean, working largely on its own for eleven days and generating roughly 13 million lines of code. Four days later OpenAI announced a Navier-Stokes result, one of the seven Millennium Prize Problems, reached in about 88 hours by roughly 10,000 concurrent agents and also published as a Lean proof. Neither result is disputed on correctness, because in both cases a compiler settled it. What is disputed is authorship: NYU mathematician Tristan Buckmaster says OpenAI raced him to the full problem after learning of his progress with Levent Alpöge, then pressured him over co-authorship, which OpenAI denies. A machine can now prove a theorem faster than a field can agree on who proved it.

Pierre de Fermat beside a screen showing his Last Theorem typechecking in Lean
The Build-to-Govern Ratio: a build layer of people and agents on one side, a tiered governance layer of teams aligning teams on the other

Framework We're Using

The Build-to-Govern Ratio

OpenAI pointed roughly 10,000 agents and zero steering committees at one problem for 88 hours. Most enterprise AI programs answer the same mandate by forming a team, then a second team to align the first, then a steering group to align both, and Salim Ismail's note to us this week puts the price of that at somewhere between a quarter and half of total effort inside a large group going into coordinating the group rather than doing the work. Coordination overhead rises with every person you add, without exception, and it never levels off at some ideal team size. So we now count two numbers before an engagement starts: how many people sit in the build layer, and how many sit in the governance layer. If the second number is bigger, you have already found your constraint, and it is not the model.

AIBES Tech Of The Week

Machine-checkable outputs

Both of this week's landmark results were accepted because they compiled, not because a committee read them and agreed. The pattern worth stealing is to make an agent emit its result in a form a deterministic checker can accept or reject: a passing test, a schema-valid record, a reconciling ledger entry, a query that returns the expected row, rather than prose a person has to read and believe. Human review capacity is the real ceiling on agent throughput, and a checker does not have one. Write the check first, then let the agent run at whatever speed it wants, because the four questions still have to resolve: what did AI do, who approved it, what did it cost, and what changed downstream? Run AI like you run finance.

Machine-checkable outputs: an AI agent emitting results a deterministic checker can accept or reject

Trending News

The headlines that fit the bigger pattern

  1. Shopify acquired Tailwind Labs on September 9, with Tailwind CSS staying MIT-licensed while Tailwind Plus and ui.sh close to new signups. Why it matters: A styling layer installed more than 110 million times a week, running under ChatGPT, X, Reddit, and Cloudflare, now has a commerce company as its steward.
  2. OpenAI will introduce Managed Agents at DevDay on September 29, a hosted platform with customizable environments, skills, and plugins, plus a self-hosting option. Why it matters: The agent runtime is becoming a purchased service, which moves the build-versus-buy line off the model and onto the whole execution environment.
  3. Meta launched Muse in the United States on September 8, a personal agent that runs in its own cloud virtual machine and can send email, book travel, and negotiate bills. Why it matters: A consumer agent with spend authority puts the approval problem in households before most companies have solved it internally.
  4. OpenAI released GPT-6 Astra on September 3, its first model to meet the Critical cybersecurity threshold, scoring 100% on ExploitBench and shipping off by default behind per-workspace admin enablement. Why it matters: When the release gate is an entitlement toggle, the admin console becomes a frontier-safety control.
  5. Server CPUs and DRAM became the scarce parts, with lead times stretching from one or two weeks to eight to twelve and AI data centers absorbing roughly 70% of global memory output. Why it matters: Capacity plans built around GPU availability are now gated by the ordinary hardware sitting around the GPU.
The Shopify logo
Henri Poincare at his desk surrounded by mathematics texts

Quote We're Pondering

"Science is built up of facts, as a house is built of stones; but an accumulation of facts is no more a science than a heap of stones is a house."
  • Henri Poincaré, the French mathematician who, in 1902, warned that gathering results and understanding them are not the same act, and whose own conjecture became one of the seven Millennium Prize Problems.

Related insights

More weekly commentary and practical AI updates from the AIBES team.

AIBES 5-Point Friday #34

Nemotron beat the best human coder alive at the world's hardest programming contest, and the ceiling this week didn't rise gradually, it moved.

Read more

AIBES 5-Point Friday #33

OpenAI's own silicon beat its supplier's on inference, Nvidia moved to buy the place open models live, and the real ceiling on AI throughput turned out to be a signature.

Read more

AIBES 5-Point Friday #32

The payments layer bought the routing layer, a frontier lab stopped its own biggest training run, and the leverage keeps moving to the boring middle of the stack.

Read more

Ready to take your business to the next level?

Get in touch with AIBES today: discovery, roadmap, and a clear ROI path.

Get in touch →