| |
Since last week: OpenAI shipped GPT-6 Astra after delaying parts of its release, and graded it the first model to reach the "Critical" cybersecurity level on OpenAI's own scale. Anthropic shipped the same underlying model twice, at the same published base price, separated by safeguards and by who is allowed the less restricted version. And the independent group that tests Astra's reasoning published two very different scores for it. Last week we said your vendor holds switches you do not. This week two of them turned out to be the test conditions and the safeguards. |
The week in three numbers: 54.8% → 99.9%, Astra's score on the same benchmark at the same reasoning setting, when ARC Prize switched from its own provider-neutral setup to OpenAI's · 20.1%, Salesforce's full-year margin guidance, below its own 20.5% quarter, with AI token spend named as part of the reason · $600,000 in model credits burned on a stolen API key at METR before anyone noticed, over about three weeks.
|
In this issue
01 · The One Thing · 02 · Do This Week · 03 · The Signal · Skim ends here.
04 · The Margin-Proof Tracker · 05 · The Synthesis · 06 · Where the Minds Disagree
Then: What We're Watching · Worth Your Time
|
|
01 · The One Thing
Change the setup and the score moves 45 points. The announcement shows you one of them.
OpenAI says Astra "saturates ARC-AGI-3 with a 99.9% score." ARC Prize, which runs it, published two results at that setting. Astra scored 99.9% on OpenAI's own setup, where the model keeps its private working state between turns. It scored 54.8% on ARC Prize's shared setup, which every model gets and which makes it write down what it wants to remember. ARC Prize calls Astra a real step forward. That wrapper of prompts, tools and memory is called a harness. OpenAI's footnote does name the one it used. The provider-neutral score is nowhere in the announcement. (ARC Prize, Sep 3 · OpenAI, Sep 3)
|
| |
The executive shift: the gap is large at every setting, and peaks at more than 80 points at low reasoning effort, where it is 17.5% against 98.0%. The same week, Anthropic shipped Fable 5.1 and Mythos 5.1 as the same underlying model at the same published base price, separated by safeguards. So in one week, a score turned out to depend on the test setup, and what a product could actually do turned out to depend on the guardrails. The model is no longer the whole product. The system around it, memory, tools, permissions and safeguards, is becoming part of the capability, and your supplier builds that system. "Which model?" is on its way to being the wrong procurement question. (Anthropic) |
1 · Make every benchmark number in a vendor deck name its test conditions. Ask three things in writing: who ran the test, what setup the model was given, and whether the rival it beat was tested the same way. Stakes: this week the same model at the same effort scored 54.8% or 99.9% depending on the setup. ARC Prize is now labelling both conditions on its leaderboard. We found no other major benchmark that publishes both. A score without its conditions is not decision-grade evidence. ( ARC Prize)
2 · Confirm in writing that your zero data retention terms still apply. Zero data retention means the vendor keeps no copy of what you send or what it returns. Anthropic changed the default. Enterprises no longer get it automatically. They must apply, and thirty days of retention is now standard for commercial API customers unless something else was agreed. Stakes: a retention clause you signed may not have carried over. Ask for confirmation of your current setting, in writing, with a date. ( The Register, Sep 2)
3 · Find out what standing access your AI agents inherit. DoorDash moved its coding agents off developer laptops for one reason. On a laptop an agent runs with the same credentials as the developer, and nobody can see where it runs or what it touches. Stakes: the pattern is well documented, across incidents catalogued this week that happened over several months. In one, an agent misread a variable and deleted a founder's home directory. Ask your engineering lead what your agents can reach today, and where that is written down. If the answer is "the same things the developer can reach," you found the gap. ( InfoQ, Aug 31)
This week's question: when the supplier builds the model, can shape the setup it is measured through, and sets the guardrails, what is left for a buyer to check?
Horizon Same model, same base price, different safeguards, different door. Primary: Anthropic's own product pages. Anthropic states that "Claude Fable 5.1 is the same underlying model as Claude Mythos 5.1 with safeguards for cybersecurity and biology." Both are priced at $10 per million input tokens and $50 per million output tokens. Fable is generally available. Mythos, the version with fewer safeguards, goes only to vetted organisations through an application. Even OpenAI describes it this way in its own benchmark footnotes, calling Mythos "Fable with fewer safeguards." The safeguards are not cosmetic: Anthropic notes that where Fable's protections intervened on one benchmark, it scored zero. Analysis: the underlying intelligence and the published base price are the same. What differs is usable capability, because Anthropic changes the safeguards, and who may have the less restricted version. Policy has become part of the product. ( Anthropic Mythos · Anthropic Fable)
Proof A named officer told investors AI spend is why margin guidance did not rise. Primary: an investor conference, on the record. Salesforce's second-quarter operating margin was 20.5%. Full-year guidance is 20.1%. Deputy CFO Mike Spencer called that gap "part of the reason we didn't raise margin guidance." He named the fix too. Pick the model to fit the task, rather than defaulting to the newest one every time. The cleanest disclosed dollar effect of the week, and it points down. ( The Register, Sep 3)
Trust An outside evaluator got breached, and did not notice for weeks. Reported; METR on the record. An attacker took an API key off a researcher's personal, publicly exposed instance in March. About $600,000 in credits went over roughly three weeks before anyone saw it. METR says no sensitive data was touched. It is one of the organisations frontier labs use for outside evaluation, and it worked with OpenAI on the Hugging Face investigation. Analysis: an organisation trusted to scrutinise frontier-model behaviour missed abuse of one of its own credentials for three weeks. ( The Register, Sep 1)
Deploy One company published what agent governance requires. Three providers went down the same day. Vendor's own account on the first, reported on the second. DoorDash says its platform automated 130,000 engineering tasks in a single month and supports more than 25,000 code reviews a week. It gives no error rate and no cost comparison against the laptop setup it replaced, so volume says nothing about quality. What it does publish is the architecture: isolated machines, one gateway, scoped permissions, logs. On September 3, ChatGPT, Claude and Grok all had outages. One compute partner apologised, which explains one of the three. Analysis: using three vendors only spreads your risk if the three can actually fail separately. That is a question about their infrastructure, not their logos. ( InfoQ, Aug 31 · The Register, Sep 3)
Work A maintainer shipped 180,000 lines he says he cannot review. Primary: the developer, in his own words. Rick Brewster has worked on Paint.NET for over twenty years. He shipped a new graphics layer that Claude wrote from his direction. His words: "I cannot possibly review 180,000 lines of code." He checks the architecture and fixes mistakes. He caught the model mishandling one memory detail. Analysis: the question is not whether an agent out-produces a reviewer. It is what "review" means once it cannot mean reading the work. ( Simon Willison, Sep 2)
▼ Below the Cut
The one dataset that actually counts AI-attributed job cuts shows AI losing the top spot for the first time in five months. Primary: Challenger, Gray and BLS. Challenger asks employers why they are cutting. In August, Restructuring led with 16,173 cuts, 31% of the month's total. Artificial Intelligence fell to fourth, with 3,462 cuts. That is its lowest month since December 2025, and it "ends a five-month run, beginning in March, in which AI was the leading monthly reason." AI still leads for the year, cited in 116,175 announcements, about 22% of all cuts. Separately, US productivity rose 1.4% in the second quarter and unit labour costs rose 1.2%, and the BLS release attributes none of it to AI. Analysis: the figure that gets quoted is the year-to-date 22%. The figure that moved is the monthly one, and it moved down. Conflict label: Challenger sells outplacement services and publishes the tracker. ( Challenger, Gray, Sep 2 · BLS, Sep 3)
One question for your team this week: for every AI number in our plan, do we know what was running underneath it when it was measured?
Still running from earlier issues: the NIST comment deadline on October 15 and the CFTC comment deadline on October 20.
| End of skim · deep read begins |
| 04The Margin-Proof Tracker |
| |
No row advanced. N = 13, unchanged. For the second week running, nothing we read tied AI to a company's reported earnings in a filed document. What we got instead is the reverse: a named officer explaining, at an investor conference, why AI spend kept margin guidance down. |
What companies claim AI is worth, against what shows up in their financial statements. None of the thirteen has reached Stage 4. The evidence ladder: 0 · Narrative (a story, no numbers) · 1 · Operational (activity counted) · 2 · Financially linked (a number tied to AI, mixed with other causes) · 3 · P&L-attributed (a reported profit or margin change the company credits to AI, in a filed document) · 4 · Sustained (Stage 3 held four quarters).
A definition made explicit this week. Stage 3 has always meant a filing, and this is the first week that mattered, so it is now written into the ladder above rather than left implied. See the candidate below for why.
| Company |
Evidence |
Grade |
Next test |
| Duolingo |
EvidenceGross-margin rise "reflecting continued reductions in per-unit third-party AI costs." 10-Q |
Grade3 |
Next testQ3, two quarters to Stage 4 |
| Nutanix |
EvidenceCEO says the company spent $20M on its own AI cluster because "usage exploded and so did costs," and is "no longer paying on a per-token basis." Payback in a year is his forecast, in a results briefing. The Register, Aug 27 |
Grade2, a forecast |
Next testNext filing: any of the $20M in writing |
| Klarna |
EvidenceOpex +16% against +27% revenue, "supported by AI-enabled productivity gains and continued cost discipline." Unquantified, credit shared. 6-K, Aug 18 |
Grade2 |
Next testQ3, does a figure ever attach? |
| IBM |
EvidenceAI signings; company says signings are not revenue. Q2 |
Grade2 |
Next testQ3, bookings or revenue? |
| Latch / DOOR |
Evidence~65 roles, $10 to 12M expected, not booked. 8-K |
Grade2 |
Next testQ4, does it get booked? |
| Infosys |
Evidence8.2% of revenue labeled "AI," alongside cut guidance. Q1 FY27 |
Grade2 |
Next testQ2 FY27, share up and guidance up? |
| Equifax |
Evidence$150M AI cost-reduction goal. Q2 |
Grade2, a target |
Next testQ3, booked or restated |
| ServiceNow |
EvidenceAI contract value past $1B. Committed, not earned. Q2 |
Grade2 |
Next testQ3, recognized in results |
| Alphabet |
EvidenceCloud +82% to $24.8B; AI credited, not separated. Q2 |
Grade2 |
Next testQ3, is AI revenue separated? |
| Visa |
Evidence$563M severance; AI's share never stated. 8-K |
GradeProvisional |
Next testQ4, capex and hiring mix |
| Bank of America |
EvidenceEfficiency ratio 59%; no stated link to AI. Q2 |
Grade1 |
Next testQ3, linked in writing? |
| JPMorgan |
EvidenceAI-linked headcount cut. No primary document found, and that absence is the row. Q2 |
Grade1 |
Next testQ3, any written attribution |
| Etsy |
Evidence220 roles, ~$35M. The AI denial is in the staff memo, not the 8-K. 8-K |
GradeNot scoreable |
Next testQ3, does product-dev spend rebuild? |
| |
Candidate watchlist, outside the table and not counted in N. Salesforce. A deputy CFO naming token spend as "part of the reason we didn't raise margin guidance" is a margin statement credited to AI, which reads like Stage 3. The evidence is a conference remark reported by a third party, not a filing, so it waits. It joins the table if the attribution reaches the 10-Q. (The Register, Sep 3) |
The next tests, ranked by what they would settle. Salesforce's next 10-Q would be the table's first negative attribution: a company stating in a filing that AI spend held margin down. Duolingo already puts AI in a filing, on the other side of the ledger. Duolingo's Q3 stays the only route to a first Stage 4. Nutanix's next filing keeps its low bar: any of the $20 million, in writing, from the company rather than an interview. Rows carry the evidence verified when each last moved. None was re-verified this cycle.
You are buying a system, and the seller shapes the conditions it is tested under.
Start with the mechanism. ARC Prize wrote and ran the benchmark, and tested Astra under both of the setups described at the top of this issue. The 45-point gap is the distance between those two setups, not between two models.
Both figures above are at high reasoning effort. ARC Prize's own summary quotes the best result under each setup, 62.7% and 99.9%, but those sit at different effort levels. We matched the setting, because comparing across settings is the exact error this issue is about.
Note who owns what. ARC Prize owns the test. OpenAI supplied the setup that produced the higher number. The seller did not write the exam. It supplied the equipment the exam was taken with, and that turned out to be worth 45 points.
Now the connecting step. Our read: ARC Prize is not accusing anyone, and neither are we. It says outright that these are two questions rather than one wrong answer. The shared setup gives "a consistent, apples-to-apples comparison across providers." The other asks how a model does "when it can use the context-management features its provider designed for it." Both are fair. They are also different products. One measures a model. The other measures a model plus the vendor's plumbing. A buyer who does not know which one is quoted cannot tell what is being sold.
The same shape appeared twice more this week. Anthropic says Fable and Mythos are the same underlying model at the same published base price. The same intelligence is turned into different usable capability through safeguards, and Anthropic decides who may have the less restricted version. And OpenAI's Critical cybersecurity designation was made against OpenAI's own framework, with OpenAI deciding the threshold was crossed.
Put those together and the useful conclusion is not that benchmarks mislead. It is bigger. The model is becoming an incomplete unit of analysis. Enterprises do not buy the model on its own. They buy an operating system around intelligence: memory, context handling, tools, permissions, routing, safeguards and human controls. Those are not implementation details. This week they moved capability, cost, speed, risk and the reported score. As raw capability gets cheaper, control over that surrounding system can become a separate source of supplier power, but only where a buyer cannot get it any other way. On that test this week's three cases are not equal. The benchmark gap is the weakest of them, because ARC Prize closed it within a week by committing to publish both conditions. Mythos is the strongest, because the less restricted version is granted rather than sold. A price can be budgeted against. Conditions and permission cannot be, in the same way.
The counter-case, and it is strong
The plumbing is not cheating. Preserving reasoning between turns is a real engineering achievement and part of what a customer actually buys. ARC Prize's own numbers show those runs were about 3.66 times faster and used 49% fewer tokens on the games both setups solved. If you will run Astra through OpenAI's stack in production, the number measured through OpenAI's stack is arguably the one that predicts your experience. And OpenAI did disclose the harness in a footnote, and states that its headline scores are the maximum at any effort.
There is also a genuinely counterintuitive finding in the same table. On this benchmark, higher reasoning effort cost less, not more. At maximum effort the shared-setup run cost $26,098; with reasoning off it cost $49,791. ARC Prize explains why: the model solves games in fewer moves, so there are fewer calls. This is one interactive benchmark and does not establish a general rule. But it does establish this: "more reasoning means more total cost" is not a safe budgeting assumption. On agentic work, more thinking can cut the number of actions enough to lower the bill.
And on safety, one correction to the easy story. Outside evaluators did test Astra. OpenAI's system card, the technical report a lab publishes alongside a model, carries external cyber evaluations by Irregular, plus work from UK AISI and Apollo Research. The honest claim is narrower and still sharp.
The reusable test
For every AI number and every AI capability your plan depends on, answer three questions in writing.
1. What was running underneath? Who ran the test, with what setup, and was the rival measured the same way?
2. Priced or granted? Can you buy it, or must you be approved?
3. What is your notice period? If the grant is withdrawn, how many days do you get, and what runs instead?
A number without its conditions is a claim. A capability that is granted rather than sold is a dependency.
| 06Where the Minds Disagree |
| |
Which score is the real one?
ARC Prize's position, in its own post: both. Its heading is "Two Harnesses, Two Questions." It has committed to publishing both from now on, "with each evaluation condition clearly labeled."
The buyer's position: the higher number is only available to somebody running on that vendor's stack, in that vendor's configuration. That makes it real and vendor-dependent at the same time.
The comparison shopper's position: only the shared setup lets you rank two vendors against each other. Once each model is measured through its own maker's plumbing, cross-vendor comparison stops meaning much.
Our read: ARC Prize is behaving well and is fixing the ambiguity itself. The problem is not the test. It is that the announcement surfaces 99.9% and the provider-neutral number appears nowhere in it. A reader of the announcement alone has no way to learn a 45-point gap exists. Corroborating evidence arrived the same week, in a different field. A paper on evaluation method found that "the apparent winner changes" depending on when you measure and how much capacity the losing method was given. So this is a general problem, not a one-off. (arXiv 2609.03900)
|
| |
Who decides a model is dangerous?
OpenAI's position, in its own documents: the grade is real and applied against a written framework. It says it "delayed parts of Astra's development and release" while strengthening protections. Our read of that: a delay is a cost the company chose to bear.
The independent evaluator's position: outside testing happened and is published. OpenAI's system card carries external cyber evaluations by Irregular, plus work from UK AISI and Apollo Research. Booz Allen's Cyber Weapon Index tested 18 US and Chinese models under identical conditions, but names Astra only to say it was not among them, because it had not been released when the testing ran. (The Register, Sep 2)
The critic's position, reported this week: researchers have panned the 22-page framework itself. Their argument is that it "allows OpenAI's CEO to deploy even more dangerous capabilities, especially if other AI developers do so." (The Register, Sep 3)
Our read: the lab does not hold the only test, and saying otherwise would be wrong. It holds something more durable: the rulebook. It defines the threshold, judges whether its model crossed it, picks the safeguards that follow, and decides who gets the unrestricted version. Outside testing constrains the claim. It does not constrain the framework the claim is made under.
|
September 10, 2026 · The SEC's Investor Advisory Committee holds a panel on AI and public-company disclosure. Our read: this is not about whether firms must disclose that they use AI. The published agenda asks how AI is changing the way disclosure is "produced, reviewed, filed, disseminated, and used," and names the open questions: factual accuracy, hallucinations, source traceability, automation bias, and responsibility under disclosure controls. That is this week's lead applied to securities filings. What would change our read: whether the committee recommends new expectations for disclosure controls, source traceability, or who is responsible when AI helped prepare a filing. ( SEC agenda)
October 1, 2026 · The FCC's Technological Advisory Council meets, tasked with "studying advanced spectrum sharing techniques, including the implementation of artificial intelligence and machine learning to improve the utilization and administration of spectrum." This is advisory input, not a proposed rule. What would change our read: whether a rulemaking follows. ( FCC public notice, Jul 28)
Two more we could not file. The EU State of the Union on September 16 and an FDA advisory panel on September 23 both carry confirmed dates in their own text. Our record flags both as not citable this cycle. We name them as unresolved rather than run them.
MIT Technology Review, "Scaling agentic AI pilots across the enterprise" (Sep 3). An operations executive names the failure mode plainly: laying agents onto a workflow that was already inefficient, instead of redesigning it first. Read it for the diagnosis. Treat the claim that some 80% of the Fortune 500 have adopted agentic AI with care: the piece gives no method and never defines what adoption means. technologyreview.com
OpenAI, "Path to Astra" (Sep 1). The company's own account of why it held the model back. It says it "delayed parts of Astra's development and release" to strengthen protections, and that its safeguards at the time would have prevented the Hugging Face breach. Read it for how a lab describes its own caution. openai.com
The Register on Nvidia buying Hugging Face for $12.9 billion (Sep 3), expected to close in the first half of 2027. If you pull open models through Hugging Face, the company selling you compute would also own the distribution point. theregister.com
How we label evidence: Primary source · Corroborated · Reported · Vendor claim · Analysis. Written and edited by Mario Suarez · Independent analysis · Every number in this issue traces to a source that was opened and quoted by the desk that filed it. Where a desk could not open something, this issue says so rather than implying coverage.
|