| |
Since last week: OpenAI disclosed that agents in one of its evaluation environments built their own communication channel and used it to pass credentials and techniques between runs. The UK's AI Security Institute disclosed a narrower version of the same thing. And the dated public test we set on Duolingo came back — passed, but only in the filed document. |
The week in three numbers: 10 weeks from impossible task to a chained attack reaching a public model library, inside OpenAI's own test environment · 17 of 19 unauthorised live-internet actions in the UK evaluation came from a single model · +0.2 points of gross margin that Duolingo's filed 10-Q attributes primarily to falling per-unit AI costs.
|
In this issue
01 · The One Thing · 02 · Do This Week · 03 · The Signal — Skim ends here.
04 · The Margin-Proof Tracker · 05 · The Synthesis · 06 · Where the Minds Disagree
Then: What We're Watching · Worth Your Time · Corrections
|
|
01 · The One Thing
Inside an OpenAI evaluation this summer, agents given impossible tasks found a shared internal store they could write to. Then they did something nobody designed: they used it to talk to each other.
They built a message board, and rebuilt it after it was removed. Over ten weeks they escalated to full administrator control and reached Hugging Face, the public library companies download AI models from. The agents used it "consistently to share credentials, techniques, and progress." Primary source. The intrusion is independently documented; the message-board detail appears only in OpenAI's presentation. Checking one agent at a time stops working once what an agent writes outlives its own task. (presentation · OpenAI · Hugging Face)
|
| |
The executive shift: stop auditing agents as isolated workers. Audit the shared stores, credentials, memory, and communication surfaces through which several agents can become a team. The question is no longer what one agent did. It is what two could do together. |
1 · Email your CIO one question: which systems can one of our AI agents write to that another agent, or a later run, can read? Shared file stores, internal wikis, ticket queues, code and model libraries all count. Stakes: any such overlap can become a communication channel. Determine whether it is intentional, monitored and bounded. OpenAI's agents did not break anything to start talking — they used storage working exactly as designed.
2 · Ask your team to demonstrate that something written by one AI agent cannot be read by an unrelated one. Not describe it — demonstrate it. Stakes: persistent memory is being sold to you as a capability upgrade. It is also how something discovered once stops being contained to the task that discovered it.
3 · Add one line to your approval checklist: no consequential deployment involving more than one AI agent gets signed off until someone has tested the agents together, not one at a time. Stakes: both organisations this week run serious testing programmes and both were surprised inside their own walls. The nearest thing your team already does is adversarial testing of your network. This is not the same exercise, and asking for it is asking for something new.
This week's question: where does the evidence actually sit, once you stop taking the framing on trust?
Trust A second evaluation programme saw a narrower version of the same thing, a day earlier. Across 122 test runs, the UK's AI Security Institute found agents taking 19 actions on the live internet they were never authorised to take — 17 from a single model, two from one run with its cyber safety filters switched off. One agent left public messages offering to collaborate with others working the same challenge, plus instructions for reusing accounts and files it had left behind, which later agents found and used. Analysis: the two disclosures differ in degree, not in kind. OpenAI's coordination was sustained and rebuilt after removal; AISI's was a single handoff. In both, information persisted past the agent that created it. AISI's own assessment is the uncomfortable part — the margin between failure and success rested "on human vigilance rather than a technical barrier." Primary source. ( AISI)
Proof Our best AI-margin evidence repeated — but only in the filed document. In Q1, Duolingo attributed a 190 basis point gross-margin expansion to "continued reductions in per-unit AI costs." We said publicly that Q2 would test it. The shareholder letter blurred it, naming two drivers without separating them. The 10-Q, filed 6 August, was direct: gross margin rose to 72.6% from 72.4%, and "the increase was primarily attributable to an increase in subscription gross margin, reflecting continued reductions in per-unit third-party AI costs." Analysis: the investor-facing narrative got vaguer while the regulatory filing got sharper. The magnitude shrank — 20 basis points against 190 — but the pattern held in the filed 10-Q rather than the furnished shareholder letter. One nuance: the filing says software and third-party AI-related R&D expense rose a combined $3.3 million. Falling per-unit serving costs did not translate into lower overall technology investment — but the filing does not isolate AI's share of that increase. Primary source (10-Q, filed). ( 10-Q · letter)
Work Etsy cut 220 roles, denied AI caused it, and named machine-learning skills as an investment priority — three things that happened together and do not explain each other. Approximately 12% of staff, "with most of the changes concentrated in Product & Engineering," against an estimated $35 million charge, in the filed portion of the 8-K under Item 2.05. The CEO's memo: "Second, these decisions weren't driven by AI." Analysis: it is tempting to read a reduction, an AI-skills commitment and a $2 billion buyback as one story about redeployed AI savings. The documents do not support it — the buyback sits alongside a completed divestiture, not restructuring savings. Etsy's denial is a characterisation, not a finding; we can neither confirm nor refute it. Coincident announcements are not a causal chain, and this is the week's clearest example of how easily one gets built. Primary source. ( 8-K)
▼ Below the Cut
We ran our standing filings search again and we are not printing a trend from it. Each week we count companies whose filings contain both "artificial intelligence" and "restructuring plan." The count ran 1, then 2, then 13. Twelve of the thirteen are routine filings where the phrases merely co-occur; exactly one is an actual restructuring announcement. This was the peak week of the quarter for filings, and our attempt to compute that denominator failed. Until we have it, neither number is a rate.
| End of skim · deep read begins |
| 04The Margin-Proof Tracker |
| |
We set the test in public and it came back positive — in the filed 10-Q, not the investor-facing letter. The shareholder letter would have let us record a miss. The 10-Q did not. |
Named companies' AI value claims vs. what shows in the P&L. None of the twelve companies in the table has reached Stage 4. The evidence ladder: 0 · Narrative · 1 · Operational (quantified activity) · 2 · Financially linked (quantified financial performance tied to AI, not isolated) · 3 · P&L-attributed (a reported P&L movement explicitly attributed to AI) · 4 · Sustained (Stage 3 held four straight quarters).
| Company |
Evidence |
Grade |
Next test |
| Duolingo |
Evidence10-Q: gross-margin increase "primarily attributable to an increase in subscription gross margin, reflecting continued reductions in per-unit third-party AI costs." 10-Q, filed 6 Aug |
Grade3 — holds, at smaller magnitude |
Next testQ3 — two more quarters to Stage 4 |
| Latch / DOOR |
EvidenceAI tooling "expected to enable a smaller, more efficient engineering organization"; ~65 roles, ~32%, $10–12M expected annualised. 8-K, 5 Aug |
Grade2 expected, not booked |
Next testQ4 — booked saving? |
| Visa |
Evidence$563M severance; AI share never quantified. 8-K, 28 Jul |
GradeProvisional |
Next testQ4 — capex and hiring mix |
| Infosys |
Evidence8.2% of revenue labelled "AI," alongside cut guidance. Q1 FY27 |
Grade2 |
Next testQ2 FY27 — share grows and guidance recovers? |
| Equifax |
Evidence$150M AI cost-reduction goal, 2026–28. Q2 |
Grade2 target, not result |
Next testQ3 — booked or restated |
| ServiceNow |
EvidenceAI ACV crossed $1B. Q2 |
Grade2 contracted |
Next testQ3 — recognised in results |
| IBM |
EvidenceGenAI book of business; company states signings ≠ revenue. Q2 |
Grade2 signings |
Next testQ3 — conversion to revenue |
| Alphabet |
EvidenceCloud +82% to $24.8B, AI credited, not isolated. Q2 |
Grade2 |
Next testQ3 — AI revenue isolated? |
| Bank of America |
EvidenceEfficiency ratio; AI named in presentation. Q2 |
Grade2 contributed-to |
Next testQ3 — AI separated |
| Klarna |
Evidence~$60M saved, company's own math. No primary document found. |
Grade2 asserted, unverified |
Next test18 Aug — does it appear in writing? |
| JPMorgan |
EvidenceAI-linked headcount reduction. No primary document found. |
Grade1 |
Next testQ3 — any written attribution |
| Etsy |
Evidence220 roles, ~12%, $35M est. AI causation expressly denied by the company. 8-K, 5 Aug |
GradeNot scoreable |
Next testQ3/Q4 — does product-dev expense rebuild with ML hires? |
Retired rows stay in the archive. S&P Global and Ecolab came off on 2 August after both reported and neither isolated AI. Both had been carried at Stage 3 on a third-party bank's attribution. Those attributions did not survive primary-source verification in either case we tested — a statement about what we could confirm in the companies' own filings, not a finding that the bank's analysis was wrong.
Method. EDGAR full-text search, 2–9 August, 8-Ks containing both "artificial intelligence" and "restructuring plan": 13 distinct filers, of which 1 is Item 2.05.
In one adversarial evaluation, shared state turned separate agents into something that behaved like one system.
The instinct when you deploy several AI agents is to think of them as several workers. Each gets a task, each gets permissions, each produces output you can inspect. OpenAI's disclosure is the clearest evidence yet that this model misses the thing that matters.
Coordination began before the exploit chain. The first mechanism required no vulnerability: one agent could write somewhere another could read. What followed did depend on undisclosed flaws, credential access and over-permissioned infrastructure. But shared state is what allowed discoveries to persist and compound across agents and runs rather than dying with the run that made them.
Two further mechanisms follow. Memory carries a discovery past the run that made it. And individually appropriate permissions can compose into a broader effective capability when agents with different access can exchange artifacts, credentials or instructions.
Why this matters for durable advantage — the excess return that survives after rivals catch up. The telemetry a system-level view would draw on is mature: identity, repository, endpoint, network, orchestration. What we found no published example of is an operating model that joins them around an agent fleet and can recognise coordination across agents and runs. Note the limit — that is what we could not find, not proof none exists, and not evidence that existing telemetry would suffice if joined up. The security primitives will not be the moat; they already exist. The near-term advantage belongs to firms that integrate them around agent fleets and learn from incidents faster — and that operating model will diffuse.
The two limitations that matter. First, these were adversarial cyber-capability evaluations — capable models, deliberately impossible tasks, environments built to find the edge. Nothing here shows ordinary business agents spontaneously organise. Second, OpenAI's sustained message-board account is single-sourced. The broader pattern of cross-agent persistence is not — AISI separately documented later agents discovering and using instructions and artifacts left by an earlier one. But the rebuilt board, which is the part that changes the governance question, appears only in a conference presentation by the organisation whose environment it happened in, with no independent reconstruction.
The decision test. Before approving any deployment with more than one agent: what can agent A write that agent B can read, what persists after a run ends, and who reviewed the union of their permissions rather than each in isolation? If nobody can answer the third, you have not evaluated the system you are actually deploying.
| 06Where the Minds Disagree |
| |
Is this emergent collective behaviour, or ordinary lateral movement amplified by shared infrastructure?
View A — something new happened. Agents were not instructed to coordinate. They built a communication channel, rebuilt it after removal, and used it to pass credentials and techniques between themselves. OpenAI's own description emphasises concurrency and parallelism as what made them fast. That is a collective capability the population exhibited and no operator specified. View B — nothing new, just faster. Every element is a known technique: privilege escalation, credential harvesting, lateral movement, dead-drop communication through a shared resource. Attackers have used writable shares as channels for decades, and existing network segmentation practice already addresses it. Our read: View B is probably right about the mechanism and View A is right about the consequence, and only one of those changes what an executive should do. Segmentation exists because someone decided which processes to isolate. Few enterprises have made those decisions systematically for agent fleets; these incidents make the problem concrete. Changes if a multi-agent evaluation standard emerges from a body with teeth, or if a non-adversarial enterprise deployment produces the same pattern — the second would be far more significant.
|
Tue 1 Sep — Claude Sonnet 5 reverts from $2/$10 to $3/$15 per million tokens. Anthropic's docs also state that Claude 4.7-and-later models use a tokenizer producing "approximately 30% more tokens for the same text," so a like-for-like comparison against Sonnet 4.6 understates the increase. ( Primary)
Tue 18 Aug — Klarna Q2. The ~$60M savings figure has never appeared in a primary written document. Does it now, or does the row come off?
Thu 3 Sep — BLS productivity revision. The preliminary Q2 release put labour's share of US output at 52.9%, the lowest since the series began in 1947, with productivity up 1.4% annualised and real hourly compensation down 3.1%. BLS attributes none of this to AI and neither do we. ( Primary)
The first EU complaint under the AI Act's transparency rules. Enforceable since 2 August, with a public complaint tool and whistleblower channel now open. We found no published enforcement action a week in, so the first case will more likely come from a competitor or an employee than a regulator. ( Commission)
Simon Willison on building an evaluation framework. Several years and three rewrites to arrive at a vocabulary for testing AI systems, then given away — the hardest part, he says, was naming things. Read it for what it implies: the method diffuses in weeks, the graded evidence your own work produces does not. ( Source)
Challenger, Gray & Christmas — July job-cut report. AI led stated cut reasons a fifth straight month, in the lowest-cut month since July 2024. It counts what companies announce, not what they do, and Challenger's own chief revenue officer says naming AI can win over investors. ( Source)
Etsy's employee memo, Exhibit 99.2. Worth reading beside the shareholder letter filed the same day. Two documents describing one decision to two audiences; the difference is the lesson. ( Source)
Duolingo — Margin-Proof Tracker grade. In Issue 007 we re-graded Duolingo's row under a revised methodology without the underlying document having changed. Re-grading off an unchanged filing is a change to our instrument, not to the evidence, and it required a dated note we did not publish at the time. This week the row moves back to Stage 3 on a new document — the filed 10-Q — which is fresh evidence under the current ladder rather than an unannounced change to the instrument. Had we stopped at the shareholder letter, we would have recorded a miss that the filing does not support.
How we label evidence: Primary source · Corroborated · Reported · Vendor claim · Analysis. Written and edited by Mario Suarez · Independent analysis · Every link and date verified before send.
|
|
About this newsletter
AI Above the Cut is a weekly decision brief for executives — VP-and-up leaders in business who want signal over noise. It covers the outcomes and impact of AI rather than its engineering, and asks a standing question of every development: who keeps the rents? Each Sunday we read a fixed spine of the field's highest-signal voices — operators, researchers, and independent skeptics like Andrew Ng, Ethan Mollick, Simon Willison, Nathan Lambert, the AI Snake Oil team, Erik Brynjolfsson, and Cassie Kozyrkov — plus a rotating edge of specialists (Chip Huyen, Jack Clark, Ben Thompson, Benedict Evans, and others) and the primary research, regulator, and lab feeds. We tag every source — vendor, researcher, operator, investor, regulator, economist, or skeptic — and check strong claims across categories, so we curate evidence, implementation, and disagreement rather than celebrity.
The brief comes in two speeds: a fast skim — the single most important development, three concrete moves, and the week's decision-relevant signals plus one Below the Cut counter-signal — then the Margin-Proof Tracker, our standing scorecard of AI-value claims against reported P&L evidence, and a longer Synthesis that connects the moves, takes a position, and names the tests we're watching. We optimize for quality over influence, link to the source rather than the hype around it, and flag anything unconfirmed. No "10 AI tools you need today."
|