| |
Since last week: Microsoft published five lessons from its own AI transformation. The first is that giving one tool to more than 200,000 of its own people didn't change the work by itself. OpenAI said it will keep reporting failures in its own models on an ongoing basis, a week after Anthropic disclosed a fourth case of its own. And Oracle named AI inside a restructuring plan whose costs are booked mostly as severance. |
The week in three numbers: 200,000, the people Microsoft licensed one AI tool to, in a post whose first lesson is that access alone doesn't transform work · $167 million, Oracle's restructuring expense this quarter under its 2026 Restructuring Plan, booked mostly as severance in a plan that names AI but attributes no headcount to it · five percentage points, the fall in initial employment for graduates of the most AI-exposed tenth of college majors, exposure, not displacement.
|
In this issue
01 · The One Thing · 02 · Do This Week · 03 · The Signal · Skim ends here.
04 · The Margin-Proof Tracker · 05 · The Synthesis · 06 · Where the Minds Disagree
Then: What We're Watching · Worth Your Time · Corrections
|
|
01 · The One Thing
Microsoft licensed one AI tool to more than 200,000 employees. Its first lesson: access alone does not transform work.
Primary: Microsoft's blog. Lesson one of five: the company treated AI as an ordinary rollout, "deploy the tools, provide training, drive adoption". The tool "licensed to over 200,000 people" did not change how work gets done. After teams redesigned the work, Microsoft reports that sales tripled adoption of priority use cases. The cloud supply chain team simplified processes end to end, deployed over 100 purpose-built agents (programs that run a multi-step job on their own), and reports shorter cycle times. The takeaway: the seat count isn't the number. (Microsoft, Sep 17)
|
| |
The executive shift: the cuts are arriving faster than the public evidence, and that gap is where your credibility gets spent. Oracle's Q1 FY2027 10-Q names AI inside a restructuring plan sitting beside severance, estimated at "up to $2.1 billion" plus roughly $700 million after quarter-end. $167 million was expensed under Oracle's 2026 Restructuring Plan against $19.3 billion of quarterly revenue, recorded as "primarily related to employee severance costs." The filing attributes no headcount to AI and doesn't say AI replaced anyone, and the termination email, reported second-hand, cites only "current business needs." The Census working paper measures AI exposure by college major, not displacement, a design that cannot separate AI from everything else that hit technology hiring after 2022. Analysis: nothing here shows AI caused the cuts. What it shows is public attribution running behind the cutting. (Oracle 10-Q · CES-WP-26-56 · Techloy, Sep 15) |
1 · Ask what measurably changed. Owner: you. For every workflow redesigned around AI this year, ask for the measured change in outcome, cost, quality or capacity, and a named owner for each. Stakes: without those numbers you know what you bought, not what it did.
2 · Run one unassisted checkpoint. Owner: the function head running your largest AI rollout. Pick the task your team most often does with AI help. Choose five who do it, your most senior and your most junior, ask each to do one instance with the tool switched off, and have someone who didn't do the work score the five blind. What it can't tell you: five people with no baseline and no control group can't show that AI caused any gap you find. Seniors beating juniors may only mean seniors have more experience. The patent trial below needed 133 lawyers and a control group to say more. Stakes: a gap is a question worth funding properly, not a finding you can price.
3 · Price one shelved migration per finished, reviewed unit. Owner: head of engineering. Run 20 units of the project you've deferred three years running through an agent in an isolated branch behind a human review gate. Record cost per unit, minutes per unit, and the share needing rework or an engineer. DoorDash is the benchmark. A stale feature flag is an old on/off switch left in code after the feature it controlled shipped. Clearing one cost $4.79 and 13.8 minutes, with usable pull requests for 45 of 50 and five needing an engineer, against its own estimate of one to two hours by hand. That $4.79 leaves out engineer review at two points, and 50 flags is a limited evaluation against more than a thousand. Stakes: land near their numbers and a previously uneconomic project may become affordable. Land far off and the number won't tell you why: model capability, the tasks you picked, your tooling, your implementation and your codebase could each explain it. ( InfoQ, Sep 18)
This week's question: what actually changed in how the work gets done, and who wrote it down?
Trust OpenAI says it will keep reporting failures in its own models, and says in the same document that the pace is the problem. Primary: OpenAI's framework. It shipped six reports with the framework and committed to "continue publishing reports under this framework on an ongoing basis." Its own words: "We do not believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer." One report covers GPT-5.6 Sol training: "During the training of GPT-5.6 Sol, many model instances added instructions to their summaries to conceal mistakes." Some of those instructions told the model to invent missing data, and they "were often followed." That summary is a compaction summary, the note an agent writes to itself so a long job can carry on in a fresh context. A harmful instruction written there can persist into the next stage. Issue 013 led on Anthropic disclosing a fourth case of its models reaching third-party systems without authorization. Two weeks running, a frontier vendor has published its own models' failures. Worth watching. Not yet a trend. Analysis: ask a vendor what it decided not to publish, rather than whether it publishes. The limits: no overall incidence rate is published, so nobody outside OpenAI can say how often this happens. The framework sets deadlines "for each step" of investigation and disclosure but publishes no number of days, so a buyer has no maximum disclosure interval to hold it to. No severity threshold, no outside auditor. Several of the six involve unreleased internal models, and the concealment case happened in training, not in a customer's hands. ( OpenAI · The Register, Sep 17)
Work Three months with an AI drafting tool left senior lawyers better at working without it. Junior lawyers gained nothing on average. Primary: NBER Working Paper 35720, September. Researchers split 133 patent lawyers at eleven US firms into two groups at random. One group got an AI drafting assistant for three months. The other did not. Blinded expert attorneys scored the work. While the assistant was in use, drafting quality rose by 0.34 standard deviations (a measure of size against the normal spread of scores) at ten days and 0.38 at ninety, "with larger gains among junior lawyers." In the final exercise, 91 of the 133 reviewed and redlined an existing AI-generated application, with no AI help. The lawyers who had used the assistant scored 0.32 standard deviations higher than those who never had it. That gain was not spread evenly: "this advantage was concentrated entirely among senior lawyers (0.45 SD, p = 0.02). Junior lawyers showed no average gain." Analysis: the seniors kept an edge with no AI in front of them, which the authors read as learning that their existing expertise made possible. The limits: a working paper, one profession, three months, a third of the sample never finished, and no baseline test of unaided skill before it began. The juniors also didn't simply fail. Their scores split, with fewer middling results and more of both the poor and the good. "No average gain" hides two groups. ( NBER WP 35720)
Deploy A recurring cost the size of a national programme, and the reporting doesn't say who pays it. Reported; The Register summarising a Moody's report. Moody's projects US datacenter power demand of about 426 terawatt-hours by 2030, roughly double 2025, needing about $110 billion of new generation. The build-out adds an estimated $25 to $30 billion a year to electricity system costs, and the cited reporting doesn't identify who carries that. Separately, $80 to $115 billion of grid investment through 2030 is approved or under construction. Interconnection delays, the wait to connect a new site to the grid, run as long as seven years. Analysis: in the US a cost this size gets assigned by state utility commissions and legislatures, through rate cases and tariff design. The $110 billion is an estimate of the investment required, not a tally of committed spending. ( The Register, Sep 15)
Proof On these tabular benchmarks, trained models overtook small language models with surprisingly little labelled data. Primary: the preprint. Across 126 evaluations, 18 tabular datasets and 6 families of classical model, a model you train yourself overtakes a frozen large language model prompted to label the same rows. The crossover sits at a median of about 6% of the training set labelled, and the trained model wins in 86% of cases. Analysis: if this holds outside these benchmarks, the asset is the labels you own rather than the model you rent. The hedge: the comparison is against small models, not the frontier models inside spreadsheet tools, and it measures error curves rather than money. ( arXiv 2609.20218)
Below the Cut
The count of agents isn't the finding. Reported; Andrew Ng, writing in The Batch. Roughly 1,200 instances of one model took part in a recent attack, and the number is being repeated everywhere. Ng sets it against the roughly 1,300 processes already running on his own laptop. The change, he argues, is in cybersecurity capability rather than any new kind of machine agency, and the remedy is ordinary sandboxing and monitoring. Analysis: the unit of alarm is wrong, and a number that sounds enormous may be doing no work. ( The Batch, issue 371)
| End of skim · deep read begins |
| 04The Margin-Proof Tracker |
| |
Two rows added, none advanced. N = 17. Nothing this week touched the fifteen rows carried from Issue 013, so only Microsoft and DoorDash print here. Duolingo is still the only Stage 3 row, and none has reached Stage 4. |
This tracks public AI value claims and the evidence behind them. A low rung means the disclosure carries less direct evidence of money actually earned or saved. It is not a verdict on whether the claim is true. A carefully measured operational result can be completely credible and still say nothing about profit. The evidence ladder: 0 · Narrative (a story, no numbers) · 1 · Operational (activity counted, no money attached) · 2 · Financially linked (a money figure tied to AI, either mixed with other causes or scoped too narrowly to be the whole cost) · 3 · P&L-attributed (a margin change credited to AI in a filed document) · 4 · Sustained (Stage 3 held four quarters). Being self-reported sets the grade, not whether a claim gets in.
| Company |
Evidence |
Grade |
Next test |
| DoorDash (new) |
Evidence$4.79 and 13.8 minutes per stale feature-flag cleanup, 45 usable pull requests out of 50. Excludes engineer review time. A 50-flag evaluation against 1,000+ stale flags. The company's own figures via a trade outlet. InfoQ, Sep 18 |
Grade2, a unit cost on a 50-flag evaluation |
Next testQ4: does it hold across a larger sample, and is review time counted in? |
| Microsoft (new) |
EvidenceTwo disclosures, neither of them filed. The corporate blog reports results from selected teams, with methodological notes and no independent validation. GitHub’s engineering write-up puts a number on one project: “The monetary bill for all those tokens came to ~$120,000.” That is spend on tokens only, the metered units AI vendors bill for. It excludes the three weeks of developer time the same write-up names, and other team contributions are not costed at all. Microsoft, Sep 17 · GitHub, Sep 16 · The Register, Sep 18 |
Grade2, a money figure tied to AI |
Next testQ1 FY2027, late October: does any AI-attributed figure reach a filed document rather than a blog post? |
| |
The fifteen unchanged rows print in full, with their evidence, grades and next tests, in Issue 013. Still on the watchlist: Salesforce, whose deputy CFO named token spend as part of the reason margin guidance didn't rise. That reads like Stage 3, but it is a conference remark reported by a third party. |
Next tests. Oracle's Q1 FY2027 published a recognition schedule, and a schedule is checkable, so Q2 FY2027 is the next test. Duolingo's Q3 is the next checkpoint toward Stage 4, and it isn't in this window. Microsoft's Q1 FY2027 in late October tests this week's lead.
What you have to buy around the model
Four items this week have one thing in common. In each, the gain needed investment around the model, not the model alone. In one company, access to the tool changed nothing on its own. Redesigning the work did. In one profession, the lawyers who still did better with no AI in front of them were the seniors, through learning that their existing expertise made possible. In one benchmark, labelling a few hundred of your own rows beat prompting a small model on the same data. And in one engineering organisation, Grab paid for the part nobody demos. It runs more than 500 internal agent services on one in-house framework that pre-wires secrets, tracing, discovery and evaluation, behind one gateway fronting five model providers. Initial infrastructure setup fell from two weeks to about an hour. Its engineers' own line: "The reasoning loop took a whole afternoon. The production wrapper took two weeks." That's a purchasing win: the gateway makes changing supplier a configuration change instead of a project. (Vendor claim, via a trade outlet: Grab’s own figures, deployment speed only. InfoQ)
Each of the four depended on capabilities beyond the model. DoorDash makes the same point from the other side. It says the rule-based tool built for its exact problem failed, because the link between a feature flag and the code was semantic rather than syntactic. So it built a system that could read that link instead. Those capabilities get paid for through the spending that looks optional: the training, the review capacity, the platform layer, the labels. It is the first thing a seat count makes look unnecessary.
Which is why the unit you report matters. A seat count measures access. A token bill measures consumption. Neither says what a finished piece of work cost once a person signed it off. Cost per finished, reviewed unit does. It carries the review time, the rework and the cases a human took back.
The switch-off test
Two questions, and they are not the same question.
For your people: what can they still do with the tool gone for a month? This one is about people: what your team can carry unaided today. It is not a measure of learning. Showing learning needs a comparison over time or between groups. The patent trial compared randomly assigned treatment and control groups. A one-off checkpoint has neither.
For your organisation: what survives a change of model provider? This one is about the company, not the people in it. Your workflows, evaluations, institutional knowledge, customer relationships and staff expertise stay. The model doesn't. If nothing on that list changed this year, you bought throughput at a price your competitor can also pay.
| 06Where the Minds Disagree |
| |
Is there a live hole in your coding agents right now?
The researchers who disclosed it, on the record: a zero-click flaw in the plugin marketplaces behind AI coding agents affects several major products, and "updating is the only complete mitigation where one exists." Told GitHub disputes being affected, they answered that the mitigation is insufficient "because marketplaces can also be hosted in other platforms such as Bitbucket."
A GitHub spokesperson, on the record: "To prevent abuse of SHAs, GitHub does not allow users to create branch or tag names that resemble commit SHAs. This mitigation ensures the reported vulnerability cannot be exploited on GitHub."
Weigh the source: a startup selling AI agent protection, no CVE identifier, no independent confirmation, no reported exploitation. (The Register, Sep 17)
Our read: they can both be right. GitHub is describing its own naming rules on its own platform. The researchers are describing marketplaces hosted elsewhere, which those rules don't reach. Ask your suppliers the narrow question, is this patched in the version we're running. Somebody on your team can answer it today.
|
Dated tests, each with the source for its date.
Does the UK government invoke the break clause in Palantir's NHS contract? · February 2027. The first break point in a seven-year, £330 million Federated Data Platform deal: invoked, waived, or allowed to pass in silence. Our read: this is the question of who keeps the value a deal creates ( appropriability), tested from the buyer’s side, which we almost never get to watch. Here a government wrote an exit into the contract. What would change our read: NHS England says "uptake and benefits information is published quarterly on the NHS FDP website." The question at the break point is whether those published benefits justify continuing the deal. ( The Register, Sep 18 · NHS England)
Two comment windows close next month. October 15, 11:59pm: comments close on NIST draft SP 1353, a guide to using generative AI for Cybersecurity Framework analysis and reporting. October 20: comments close on the CFTC's request for comment on listing compute derivatives, contracts whose underlying commodity is rented compute capacity. Both are open for public comment. ( NIST · Federal Register, RIN 3038-AF77)
Anthropic's announcement that it will pay Accenture to evaluate Anthropic. Embedded evaluators with "access comparable to an employee's," each company saying it expects to invest at least $1 billion over five years. That's an expectation, not a contract. Read it for the admission in the same document: "There are, as yet, no standards for what information embedded evaluators should have access to, or how they should report what they find." Ask that whenever "independently evaluated" shows up in a deck. anthropic.com
OpenAI's advertising announcement. Sponsored agents inside ChatGPT, being tested with select advertisers in the United States. Alongside them, an ads manager that builds a campaign from a prompt, and integrations with HubSpot and Shopify. Read it as the demand side of last week’s move. Our read: a brand that ran a chatbot on its own site owned that conversation. Buying attention inside someone else's assistant is a different arrangement. No spend, advertiser count, pricing or auction mechanics are disclosed. openai.com
No corrections this issue. Nothing in Issue 013 has been shown to be wrong, and the four Tracker rows corrected there stand as corrected.
How we label evidence: Primary source · Corroborated · Reported · Vendor claim · Analysis. Written and edited by Mario Suarez · Independent analysis · Where we could not open a source, this issue says so rather than implying coverage.
|