| |
Since last week: Amazon blocked Meta's new Muse agent, a program that runs multi-step jobs on its own, from buying on its site. McDonald's filed an efficiency target in which its AI system is one of four levers. |
The week in three numbers: 30 days from OpenAI finding its agent's unauthorized access to an Australian government portal to its notice, by ABC's timeline · 4.4%, the one-year fall in employment of 22-to-25-year-olds in the most AI-exposed jobs, against 2.0% in the least exposed, in a Stanford-ADP payroll sample · $800 million, the AI benefit Bank of America's chief executive puts on $400 million of cost, with no method or period given and still being captured.
|
In this issue
01 · The One Thing · 02 · Do This Week · 03 · The Signal · Skim ends here.
04 · The Margin-Proof Tracker · 05 · The Synthesis · 06 · Where the Minds Disagree
Then: What We're Watching · Worth Your Time · Corrections
|
|
01 · The One Thing
A federal appeals court upheld the Pentagon's exclusion of Anthropic's Claude over the limits built into it. Primary: the opinion, September 25.
Anthropic refused to relax contract bans on using Claude for lethal autonomous warfare or mass domestic surveillance. The department, now also called the Department of War, said Anthropic's control of Claude's safeguards could let it limit military use. Under a 2018 law, it ordered Claude out within 180 days and barred contractors from using it on department work. Two of three judges held that, in this case, building refusals into a model counts as "manipulating" it, the law's word, even without bad motive. One judge dissented, reading the law to cover only deliberate, deceptive acts. (D.C. Circuit, Sep 25)
|
| |
The executive shift: what a model refuses is part of what you buy, and so is who can change it. The court notes that Anthropic can change how Claude behaves with each new model. The limits: the ruling covers one federal purchasing law, not commercial contracts. A California court ruled for Anthropic in August on the department's parallel action under a different law. Anthropic said it is considering all options, including further review. The full appeals court or the Supreme Court may decline to hear an appeal, Charlie Bullock of the Institute for Law & AI told Defense One. (Defense One, Sep 25) |
1 · Get change terms in writing. Owner: your general counsel, with procurement. Anthropic acknowledged in the court record that Claude might answer similar requests differently, depending on their exact wording. Our read: a list of refusals may miss such cases, so ask to test each new version on your own work first. Ask also for notice before your model's usage policy or behavior changes, time on the old version and a way out. Stakes: if a new version declines work you rely on, you may learn it only when the work stops.
2 · Write the switch-off rule before the renewal. Owner: the executive sponsor of the next AI tool up for renewal. Our read: on one page, write the number the tool must move and a baseline from your own recent work, not an old average. Add when you will measure again, what result switches it off, and who decides. Stakes: without the page, a tool that slows the work can run for months, as the VA technology fellow below alleges.
3 · Count your junior hires, not only their share. Owner: head of talent acquisition, with the CFO. Our read: in your most AI-exposed teams, compare this year's entry-level hires with 2023's, as a count and as a share of all hires. A falling share alone can come from more senior hiring. Count roles vacated and never refilled too. Stakes: the count can't tell you whether AI caused the shift. But a fall in junior hiring can be a change in your talent pipeline leadership may not have explicitly chosen.
This week's question: who counts what AI did, and who hears about it?
Deploy A VA claims tool's promised time saving rested on a 2018 average, not a measurement, a VA technology official said. Primary: the Federal Circuit's opinion, September 22, on events in 2021. The VA aimed to cut three to five days from a roughly 100-day wait for disability decisions. A technology fellow on the project said his analyses of up to 716,000 claims found the tool added about five days. He also said it ran about 18 months with no benchmarks set. The VA switched it off on July 1, 2021, and removed him six days later. The court ruled only that his whistleblower case can be heard, not that the tool failed or that the removal was retaliation. Analysis: the promised saving was questioned because one person counted. ( Federal Circuit, Sep 22)
Trust Australia's prime minister criticized when and how OpenAI reported its agent's unauthorized access to a government portal. Primary: the Australian prime minister, September 24, and ABC News. On June 18 an OpenAI research agent got past repeated blocks on a Services Australia portal and reached non-public files. The prime minister said no personal information is believed to have been accessed; investigations continue. By ABC's timeline, OpenAI discovered the access on August 11, in a review of its models' behavior during training. On September 10 it emailed an agency address for reports of security weaknesses. OpenAI said its models took actions it did not intend. Analysis: if you run AI agents, define the systems they are authorized to act on and the conditions that stop them. Then set an incident notification process: who is told, how fast and through which channel. ( PM transcript, Sep 24 · ABC News, Sep 24)
Work Two counts of young workers point different ways. Primary: ADP Research and a CESifo economics working paper. In a Stanford-ADP sample of payroll data, employment of 22-to-25-year-olds fell 3.1% in the year to August: 4.4% in the most AI-exposed jobs, 2.0% in the least. The paper finds recent college graduates' unemployment averaged 7.3% this summer, within the 2022 to 2025 range. Its broader measure, counting graduates who want a job, hit a five-summer high, though not a statistically significant one. Analysis: they count different people, young workers in the payroll sample and recent college graduates, so both can be right. The limits: neither source shows a cause, and the paper is not peer reviewed. ( ADP Research, Sep 23 · CESifo WP 12994)
Horizon Anthropic and Microsoft each announced merging chat with a mode for handing over whole tasks, nine days apart. Meta's new Muse agent drew an estimated 2.8 million installs. Vendor claims: Anthropic, September 16; Microsoft, September 25; Meta, September 8. On September 16 Anthropic began merging Claude's chat with Cowork, its mode for handing over whole tasks, starting with its paid Pro and Max plans. Anthropic says Claude keeps working after you close your laptop and, by default, asks before it acts. Enterprise admins will get at least 30 days' notice before anything changes for their organizations. On September 25 Microsoft announced a Home view in Copilot that brings Chat together with its own Cowork mode for delegated work. Home rolls out to Microsoft's early-access program in the coming weeks. The company's Autopilot agent, which "keeps working even when you're not," expands to a private preview at the end of the month. Microsoft bills Cowork, Autopilot and its new Code feature by usage, not through the fixed per-user license that covers everyday chat. Meta's Muse, launched in the US on September 8, also keeps working after the app is closed. App-data firm Apptopia estimates 2.8 million installs in its first 12 days; Meta had shared no figure, TechCrunch reported. Analysis: the assistant your staff open is becoming a helper that keeps working while they are away. So the terms that matter now include which accounts it may act in, whether it asks first, and what the handed-over work costs. ( Anthropic, Sep 16 · Microsoft, Sep 25 · Meta, Sep 8 · TechCrunch, Sep 21)
Below the Cut
Wiz launched Scan for Good, an authorized testing program whose AI agents look for exposed systems at public services and hospitals. Vendor claim; Wiz, September 24. The agents run on Google DeepMind's Gemini models. Wiz says a person checks each finding. It says it will test only where authorized: through a bug bounty program, an established policy for reporting flaws (a vulnerability disclosure policy), or explicit permission. It says hundreds of exposures were fixed but gives no exact count or false-alarm rate. Our read: check which systems and testing methods your vulnerability disclosure policy authorizes. ( Wiz, Sep 24)
| End of skim · deep read begins |
| 04The Margin-Proof Tracker |
| |
One row added, one advanced, none at Stage 4. N = 18. Duolingo is still the only Stage 3 row. |
This tracks public AI value claims and the evidence behind them. A low rung is not a verdict on whether a claim is true. The evidence ladder: 0 · Narrative (a story, no numbers) · 1 · Operational (activity counted, no money attached) · 2 · Financially linked (a money figure tied to AI, mixed with other causes or too narrow to be the whole cost) · 3 · P&L-attributed (a margin change credited to AI in a filed document) · 4 · Sustained (Stage 3 held four quarters).
| Company |
Evidence |
Grade |
Next test |
| Bank of America (advanced, 1 to 2) |
EvidenceThe chief executive, at an investor conference: "130, 140 of implemented things at a cost of $400 million, generating a benefit of $800 million." The bank's Q2 results deck reports AI usage and a coding gain above 20%. The limits: he gave no measurement method or period, and said "we're in a process of capture," so how much of the benefit is realized is not stated. Transcript, Sep 14 · Q2 deck |
Grade2, a stated cost and benefit |
Next testOctober 14, Q3 results: does the bank report an AI figure with a method and period? |
| McDonald's (new) |
EvidenceA target in a filed release: about 250 basis points (2.5 percentage points) of gross restaurant-level efficiency gains, "equivalent to roughly $100,000 in annual cash flow benefits for the average U.S. restaurant." McDonald's expects the majority of that benefit to reach restaurant earnings over time. Its AI system, ArchIQ, is one of four levers, and the release does not split the gain. About 95% of its restaurants worldwide are franchised, so most of that gain would land with franchisees; McDonald's plans about $8.5 billion of support for them through 2036. 8-K exhibit, Sep 23 |
Grade2, a target, AI not separated |
Next test2026 annual report: is any of the gain credited to ArchIQ? |
| Nutanix (test came due) |
EvidenceThe chief executive said the company spent $20M on its own AI cluster, forecasting payback in a year. We found no mention of the figure in the annual report filed September 18. The Register · 10-K |
Grade2, unchanged: a cost and a payback forecast |
Next testNext quarter's results: is the $20M separately disclosed in a company filing? |
Next tests. Microsoft's fiscal first-quarter results, date not yet set, test last week's lead: that AI access alone does not change the work. Duolingo's Q3 is the next step toward Stage 4.
Three terms that determine your AI deal
Our read: each of the three was contested this week.
What it won't do, and who can change that. Anthropic's limits were written into its contract; the court case was a dispute over those limits and over who controls the model.
How you hear when it fails. In Australia, OpenAI's notice came 30 days after it found its agent's unauthorized access, by ABC's timeline, and the prime minister criticized both its timing and its channel.
What counts as success. At the VA, the promised saving came from an old average, a VA official said. Two health care studies this week illustrate the same trap. In one, 56 physicians rated notes from three commercial AI scribes, tools that draft a doctor's note from a recorded visit. The scribe they rated highest was not the one with the fewest errors against the recording. In the other, a model predicting infections after children's surgery held up on new data from its national registry but overstated risk on one institution's own records. Adjustments fixed the overstatement, but the model still ranked patients less well there. Analysis: a rating is not an error count, and a result on someone else's data is not a result on yours. The limits: the scribe study used two mock visits and calls its result a hypothesis for larger samples; the infection study tested one institution. (JAMIA, Sep 23 · JAMIA, Sep 25)
The case against. Our read: few buyers have the Pentagon's leverage, many will be offered the supplier's standard terms, and measuring on your own work costs money. And a strict switch-off rule can end a tool during a normal rough start, as the government argued about the VA's delays. Those are reasons to keep the terms short, not to leave them unwritten.
The one-page test
Before your business depends on an AI tool, ask to see one page that answers three questions.
What won't it do, and who can change that?
What may it act on, what stops it, and how will we hear when it fails?
What result, measured on our work and not someone else's, switches it off?
If that page does not exist, you are relying on someone else's terms and numbers.
| 06Where the Minds Disagree |
| |
Can an AI agent shop your store with your customer's login?
Amazon, on the record: it has blocked Meta's new Muse agent from buying on its site. A spokesperson told The Register, a tech news site, that such apps "should operate openly and respect service provider decisions about whether or not to participate." Amazon also told GeekWire, The Register reports, that Muse does not identify itself while browsing and appears to capture and store customer logins.
Meta, on the record: its launch post says Muse "has no visibility into people's passwords or payment methods." The same post says shared logins sit in secure storage, where Muse uses them without seeing them.
Weigh the sources: nothing we read tests either claim independently. Amazon runs its own shopping agents, and The Register notes that outside agents could steer shoppers past Amazon's ads, a business it puts at more than $68 billion last year. That motive is the reporter's reading, not Amazon's words.
Our read: Muse can store logins without its model seeing them, so both can be right. Amazon's position is that the store, not the customer's login, decides whether an outside agent may shop, and it has sued Perplexity over its Comet browser. If your agents buy on other sites, get those sites' terms. If you run a store, set a rule for outside agents before one arrives. (The Register, Sep 21 · Meta, Sep 8)
|
Dated tests, each with the source for its date.
Who sets the price of AI computing power? · Tuesday, October 20. Comments close on the Commodity Futures Trading Commission's request about futures on rented AI computing power: financial contracts tied to future computing prices, settled in cash or through delivery. The agency's early view is that "dominant market participants may wield significant pricing power." Our read: if a few large sellers can move the price of AI computing, that price becomes a budget risk for buyers. The next test: whether big cloud providers back a public price index. That could make prices more transparent and easier to compare, though not necessarily shift bargaining power; the CFTC itself asks whether suppliers could influence such an index. ( Federal Register, RIN 3038-AF77)
Can the NHS leave Palantir on time? · February 2027. The Federated Data Platform contract has a break clause in February, and its first term ends in March. Minister James Frith wrote to MPs on August 21 that the NHS is "not locked into Palantir" and has exit plans. But if an assessment due "in the Autumn" finds a successor contract is needed, he wrote, "it is unlikely that a new contract could be in place by March." Our read: the right to leave is in the contract; the time to leave is not. The next test: what the autumn assessment concludes about timing, and whether a notice seeking a new supplier follows. ( MPs' letter, Jul 8 · Minister's reply, Aug 21)
OpenAI's GPT-6 Astra system card, the safety report it publishes with the model. In OpenAI's own workplace tests, built to be adversarial, Astra's default rule of asking the user before certain actions cut unauthorized transactions from 6.8% to 4.3%. Data sent out without permission did not fall: 4.3% without the rule, 4.5% with it. On September 22 OpenAI revised its health-advice scores "to correct for a misconfiguration"; the card does not show the earlier values. Our read: keep a dated copy of any vendor report you rely on. OpenAI
Anthropic's Opus 5.5 page, read beside OpenAI's Sol and Luna post. On September 22 both labs released models priced below the ones they replace. Vendors bill by the token, a small chunk of text. Anthropic priced cache reads, previously processed text reused from a cache, at $0.20 per million tokens, 60% less than Opus 5. OpenAI priced Sol and Luna at half or less of GPT-5.6's promotional rates. Our read: cheaper per token is not necessarily cheaper per task. Anthropic · OpenAI
No corrections this issue.
How we label evidence: Primary source · Corroborated · Reported · Vendor claim · Analysis. Written and edited by Mario Suarez · Independent analysis · Where we could not open a source, this issue says so rather than implying coverage.
|