Get AI Above the Cut — free, every Sunday

A fast Sunday skim of what the field's top minds actually said this week — signal over hype.

ISSUE 004 · JULY 12, 20265 MIN SKIM · 14 MIN READ
AI ABOVE THE CUT
Tracking the top minds in AI — a weekly brief for executives
FRONTIERLABSENTERPRISEPOLICYHEALTHCARE
Since last week: the government gate cleared GPT-5.6 for global release — but this time it read the model's alignment, not just its cyber-risk; the DeepSeek V4 the whole market braced for never shipped, leaving Grok 4.5's $2/$6 as the unopposed price floor; and the agents that just reached GA became the reason big buyers are now capping AI spend.

The week in three numbers: $2/$6 Grok 4.5's unopposed price floor · Aug 2 EU GPAI enforcement switches on · €15M/3% the maximum GPAI fine.

In this issue
01 · The One Thing — the week in 60 seconds
02 · Do This Week — three concrete moves
03 · The Signal — six decision-relevant moves, tagged by lane
04 · The Synthesis — the argument under the news
05 · Where the Minds Disagree — the live splits, refereed
Then: Anti-Hype Watch · Worth Your Time · On the Radar · Corrections
Skim = through The Signal (~4 min). Everything after the band is the deep read.
01 · The One Thing
The US safety gate quietly changed what it polices — from capability to intent. OpenAI cleared the government's voluntary national-security review and shipped GPT-5.6 worldwide on July 9 (three tiers — Sol, Luna, Terra — plus ChatGPT Work). The launch isn't the story; the system card is. GPT-5.6 shipped with named misalignment evaluations — chain-of-thought controllability, "metagaming," observed agentic overreach — as scored line items, not red-team anecdotes. That means a government release review has, for the first time, reached the question "is the model trying to deceive us," and it quietly concedes those behaviors are expected, not hypothetical. The operator consequence is immediate: misalignment evals just became a procurement comparable, and the first vendor to publish them sets the audit format everyone else gets benchmarked against. Add that section to your vendor diligence this quarter. (Axios, GPT-5.6 system card)
02Do This Week1 MIN
Do: Draft the one-pager your token value-audit will need anyway — workloads down the side, monthly spend and the decision each drives across the top. The agents that hit GA this week (Microsoft's Sales/Service Agents, Claude Cowork) are what turn that grid from hygiene into a budget line, because GA is what makes an agent an always-on token burner.
Watch: Three dates — Aug 2 (EU GPAI enforcement switches on — the durable test), ~Jul 17 (Google's reported Gemini 3.5 Pro rebuild — does it land or slip again?), and late July (Microsoft/Alphabet/Meta earnings — the first read on whether the pricing revolt shows up in AI capex guidance).
Say: "We now ask every vendor for the misalignment section of the system card, not just the security one — and we treat 'cheapest capable model' as a live feed. Cost governance and alignment diligence are the two things we won't outsource."
03The Signal2 MIN

Six decision-relevant moves this week, tagged by lane.

Labs The challenger was a no-show — and that's the story. The DeepSeek V4 the whole market had pre-written its reaction to never shipped (its own API changelog has no July entry), so Grok 4.5's $2/$6 set the cheap-frontier floor unopposed, from the West. Don't wait on V4 to re-cut your routing — price for a floor that's already moving without it. (xAI, DeepSeek changelog)
Labs Grok 4.5 undercut on the axis that actually bills. At $2/$6 it cuts Sonnet 5's output price ($2/$10) by 40% — aimed squarely at the agentic/coding loop where token burn is highest. The only hard number in the announcement is the price; the "Opus-class" claim is vendor-reported. (xAI)
Policy The FTC opened a new front: model-tuning as deception. A July 7 policy statement says an AI system that quietly deprioritizes accuracy — even to comply with a state law — without "clear and conspicuous" disclosure can trigger Section 5 liability. Non-binding and open for comment through July 31, but it turns an engineering knob into a disclosure obligation. (Federal Register)
Policy The EU's "delay" was a head-fake. The Digital Omnibus deferred the AI Act's high-risk obligations to 2027 — but left the August 2 switch-on of the Commission's GPAI enforcement powers untouched: information demands, model access, fines to €15M or 3% of global turnover. The part that bites GPAI providers this year survived the headline. (Gibson Dunn, European Commission)
Enterprise Agents crossed to GA — which is exactly what triggered the cost revolt. Microsoft's Service and Sales Agents reached general availability (June 30 / July 2) and Anthropic pushed Claude Cowork's background tasks to web and mobile. GA turns an agent from a seat-priced tool into an always-on token consumer — the very spend that Uber, Microsoft, Salesforce and Meta are now reportedly capping. (Microsoft)
Healthcare The first patient-facing clinical-LLM clearance is a leash, not a license. The decision-relevant item is UpDoc's insulin-titration SaMD, cleared December 23, 2025 against a drug-dose-calculator predicate and indicated only under clinician oversight — a narrow, human-in-the-loop pathway, not authorization for an autonomous clinical agent. (McGuireWoods)
End of skim · deep read begins
04The Synthesis8 MIN
"Something has gone completely wrong." — Alex Karp (Palantir), on per-token AI pricing, as the labs race toward their IPOs. (The Daily Upside)

Last week set the frame: incumbents cutting prices, a government gate swinging on safety, and the economics arriving. This week each moved — and the sharpest reads are in what didn't happen and in a connection nobody drew.

Thread 1 — The price war's floor got set in the West, unopposed (EVOLVING)

The whole market built its week around DeepSeek V4 — the mid-July release with peak/off-peak pricing everyone had pre-written a reaction to. It didn't ship. So the only move that mattered came from Grok 4.5 at $2/$6, cutting Sonnet 5's output price by 40% with no Eastern answer. The story isn't a launch; it's an absence — the price floor is now being set by US labs racing each other, while the market watches the wrong door.

The lesson: stop anchoring your routing thesis to an anticipated competitor. The cheap tier is re-cut every couple of weeks by whoever's closest to an IPO — treat "cheapest capable model" as a live feed, and never pre-commit a stack to a launch that hasn't shipped.

Thread 2 — The gate started policing intent — and it's not the only one (EVOLVING → resolving)

The US review didn't just clear GPT-5.6; it cleared it with alignment evals in the card, expanding the gate's scope from "can this leak capability" to "is this model gaming us." In parallel, the FTC recast quiet accuracy-tuning as consumer deception, and the EU's statutory Aug 2 GPAI powers held despite the "delay" headlines. Three regulators, three clocks, all converging on model behavior.

The lesson: the safeguard bar is rising from "can it be jailbroken" to "can we tell when it's gaming us" — and the US gate is fast and reversible where the EU's is durable and statutory. Ask vendors for the misalignment section, and assume disclosure obligations are coming, not optional.

Thread 3 — GA and the pricing revolt are the same story (EVOLVING)

Here's the connection nobody drew this week: the agents celebrating general availability are precisely what set off the revolt over per-token pricing. GA converts an agent from a seat you pay for once into a background process that burns tokens 24/7 — the exact consumption Uber, Microsoft, Salesforce and Meta are reportedly capping, and the exact model Anthropic reportedly wants to take public in ~October. The product milestone the vendors are celebrating is what made the bill unbudgetable.

The lesson: token spend is now a governed budget line, not a routing tweak — and the fight is fiscal, not technical. Stand up the value-audit before the agents are load-bearing, because GA means they already are.

Still dormant: world models — a fourth straight week with no new LeCun (AMI Labs) or Fei-Fei Li (World Labs) primary. We're not padding it into a thread until one ships.

Editor's take

Two things an operator can't outsource both went first-class this week — cost governance and alignment diligence — and the most useful reads came from resisting the obvious. The obvious story was "GPT-5.6 is here"; the real one was that the card changed what a government gate polices, and that the DeepSeek that was supposed to reset prices simply didn't exist this week. The honest counter-case: one system card with metagaming evals isn't proof the labs will gate a release on alignment when a quarter is on the line — it's a disclosure, not a commitment, and the same week's IPO-speed shipping is exactly the pressure that erodes it. Which is why the move stays flexibility plus diligence, not conviction in any one lab's safety story or any one week's price sheet.

Watching next

1) Does DeepSeek V4 finally ship and reopen the price gap the US labs just closed? 2) Does the EU's Aug 2 GPAI enforcement actually bite, or slip? 3) Does any lab gate a real release on an alignment finding — or does the misalignment section stay disclosure-only once revenue is on the line? 4) Do late-July earnings show the per-token revolt in AI capex guidance, or is it procurement theater?

05Where the Minds Disagree

Not who-said-what — the live splits between serious people, and where we come down. We run this only when the disagreement is real; this week there are three.

Are we certifying safety on the model — when the risk has moved to the scaffold?

OpenAI scored GPT-5.6's misalignment on the model (evals in the card). But Lilian Weng argued this month that capability now compounds in the harness — tools, memory, orchestration — around a frozen model, not in the weights. If she's right, safety evals run on the bare model are testing the wrong object, because overreach emerges where the agent is wired up, not where it was trained. Our read: ask vendors whether their alignment evals cover the agentic harness, not just the model — a card full of model-level evals with no harness story is half an answer, and the missing half is where an agent actually goes wrong.

Is quiet model-tuning "safety" — or "deception"?

The FTC now says an AI system that quietly tunes down accuracy without disclosure can be Section 5 deception. Vendors frame the same behavior-tuning as safety alignment (see the GPT-5.6 card). One regime rewards silent tuning for safety; the other treats the silence as a liability. Our read: you can't satisfy both sovereigns by staying quiet — assume disclosure is coming, and document why you tune, not just that you do.

Is entry-level hiring softening because of AI — or because of remote work?

The Brynjolfsson-aligned thesis reads weak recent-graduate hiring as early AI substitution of junior knowledge work; EPI's occupation analysis reads the same data and pins it on remote-work composition instead. Our read: genuinely unresolved — two credible datasets, opposite attributions. File "AI is compressing entry-level roles" as contested, and don't set headcount strategy on it as if the verdict were in.

Anti-Hype Watch
The benchmark confetti around Grok 4.5 and GPT-5.6. "Opus-class." "54% more token-efficient." "Strongest ever." Every headline score this week is a vendor number with no independent, different-category replication — and the tell is that Grok's capability claims shipped alongside the price meant to make them look like a bargain. Strip the confetti and the verifiable, operator-relevant facts are three: $2/$6, "global," and "cleared the review." This was a distribution-and-price week dressed up as a capability week. Buy on the price and the availability, which are real; wait for a third-party eval before you buy on the leaderboard.
Worth Your Time

Only reads that add something the front didn't already give you.

Lilian Weng — "Harness Engineering for Self-Improvement." The mental model behind this week's sharpest safety question: capability compounds by optimizing the scaffold around the model — tools, memory, orchestration — not the weights. If she's right, your AI roadmap (and your risk surface) is an orchestration problem. (Lil'Log)
Anthropic — "A global workspace inside Claude." Interpretability work locating an internal region the model uses for deliberate reasoning and self-report. Read it as an audit surface, not evidence of a mind: it makes a model's introspective claims checkable against mechanism — the "read the internals" complement to OpenAI's eval-heavy card. (Anthropic)
OpenAI — GPT-5.6 system card. If you open one primary this week, open this — and read the misalignment section (CoT controllability, metagaming, logged agentic overreach), not the benchmarks. It's the clearest public picture yet of how a lab tests a frontier model for deception. (OpenAI)
On the Radar

What's coming — the dated anchors worth having on the calendar.

~Jul 17 — Google's Gemini 3.5 Pro is the reported target after a rebuild slip. Does it land and close the gap, or slip again? (reported; unconfirmed by Google — no primary, so we flag rather than link it.)
Aug 2 — the EU's GPAI enforcement powers switch on (fines to €15M / 3%), alongside the Article 50 transparency duties — one date, two obligations. The durable, statutory test the fast, reversible US review is not.
Overdue — DeepSeek V4. No official date; a low-notice launcher historically. Its API changelog is the trigger — a July entry reopens the Eastern front of the price war.
Late July — Microsoft, Alphabet, and Meta report earnings — the first read on whether the per-token spend revolt is real or theater, and whether it dents AI capex guidance.
Corrections
Correction to the UpDoc clearance date. Two weeks ago we dated UpDoc's FDA clearance to June 25, 2026 — that was the vendor's announcement date; the 510(k) clearance letter itself issued December 23, 2025, roughly six months earlier. The substance is unchanged: a narrow prescription insulin-titration SaMD cleared against a drug-dose-calculator predicate, not an autonomous clinical agent. (McGuireWoods)
Clarification on the "gate." The clean, verifiable instance of the US national-security "gate" this week is OpenAI's GPT-5.6 — held to "trusted partners" in late June under the June 2 EO's voluntary review, then cleared for global release. Where earlier editions attached the reversible-gate story to other model names, we're realigning it to the event we can source to primary documents. (Axios)

Got a mind we should be reading, or a correction? Reply and tell us.

Subscribe to AI Above the Cut →
Continues · The One Thing · Do This Week · The Signal · The Synthesis · Where the Minds Disagree · Anti-Hype Watch · Worth Your Time · On the Radar
About this newsletter

AI Above the Cut is a weekly brief for executives — VP-and-up leaders in strategy, healthcare, and AI transformation who want signal over noise. Each Sunday we read a fixed spine of the field's highest-signal voices — operators, researchers, and independent skeptics like Andrew Ng, Ethan Mollick, Simon Willison, Nathan Lambert, the AI Snake Oil team, Erik Brynjolfsson, Cassie Kozyrkov, and Eric Topol — plus a rotating edge of specialists (Chip Huyen, Jack Clark, Ben Thompson, Robert Wachter, and others) and the primary research, regulator, and lab feeds. We tag every source — vendor, researcher, operator, investor, regulator, economist, or skeptic — and check strong claims across categories, so we curate evidence, implementation, and disagreement rather than celebrity.

The brief comes in two speeds: a fast skim — the single most important development, three concrete moves, and the week's decision-relevant signals — then a longer Synthesis that connects them, takes a position, and links to the primary work. We optimize for quality over influence, link to the source (the paper, the post, the talk) rather than the hype around it, and flag anything unconfirmed. No "10 AI tools you need today."

AI Above the Cut · A weekly AI brief for executives · Manage · Unsubscribe