Podcast episode
AI:AM: Was Trump-Xi Anything? What Counts as Utopia? + AWS GPUs Cost 3X & AI Diagnoses Rare Diseases
cloud-costs evals guardrails inference tool-use
TL;DR
A dense AI-focused weekly roundup from The Cognitive Revolution covering four distinct threads: US-China AI diplomacy and compute verification, hyperscaler GPU pricing premiums (2–3× over neo-clouds), OpenAI Dev Day model releases, and a compelling deep-dive on Gamow Labs using AI agents to diagnose rare genetic diseases in children. High signal-to-noise for AI infrastructure operators, applied-AI builders, and anyone tracking AI safety/governance dynamics.
What was covered
-
US-China AI diplomacy & verification: Jeremie and Edouard Harris (Gladstone AI) assessed the Trump-Xi meeting and prospects for an AI "red phone." Key takeaways: historical nuclear hotlines often go unanswered during crises; the most realistic near-term verification mechanism is thermal/satellite detection of large data centers; verification startups need IC pre-vetting years before a crisis, not after. Ed Harris floated an idea: use the OpenAI-Anthropic competitive dynamic as a test bed for trust-and-verify mechanisms that could later scale to a US-China framework.
-
Compute economics — hyperscaler GPU pricing: Steve Hou (Silicon Data, which builds GPU price indexes) confirmed hyperscalers consistently charge 2–3× or more vs. neo-clouds for apparently identical H100 capacity. He attributed this to product bundling, software/compliance stacks, and enterprise customer stickiness — not raw chip performance differences. Silicon Data's proprietary LLM expenditure-weighted index dropped ~60% (from ~$4 to ~$1.60 per million tokens) from mid-June through mid-July, then bounced slightly, driven by usage-mix shifts toward more expensive frontier models rather than listed price changes.
-
OpenAI Dev Day takeaways: Co-host Prakash Narayanan highlighted two new models — GPT 6.1 SOL (roughly GPT-6 SOL quality at slightly lower price) and GPT-6 Astra Ultrafast (8× speed vs. prior "fast" mode, 2×). Ultrafast enables real-time interactive software co-building (e.g., game development in-flow). "Sign in with ChatGPT" for third-party apps was flagged as a significant ecosystem move reducing token-cost friction for user onboarding.
-
AI-assisted novel writing: Author Joel Borgen co-wrote The Receipt Horizon using Claude and ChatGPT Pro (GPT-5 onward as primary workhorse). Key finding: context-length expansion was the biggest unlock; prose still requires heavy human editing to remove "AI ticks"; model musical score reading went from 0% to near-perfect overnight with OpenAI's Astra after training on musical scores.
-
Gamow Labs — AI genomic diagnostics: Founder Daniel McKinnon described using AI agents (initially OpenAI o3, now multi-model ensembles) to reanalyze unresolved rare-disease genomes. His pipeline rediscovered a 91-kilobase deletion 1 million bases upstream of a gene that a top prenatal lab had filtered out. Vanilla Claude Opus 5.5 in Claude Code scores ~50% on their RareBench variant-ranking benchmark vs. ~10% for the best traditional ML tool (Lyrical). Gamow recently acquired a biology lab to run functional studies and close the RL feedback loop between wet-lab data and model predictions.
-
Model harness/evals signal: McKinnon detailed a failure mode in Grok 4.6 (then state-of-the-art on RareBench) where the model hallucinated a gene name mid-reasoning. Gamow's harness caught and prevented this. Astra initially stalled entirely because tool-call behavior was misconfigured. Both fixed via trace analysis and harness updates.
-
Pacing AI progress — ethical cost: Nathan Labenz closed with a reflection: the 50% RareBench score means half of unresolved rare-disease cases remain unanswered today, and deliberate AI pacing has a real human cost that should not be treated as zero.
Notable claims & predictions
-
Edouard Harris (Gladstone AI): "Certain kinds of transparency can be stabilizing — giving China enough vision into what we're doing to see we are not doing the thing they fear most — and may not be costly, because the Chinese are already all up in our systems." Implies reciprocal transparency could be a low-cost confidence-building measure.
-
Jeremie Harris (Gladstone AI): In a post-incident scenario with no prior verification infrastructure, the only available ask may be "turn off all clusters above a certain size" — describing this as "an insane ask" that would cost billions in GPU depreciation, making pre-vetted verification technology now an urgent economic and strategic priority.
-
Daniel McKinnon (Gamow Labs): "O3 is the first model where AI-driven genomic interpretation became possible at all — it scores 0% on RareBench; Opus 5.5 vanilla scores ~50%, blowing out the best traditional ML tools at ~10%." Suggests the interpretation layer that blocked precision medicine for 25 years post-Human Genome Project may now be solvable.
-
McKinnon on vertical AI longevity: "Vanilla o3 would not do this task. Now [frontier models] are 20–30 percentage points better. That layer [of scaffolding] is definitely shrinking." Candid acknowledgment that the harness/tools moat erodes as base models improve.
-
Prakash Narayanan on Dev Day: "Latency at the same intelligence is probably the most important variable as you clear capability hurdles — can the model build a game with you in the moment while keeping you in flow?" Frames ultra-low-latency inference as the next major applied-AI unlock.
-
Steve Hou (Silicon Data): Proprietary LLM expenditure-weighted index fell more than 50% from mid-June to mid-July 2025, then bounced — driven by usage-mix shifts, not listed price cuts. Meta and xAI (Grok) pricing aggression is pulling the blended market rate down.
Why this matters for AI operators
-
Inference pricing is still highly heterogeneous: The 2–3× hyperscaler premium over neo-clouds is persistent and structural, not arbitrageable in the short term due to enterprise lock-in and bundled services. Operators choosing cloud vendors for inference workloads should model this explicitly; the blended token cost market dropped 50%+ in ~six weeks, then rebounded on usage-mix shifts — budget forecasting for AI workloads requires tracking both listed prices and consumption mix.
-
Vertical AI harness economics are real but eroding: Gamow Labs' RareBench data quantifies the moat: traditional ML tools at 10% vs. vanilla frontier models at 50%, with Gamow's harness adding further lift. But McKinnon explicitly acknowledges the gap is shrinking 20–30 percentage points per model generation. Applied-AI builders should plan for the scaffolding layer to thin
Full analysis
Daniel McKinnon's Gamow Labs is the thing worth stopping on this week. His team points AI agents at the genomes of kids whose rare diseases the standard clinical labs gave up on, and it is finding answers. On their own test, called RareBench, the old-school software tools score about 10%. Claude Opus 5.5, straight out of the box with no special tuning, scores about 50%. That is a jump most of us have never seen in a vertical this hard.
The second thread that matters: renting the same NVIDIA H100 chip from AWS, Azure, or Google costs two to three times what it costs from a smaller specialist cloud. Same silicon. That gap is a bill most buyers are quietly paying.
What's being decided. Nothing, for the reader, is hard to undo here. There's no shutdown date, no contract cliff. What the episode actually surfaces is two slow-moving questions you can act on at your own pace: how much are you overpaying for compute, and how long does a hand-built AI "harness" (the software wrapper that forces a model to use tools properly and checks its work) stay valuable before the base model swallows it.
The Skeptic. Be careful with RareBench. It is Gamow's own test, scored by Gamow, with no outside party checking it. A 50% that beats 10% is a great slide, but McKinnon tells on himself: the gap between a plain model and his scaffolded version is shrinking 20 to 30 points per model generation. So the moat is real today and visibly melting. And the 50% cuts the other way too. Half of these desperate cases still get no answer. That is a research tool with promising hit rates, not a diagnostic you hand a frightened parent. Also note the model names flying by here, GPT-6.1 SOL, Astra Ultrafast, Opus 5.5, Grok 4.6, are the podcast's own shorthand, not products you can go buy under those labels today.
The Researcher. Start with the failure McKinnon caught, because that is more instructive than the headline score. Grok 4.6, their best performer at the time, silently swapped one gene name for another mid-reasoning and nearly shipped a wrong diagnosis. His harness caught it. This is what "vertical AI" actually buys you right now: not raw smarts, the base model has those, but a wrapper that catches the quiet, domain-specific lies a general model tells. The gene found 1 million base pairs upstream of its target, far outside the roughly 1,000-base window most labs even look at, is the real result. The model wasn't smarter than the lab. It just didn't throw away the data the lab's filter discarded.
The Compute Pragmatist. The 2 to 3x hyperscaler premium is the most bankable thing in this episode. Steve Hou of Silicon Data, which builds price indexes for rented GPU capacity, says it plainly: you pay the premium for bundled software, compliance paperwork, and the fact that you're already a customer, not for a faster chip. If your AI workload is batch work, overnight jobs, model evaluations, anything that doesn't need to sit inside your existing cloud's security setup, you are leaving real money on the table by defaulting to AWS. The other number: blended token prices fell from about $4 to $1.60 per million tokens in roughly six weeks, then bounced. But Ho is honest that the bounce came from people using pricier frontier models, not from anyone raising list prices. Costs are still falling. Your bill may not be, because you keep reaching for the better model.
The Open-Source Advocate. Quiet but important: the big wins here run on closed frontier models, Claude, OpenAI's o3, Grok. No open-weight model is named as carrying the diagnostic load. For a field that loves to say the open models are "within 80%," rare-disease genomics is a domain where the last 20% is a kid's diagnosis, and the builders reached for the closed labs anyway. The reusable lesson is the harness-plus-tests layer McKinnon built. That's the part you own regardless of which lab wins, and it is the part worth investing in.
The Builder. What would I actually do Monday? Two things. First, run a cost test: take one batch workload off your hyperscaler and price it on a specialist GPU cloud. If Ho's 2 to 3x holds for your shape of job, that's found money with an easy rollback. Second, steal Gamow's pattern. The durable value they built is a wrapper that forces the model to use real tools and validates every output against ground truth before anyone acts on it. McKinnon's harness caught a hallucinated gene name that would have sailed through a raw API call. Whatever you're shipping, the thing that keeps paying off is the checking layer around the model, because the base model keeps getting better but still needs someone checking its work.
Where they split. The Skeptic and the Builder look at the same harness and see opposite futures. The Skeptic reads McKinnon's own admission, the gap shrinks 20 to 30 points a generation, and concludes the scaffolding is a depreciating asset. The Builder says the checking layer is exactly what survives, because catching domain-specific lies is a permanent job no base model eliminates. Both are right about different parts: the clever prompts and workarounds erode fast, the validation-against-truth layer compounds. The Compute Pragmatist and the Open-Source Advocate part ways on where the leverage is. One says save money on the chips. The other says own the harness, because the model choice barely matters. For most readers, the chip savings are available this quarter and the harness is the long game.
What it hinges on. Two beliefs. One, that the hyperscaler premium is structural and will not close on its own, which Ho argues convincingly. Two, that the validation layer around a model holds value as base models improve, which McKinnon half-confirms and half-undercuts. Verify the first yourself with a single priced-out batch job. Treat the second as true for the checking layer and false for the clever-prompt layer.
Prediction: Silicon Data's GPU price index will still show hyperscaler H100 rental rates at 2x or more above neo-cloud rates when they publish their index covering Q1 2027.
Confidence: Medium. The premium is structural lock-in, held in place by bundled compliance tooling and enterprise relationships, not by any chip advantage the hyperscalers actually have.
Why: Steve Hou attributes the 2 to 3x gap to bundled compliance tooling, software, and enterprise relationships that keep customers from switching, not to any chip advantage, which means the premium persists precisely because it is not about the hardware. For it to close, AWS, Azure, and Google would have to cut prices into a segment where their customers are stuck anyway, and companies do not discount captive buyers. The opposite outcome, the gap closing, would require the hyperscalers to decide their enterprise lock-in is worth less than a price war against smaller clouds they currently out-earn, which runs against every incentive they have.
Revisit by 2027-04-06: We're right if Silicon Data's published index for Q1 2027 shows hyperscaler H100 rates at 2x or more over neo-clouds. We're wrong if the ratio falls below 2x.
Comments