Refacto AI

Podcast episode

Agent Wars!

agent-framework agents antitrust evals model-pricing

Meta's Muse app hit #1 on the US App Store, knocking ChatGPT off the top spot. It's a personal shopping agent (software that browses and buys on your behalf) that reads your email, calendar, and credit card to buy groceries, cancel subscriptions, and book services. Amazon immediately blocked it. Shopify welcomed it. That split is the episode.

The fight over platform access is where the real money is. Amazon runs a $76 billion ad business built on humans scrolling past sponsored listings; an agent skips every one. So the emerging deal structure is access fees or revenue share, which favors Meta, OpenAI, and Google, who can negotiate and absorb the cost. Smaller agents get blocked or priced out. Separately, Treasury Secretary Scott Bessent ruled out a government liability shield for AI labs, which means labs will push legal risk down into your deployment contracts. Check that indemnification clause.

The question everything rests on: do Muse users actually complete purchases at scale, or is this a chart-topping download that nobody trusts with a credit card? Amazon's reaction suggests the platforms think it's real.

Full analysis

Meta's Muse just did the thing everyone said a personal agent would never do: it hit #1 free on the US App Store and knocked ChatGPT off the top spot. The agent buys your groceries, cancels your subscriptions, and books your services using your email, calendar, and credit card. Amazon promptly blocked it. Shopify signed a deal to welcome it. That fight, over who lets an agent shop and at what price, is the real news here. The Grok 4.7 launch and the liability policy signal are the supporting cast.

Here's what it means for anyone building or buying AI. This is easy to undo for you as a buyer today. It is very hard to undo for the platforms deciding their access rules right now.

The Skeptic

One app hitting #1 free proves people downloaded it, not that they trust it with their credit card. Free App Store ranks are a vanity metric. The question nobody answered: how many Muse users actually let it complete a purchase without watching every step? Adam Foroughi of AppLovin nailed the doubt. On a $50 order, saving 20% doesn't matter because the shopper enjoys the shopping. Agentic commerce works for the boring reorder of paper towels and dog food. It does not work for the browse-and-discover buying that drives most retail margin. And Muse reading your email, calendar, and card at once is a single account takeover away from a very bad week. That risk lands on the user, not Meta.

The Compute Pragmatist

Follow the money and Amazon's block makes perfect sense. Amazon runs a $76 billion-a-year ad business built on humans scrolling past sponsored listings. An agent skips every one of those ads and buys the cheapest match. Skipping sponsored listings is not a rounding error; it cuts directly into the core P&L. So the emerging deal shape is access fees or revenue share to let an agent transact on your platform. That structurally favors the biggest agent providers. Meta, OpenAI, Google, Anthropic can negotiate and eat the cost. A startup agent cannot. If you are building a consumer shopping agent on a small team, price in that half your target inventory may charge you for the privilege, or block you outright, the way Amazon just sued Perplexity for routing around its blocker.

The Researcher

Grok 4.7 is the case study in why you never buy on a benchmark alone. xAI claimed a 6-point gain on Cursorbench (a coding test) and a strong score on AA Briefcase, a benchmark for multi-hour white-collar work. Then developer Theo actually ran it: 30 to 80% worse token efficiency, meaning it burns far more of the paid text units to do the same job, landing at more than 2x the real cost of Grok 4.6. Artificial Analysis, an independent tester, ranked it seventh overall and fourth on coding agents. Elon Musk himself conceded third place behind OpenAI and Anthropic. A model can climb a leaderboard and still cost you double in production. Test on your own workload before you sign anything.

The Enterprise Buyer

Scott Bessent, the Treasury Secretary reportedly running US AI policy, said labs will not get a government liability shield. His line: a lab claiming a 10% chance of an extinction-level event while asking the government to absorb its legal risk is "good business for them, bad business for the American people." Read what that does to your contracts. If the labs can't offload liability onto Washington, they will push it down to you in the terms of service. When you deploy an agent that spends money or takes actions for a customer, and it buys the wrong thing or leaks a card, the indemnification language decides who pays. Check that clause now. The regulatory cover some deployment plans assumed is not coming.

Tensions

Three real disagreements. The Skeptic says agentic shopping is a narrow reorder use case that won't dent real retail; the Compute Pragmatist says Amazon's $76 billion defensive crouch proves the platforms think it's an existential threat. Both can't be right about the size. If it's niche, Amazon overreacted. If Amazon's right, the Skeptic is underrating the shift.

Second: the Researcher wants you to distrust every benchmark after Grok 4.7, but buyers still have to choose something. The answer isn't paralysis, it's your own test harness on your own tasks. Vendor numbers are marketing until you reproduce them.

Third: the access fights favor the giants, which means the open, build-it-yourself agent story gets harder in commerce specifically. You can run an open-weight model on your own hardware all day. You still can't force Amazon to let it check out.

What it hinges on

Does agentic commerce actually convert at scale, or is it a download-and-abandon novelty? That's the one belief everything rests on. If Muse users complete real purchases at real volume, Amazon's block was rational and every ad-funded marketplace has to pick a side. If they don't, this is a chart-topping toy and Shopify's bet ages badly. Before you build anything commerce-facing on an agent, run the boring test: can it complete ten real checkouts across five merchants without a human touching it, and how many of those merchants blocked or charged you?

Prediction: Before OpenAI's next Dev Day, at least one more major commerce platform besides Amazon (Walmart, Target, eBay, or Etsy) will publicly restrict or block unauthorized AI shopping agents from transacting on its site.

Confidence: Medium -- the ad-and-discovery revenue math that drove Amazon's block applies to every large marketplace.

Why: Amazon blocked Muse to protect a $76 billion ad business that depends on humans seeing sponsored listings, and every large marketplace monetizes discovery the same way, so the incentive to block is not unique to Amazon. Meta now has a #1 app pointing agents at that inventory, and Shopify just declared the opposite side by welcoming Muse, which forces every other platform to pick. The marketplaces with their own ad businesses (Walmart Connect, eBay, Etsy's promoted listings) lose the most from agents that skip ads and buy the cheapest match, so restricting is the self-interested move. The opposite outcome, everyone opening up like Shopify, only happens if a platform decides catalog access beats ad revenue, which the ad-dependent players have no reason to believe yet.

Revisit by 2027-03-28: We're right if Walmart, Target, eBay, or Etsy publicly restricts or blocks unauthorized AI shopping agents in its terms of service or via technical blocking. We're wrong if none of them does and no comparable large marketplace follows Amazon.

Comments