Podcast episode
The Biggest AI Deployment Nobody Talks About | Samsara CEO Sanjit Biswas
cost-compression edge-inference inference open-weights reliability
Sanjit Biswas, CEO of Samsara, sits down with Matt Turck to talk about what running AI at genuine scale actually looks like: millions of commercial vehicles, 25 trillion sensor data points a year, and dash cams doing on-device inference (processing video locally, without sending it to the cloud) across 99% of US roads daily. This is not a startup demo.
The two most useful threads: Biswas runs small, proprietary models on the device and routes the hard reasoning to cloud-based vision-language models (VLMs that read video and text together). That distill-and-deploy split is the reusable idea. He also benchmarks Cerebras, an alternative chip architecture, at 800 to 1,500 tokens per second against roughly 100 on a fast GPU. Impressive, but those numbers are single-stream. Serving thousands of concurrent requests looks different.
The durable point is the data. Twenty-five trillion physical-world sensor readings is something no competitor can scrape. That is the same logic as first-party audience data surviving the cookie. The model is a commodity; the data is the moat.
Full analysis
Sanjit Biswas runs the biggest physical-world AI deployment almost nobody in the model-benchmark discourse talks about. Samsara: millions of vehicles, 25 trillion data points a year, dash cams doing edge inference on 99% of US roads daily. The episode is a field report on what actually gates AI in production. Not capability. Cost and reliability at inference time.
For a team shipping AI, this is a Type 2 read most of the way through. Nothing here forces a decision this quarter. But two threads are Type 1 in disguise: where you place your inference (GPU vs. wafer-scale), and whether you're accumulating proprietary data that survives when the frontier models commoditize. Those compound. Get them wrong and you're repricing your whole stack in a year.
The Skeptic. The 10x Cerebras number is a vendor demo until you run your own traffic through it. Biswas cites Gemma 4 at ~100 tokens/sec on a fast GPU versus 800 to 1,500 on Cerebras. Great. Ask what batch size, what context length, what concurrency. Wafer-scale wins on single-stream latency and falls apart on the economics when you need to serve a thousand concurrent requests, which is exactly the ad-tech shape. The grid-tripling stat is also secondhand, one utility exec, unverified in the episode. And "prevented 380,000 crashes" is a counterfactual nobody can audit. For a PM: fast on a slide is not fast on your bill.
The Researcher. The interesting architectural tell is the split. Small models on-device (CNNs, a JEPA-style video model, tens-of-millions-of-parameter proprietary nets), VLMs (vision-language models that read video and text together) in the cloud for the hard reasoning. Biswas distills from open weights down to the device. That distill-and-push pattern is the real reusable idea, not the token count. His agent honesty is the part worth stealing: the ceiling isn't whether the agent finds the answer, it's whether it finds it in ten seconds versus an hour, and whether it loops. Reliability and latency, not IQ.
The Open-Source Advocate. This is the quiet win for open weights. Samsara does ~$2B ARR, is profitable, and runs OpenAI, Anthropic, Google APIs and open-weight models and trains its own from scratch, all at once. Nobody's locked in. Gemma 4 is the model they benchmark on novel hardware, because you can't put a closed API on a Cerebras wafer you control. For a PM: the frontier API is where you prototype, the open model is what you own and move. The map onto ad-tech is direct, this is the same multiplicative logic as bid factors: you keep the levers you can retune, you rent the ones you can't.
The Compute Pragmatist. The binding constraint here is cost-per-token, full stop. Full-shift driver coaching via video tokenization "works and it's doable today," just too expensive, and Biswas bets the curve makes it viable in one to two years. That's a bet on inference deflation, and it's a good one. The grid signal underneath it matters more than the crash stat: if a utility really triples capacity in five years versus 125 prior, 90% data-center-driven, then power, not chips, is the ceiling. The alternative-silicon story (Cerebras, Groq, the rest) is a real hedge against NVIDIA rent, but only for the workloads shaped like theirs. Batch, latency-sensitive, single-stream. Real-time auction serving is not that shape.
Where they part ways
The Skeptic and the Compute Pragmatist split on the 10x. One says prove it on your concurrency; the other says even a 3x real-world gain reprices your roadmap. Both right, and the resolution is a load test, not an argument.
The Researcher and the Open-Source Advocate agree on the distill-and-push pattern but disagree on the moat. The Researcher says the architecture is copyable by anyone. The Open-Source Advocate says the weights are, but Biswas's actual point is the data isn't. "You can't crawl Reddit and find out what happened on a construction site." Twenty-five trillion sensor points is the thing that doesn't commoditize.
That's the hinge. Every AI team is downstream of the same three or four frontier APIs. The differentiator is not the model. It's whether you sit on data nobody else can scrape. In ad-tech you already know this instinct, it's why first-party data survived the cookie. Same mechanism, physical sensors instead of pixels.
What it actually hinges on
Three beliefs. One, does non-GPU inference give your traffic a real multiple, not a single-stream demo multiple. Two, does inference cost fall enough in 12 to 24 months to move ambient use cases from "doable but too expensive" to shipped. Three, are you accumulating proprietary data that holds value when the base models are free.
The council leans clear on two and three: yes on the cost curve, yes on data as the durable edge. It's split and unproven on one. So before anyone commits budget to alternative silicon, run your own concurrency-matched load test against a GPU baseline. The number that survives that test is the only one worth planning around.
Prediction: By the time Samsara's next Beyond conference lands (roughly June 2027), full-shift AI driver coaching via continuous video tokenization will be a generally available, priced Samsara product, not a "doable but too expensive" preview.
Confidence: Medium. Inference cost curves reliably clear the gap Biswas named.
Why: Biswas said the feature works today and the only blocker is cost-per-token, and he put the timeline at one to two years himself. Inference prices for open-weight and distilled models have fallen faster than that window every year running, and Samsara controls its own on-device stack plus alternative-silicon options, so it isn't waiting on a frontier lab's pricing. The opposite outcome, that it stays a preview, would require inference costs to stall, which nothing in the last three years of the market suggests. The main risk to the call is packaging: they could ship the capability inside a bundle rather than as a named SKU, which would make it technically true but harder to score.
Revisit by 2027-06-30: We're right if Samsara lists continuous full-shift video coaching as a shipping, priced product. We're wrong if it's still gated as a pilot, preview, or "contact us" feature.
Comments