Refacto AI

Podcast episode

Who Feeds the GPUs? Inside AI's Hidden $30B Layer | Renen Hallak, VAST Data

cloud-costs gpu-supply inference open-weights privacy

Renen Hallak, founder and CEO of VAST Data, joined Matt Turck to make the case that the storage and infrastructure layer underneath AI is the actual bottleneck, and the actual business. The headline product is DataEnclave, built with NVIDIA: a system that lets a bank, hospital, or defense contractor run a frontier AI model on its own hardware without the model maker ever seeing the company's data, and without the company ever seeing the model's underlying weights (the mathematical parameters that encode everything the model knows).

The most concrete claim is that one unnamed customer went from 500 petabytes to 2 exabytes of storage in a single quarter. Hallak also says OpenAI solved a major unsolved math problem using 10,000 AI agents working in parallel, with two more close behind. No paper, no independent confirmation.

The storage infrastructure argument is plausible. The anecdotes are coming from the vendor who bills for them. And DataEnclave only matters if a named frontier lab actually commits to it, which none has yet.

Full analysis

Renen Hallak, founder and CEO of VAST Data, spent an hour making the case that the storage layer under AI is where the real money and the real bottleneck now live. The headline product is DataEnclave, built with NVIDIA: it lets a company run a frontier model inside its own data center without ever seeing the model's weights, and without the model maker ever seeing the company's data. That's aimed straight at banks, hospitals, and defense shops that have been frozen out of good models because they can't ship sensitive data to someone else's cloud.

What's actually being decided for the reader: whether to keep sending your proprietary data to OpenAI or Anthropic through their APIs, or to start planning for a world where the good model runs on hardware you control. This is a hard-to-undo architecture choice. Nothing forces it this quarter. But if DataEnclave-style deployment becomes real, it changes who you sign with and where your data sits.


The Skeptic

Hallak is selling storage, so of course storage is the hero of every story. Watch the demand anecdote: one customer went from 500 petabytes to 2 exabytes in a quarter. That's one customer, unnamed, told to you by the vendor who bills for it. Could be real. Could be a customer padding a reservation because capacity is scarce and they want to hold their slot.

Then there's the Navier-Stokes claim. Hallak says OpenAI "solved" a Clay Millennium math problem with 10,000 agents, and that two more are close. There is no paper, no verification, no Clay Institute award. Repeating an unconfirmed rumor as "the first concrete proof AI solves what humans can't" is exactly the kind of line that should make you tighten your grip on your wallet. The infrastructure argument might be sound. The hype around it is doing overtime.

The Builder

DataEnclave is the one thing here I'd actually act on. If you're in finance, health, or anywhere with data you legally can't move, the current answer to "can we use Claude on this?" is no. A confidential-computing box, co-signed by NVIDIA, with Cisco and Supermicro building the appliances, changes that answer to maybe. That's a real unlock.

But "multiple model builders joined as partners" is doing a lot of hoping. Nobody named OpenAI or Anthropic as committed. Encrypted weights running on your metal is a genuine engineering problem: key management, attestation that the box is what it claims, and who's on call when inference stalls at 2 AM inside your VPC instead of the lab's. Ask for the list of committed model partners before you budget a single dollar against this.

The Open-Source Advocate

The most useful idea in the whole episode is the one Hallak almost buries: your company's real value ends up baked into fine-tuned weights, and if you rent everything from OpenAI, you're handing that loop to them. That's correct, and it's the strongest argument for open models like Llama or Qwen or Mistral that you can fine-tune and own outright.

Here's the tension he skips. DataEnclave exists so you can run someone else's closed model on your hardware without seeing its weights. That keeps you dependent on the lab. If ownership of your fine-tuned intelligence is the goal, an open-weight model you actually control does more for you than a sealed box running Anthropic's parameters you're forbidden to inspect. DataEnclave solves the lab's trust problem, not yours.

The Compute Pragmatist

The line worth taking seriously: neoclouds are sold out 18 months forward, and the binding limit is land, power, and chip supply, not demand. If that's even half true, the price you pay for GPU time and storage isn't dropping soon, and your cloud vendor's roadmap is a queue, not a menu.

Hallak's swipe at AWS, Azure, and Google is self-serving but not empty. The neocloud pitch is that an AI-native stack wins workloads the hyperscalers can't match yet. VAST is profitable because it sells software margins, not boxes, which is rare in this layer and gives the claim some weight. For a buyer, the takeaway is dull but real: don't assume your hyperscaler contract is the cheapest or fastest path for heavy inference over the next two years. Price the neocloud option.


Where they disagree

The Open-Source Advocate and the Builder split on what DataEnclave is for. The Builder sees a market unlock for regulated buyers. The Open-Source Advocate sees a leash: it keeps you renting a closed model, just on your own floor. Both are right, and which one applies to you depends on whether you want to own your fine-tuned intelligence or just legally use someone else's.

The Skeptic and the Compute Pragmatist split on the demand numbers. The Skeptic says vendor anecdote. The Pragmatist says even discounted, "sold out 18 months forward" matches what everyone else in the supply chain is reporting, so the direction is trustworthy even if the specific exabyte figure isn't.

What it hinges on: whether real, named frontier labs commit to DataEnclave-style deployment. If OpenAI or Anthropic publicly ship a "run our model on your hardware, encrypted" option, this whole thesis has legs and the regulated-industry market opens up. If it stays "multiple unnamed partners," it's a nice demo that never leaves the slide.

What to verify before you act: ask VAST or NVIDIA which model makers have committed to DataEnclave in writing, not just at launch. Then price your heaviest inference workload on a neocloud against your current hyperscaler bill. That test costs you a spreadsheet and tells you whether the architectural-fit argument is real for your traffic.


Prediction: By NVIDIA's GTC 2027 keynote in March 2027, no top-tier frontier lab (OpenAI, Anthropic, Google DeepMind) will offer a generally available product that runs its flagship closed model on customer-owned hardware with weights sealed from the customer.

Confidence: Medium. The incentive to protect weight secrecy and API revenue runs directly against it.

Why: DataEnclave's pitch depends on frontier labs agreeing to ship their crown-jewel weights, even encrypted, onto hardware they don't control, and the launch named "multiple model builders" without naming a single frontier lab. The labs' whole business is metered API access, where they see every query and keep the weights locked in their own data centers. Letting an encrypted copy sit on a bank's floor weakens both the revenue model and the secrecy that protects a multi-hundred-million-dollar training run, so the ones with the most valuable models have the least reason to be first. The likelier outcome is that smaller or open-leaning model makers use DataEnclave as a wedge while the top labs watch, which is a real business but not the market-opening moment Hallak is selling.

Revisit by 2027-03-30: We're right if, by GTC 2027, none of OpenAI, Anthropic, or Google DeepMind has a generally available offering running its flagship model on customer-owned hardware with weights hidden from the customer. We're wrong if any one of the three ships such a product for general availability before then.

Comments