Refacto AI

Podcast episode

The New Enterprise Battle Over Who Owns the Model

build-vs-buy fine-tuning inference model-pricing open-weights

TL;DR

The episode centers on Thinking Machines Lab's new open-weight model Inkling (975B total / 41B active parameters, MoE architecture) and its fine-tuning platform Tinker as a potential enterprise alternative to closed-model providers. Host Nathaniel Whittemore frames this alongside Microsoft's move to sell against OpenAI and Anthropic using its own MAI models, Cursor's ambitions to become a frontier model developer, and Apple's search for an AI server chip acquisition.


What was covered

  • Thinking Machines Lab (TML) releases Inkling: A 975B-parameter mixture-of-experts (MoE) model — where only 41B parameters activate per inference pass — with a 1M-token context window and native multimodal (text, image, audio) support. Bootstrapped from Kimi K2.5 synthetic data but otherwise pre-trained from scratch. Paired with TML's Tinker fine-tuning API, the pitch is enterprise data sovereignty: companies tune the model on their own data, on their own infrastructure, with no leakage to the model lab.

  • Inkling benchmark positioning: Mid-tier open-weight performance. Humanity's Last Exam: 29.7% (no tools) / 46% (with tools) — ahead of Nemotron Ultra, roughly matching Kimi K2.5, ~10 points behind Kimi K2.6 and GLM 5.2, ~20 points behind Fable 5. SWE-Bench Pro coding: 54.3%. Artificial Analysis intelligence index: score of 41, ranked 19th globally. Notably strong token efficiency — ~2/3 the tokens per task vs. Kimi 2.6/DeepSeek V4 Pro.

  • Microsoft arming sales teams against OpenAI and Anthropic: Bloomberg reports Microsoft instructed sales staff in a department-wide meeting to emphasize MAI model efficiency and cost advantages. Initial focus is selling against Claude, with one exec reportedly telling reps Claude is "slower and less accurate" in Microsoft's Office suite. Microsoft has already begun switching some Copilot functions to MAI models as a cost-cutting measure. Satya Nadella's public messaging frames closed-model providers as data security risks to enterprises.

  • Cursor/xAI model ambitions: CEO Michael Trewell told staff at an all-hands in May that Cursor aims to be a "top-tier model developer" beyond coding, targeting a state-of-the-art model by end of 2024 and a "significant compute advantage" by 2027. The first joint Cursor/xAI output is Grok 4.5, described as competitive especially for advanced architectures. Cursor is also rumored to be building a Claude Code competitor.

  • NVIDIA Cosmos 3 Edge: A 4B-parameter small model for edge robotics hardware (Jensen platform), functioning as both a world model and a vision-language model (VLM). Part of NVIDIA's expanding physical AI push, including an expanded Toyota partnership covering robots, smart cities, and automated factories.

  • Apple hunting for AI server chip acquisition: Apple's server-grade chip (codename Baltra) has been delayed; the company currently relies on M2 Ultra chips for AI servers and outsources capacity to Google Cloud on NVIDIA hardware. Apple has reportedly approached semiconductor startups. Candidate pool is thin — NVIDIA acquired Groq for $20B, SiPearl valued at ~$40B, Broadcom at $1.8T would be a merger not an acquisition. Tenstorrent (CEO Jim Keller, formerly Apple) mentioned as a natural fit.

  • Anthropic IPO on track for fall: Bloomberg reports Anthropic has appointed investment banks, is meeting investors, and has added a revolving credit facility of several billion dollars. Target window: September–October. OpenAI, by contrast, is reportedly unlikely to hit Altman's $1T market cap target and is now leaning toward a 2026 IPO.


Notable claims & predictions

  • Jack Morris (researcher): "This is the only open-weight model that's trained without distilling from OpenAI or Anthropic. Kimi distills, GLM distills, Qwen distills, Nemotron distills… basically a fully different tech stack." (Note: Community corrections acknowledged a small Kimi K2.5 bootstrap and cited Llama 3.1 as another non-distilled precedent.)

  • Simon Smith (ML practitioner): "Thinking Machines is basically a bet against the bitter lesson… fine-tuning is way more effort than people think. That effort is ongoing to address new data edge cases and model updates. Models can lose capabilities… and ultimately a big general model with a bit of context comes along and beats your hard work." — A direct challenge to the fine-tuning-as-moat thesis.

  • Sriram Krishnan (a16z, former White House AI advisor): "Organizations are increasingly looking for control over how their data is used and are willing to trade off some access to frontier-level tokens for this control… companies have now actively shifted from 'how do we get our people to use tokens' to being uncomfortable with their token cost ballooning without a clear line to revenue."

  • Nathaniel Whittemore (host): "A few months ago, one would be forgiven for thinking that pretty much every enterprise was just asking whether it should sign up with OpenAI or Anthropic. Now we are talking about so much more choice — not only in terms of models but in terms of the harnesses in which the model lives."

  • Microsoft executive (unnamed, per Bloomberg): Instructed sales staff to describe Claude as "slower and less accurate" with insufficient security integrations for Microsoft Office — framing the MAI push as a Microsoft-stack integration advantage rather than a raw capability argument.


Why this matters for AI operators

  • The enterprise model-ownership battle is real and accelerating. Microsoft selling against its own model-lab partners (OpenAI, Anthropic) while simultaneously pushing Frontier Tuning, and TML offering an open-weight fine-tuning alternative with Inkling/Tinker, signals that the enterprise procurement question is no longer just "OpenAI or Anthropic" — it's becoming "who holds the weights and the trained data." Operators evaluating multi-year AI contracts need to factor data-residency and competitive-exposure risk explicitly.

  • Fine-tuning economics are genuinely contested. Simon Smith's critique — that fully loaded fine-tuning costs (data curation, pipeline maintenance, infra ops, model-drift management) are routinely underestimated — is a material risk for any enterprise buying into the fine-tuning-as-moat narrative. Operators should model total cost of ownership, not just per-token inference savings, before committing to continuous fine-tuning programs.

  • Token efficiency is emerging as a differentiated metric. Inkling's ~2/3 token-per-task ratio versus Kimi 2.6 and DeepSeek V4 Pro matters directly to inference economics at scale. As token budgets balloon (Sriram Krishnan's point), models that are architecturally

Full analysis

The story underneath all five news items is the same one: the enterprise AI question has stopped being "OpenAI or Anthropic?" and become "who holds the weights, and who touches your data?" Thinking Machines Lab's Inkling (a 975B-parameter mixture-of-experts model where only 41B fire per query) plus its Tinker fine-tuning API is the purest expression of that pitch — tune on your own infra, no leakage to the lab. Microsoft is running the same play from the other direction, telling its own sales reps to sell against Claude with its cheaper MAI models. This is a Type 1, hard-to-reverse question for anyone signing a multi-year model contract, and a Type 2, cheap-to-test question for anyone just adding a fine-tuning pipeline to one workload. The forcing function is real: Anthropic's fall IPO and ballooning token bills are pushing buyers to re-examine lock-in now.


The Skeptic — The load-bearing claim here is that fine-tuning your own open-weight model is a moat. Simon Smith already torched it: the real cost isn't the tune, it's the forever — re-curating data for new edge cases, babysitting drift, re-tuning every time the base model updates, and watching a bigger general model with a scrap of context lap your hard work six months later. Inkling ranks 19th globally at a 41 intelligence index, ~20 points behind Fable 5. You're volunteering for a maintenance treadmill to run a mid-tier model. For the PM: "own your model" often means "own a second full-time engineering project that a frontier API would have made obsolete anyway."

The Researcher — The genuinely novel thing isn't the benchmarks, it's the provenance. Jack Morris's claim — Inkling is trained without distilling from OpenAI or Anthropic, unlike Kimi, GLM, Qwen, Nemotron — matters because a fully independent tech stack means independent failure modes and no legal contamination from a competitor's outputs. But read the correction: there was a Kimi K2.5 synthetic-data bootstrap, so "from scratch" is doing some marketing work. The number I'd actually bank on is token efficiency: ~2/3 the tokens per task versus Kimi 2.6 and DeepSeek V4 Pro. That's a real, measurable architectural win. For the PM: fewer tokens to finish the same job means a smaller bill, and that's verifiable, unlike "our fine-tune is a moat."

The Open-Source Advocate — This is the good timeline. An open-weight model at 46% on Humanity's Last Exam with tools, a 1M-token context, native multimodal, and a real fine-tuning API — you can run the whole thing on your own hardware. The GKE security-blueprint and model-migration pain that Google Cloud keeps publishing about? That pain is why holding your own weights is attractive: no forced migration when a vendor deprecates a model out from under you. But be honest about the license and the reproducibility before you build on it. For the PM: open weights mean the rug can't get pulled — the model you shipped on is the model you keep.

The Compute Pragmatist — Follow the chips and the money and the story tells itself. Apple can't ship its Baltra server chip, is renting Google Cloud NVIDIA capacity, and is chip-shopping in a market where NVIDIA just bought Groq for $20B. Cursor wants a "significant compute advantage by 2027." Everyone is racing to own silicon because inference at scale is where margins live or die. Inkling's 41B active params (not 975B) per pass is the whole point — MoE means you pay for a fraction of the model each query. For the PM: the sparse-activation trick is why a giant model can be cheap to run, and token efficiency compounds that — it's the difference between a sustainable inference bill and a runaway one.

The Builder — What ships Tuesday? Not Inkling as your primary model — it's 19th and you'd be maintaining it. But Tinker as a pipeline for one narrow, high-value workflow where data can't leave your walls? That's a real weekend prototype. The honest move is what's already happening in the room: someone upgraded Claude Haiku to Sonnet and got "meaningful quality improvement" for near-zero effort. That's the baseline every fine-tuning project has to beat. Microsoft swapping Copilot functions to MAI for cost, then telling reps Claude is "slower and less accurate" in Office — that's an integration argument, not a capability one. For the PM: the cheapest win is usually a better prompt or a bigger off-the-shelf model, not a custom tune.


Where they split:

  • Researcher vs. Skeptic on independence. The Researcher thinks a non-distilled, independent tech stack is a durable asset (clean IP, uncorrelated failures). The Skeptic thinks it's irrelevant if the model is 20 points behind frontier and you're stuck maintaining it — provenance doesn't pay the maintenance bill.
  • Open-Source Advocate vs. Builder on ownership. The Advocate sees weight ownership as insurance against forced migration. The Builder points out that the room's actual win came from renting a better closed model (Haiku→Sonnet) with zero ops cost — the exact opposite of the ownership thesis.
  • Compute Pragmatist vs. Skeptic on where the moat is. The Pragmatist says the real edge is inference economics — token efficiency and sparse activation. The Skeptic says none of that matters if a bigger general model plus a little context beats your tuned setup regardless of how cheaply it runs.

What it actually hinges on: two testable beliefs. First — does owned fine-tuning beat a big general model plus RAG on your specific task, and stay ahead across base-model updates? Second — is data sovereignty a hard requirement (regulated data, competitive exposure) or a nice-to-have you're paying a maintenance tax for? If sovereignty is a genuine constraint, Inkling/Tinker and Microsoft's Frontier Tuning are worth a bake-off. If it's not, the council leans hard toward the boring answer: rent the best closed model, add context, and revisit only when the token bill has a clear line to revenue that a cheaper open model would improve.

De-risk before committing: run the Skeptic's experiment explicitly — tune Inkling on your real task, then benchmark it against Claude/Gemini-plus-RAG on a held-out set, and re-run that comparison after the next base-model release to see if your tune's edge survives an upgrade. Model total cost of ownership including drift maintenance, not just per-token savings. Get license and indemnification terms in writing before any weight lives in production.


Prediction: By the end of Q1 2027 — one to two frontier release cycles out — no independent benchmark will show a general-purpose fine-tuned Inkling beating a top-3 closed model (Claude/Gemini/GPT) plus retrieval on a broad enterprise task; its wins will stay confined to narrow, data-sovereignty-constrained niches.

Confidence: Medium — mid-tier base plus the well-documented fine-tuning-drift trap.

Why: Inkling enters at 19th globally (index 41), roughly 20 points behind Fable 5 on Humanity's Last Exam, so the fine-tune has to close a large capability gap and then hold it. The mechanism working against it is the one Simon Smith names and the Google Cloud migration article confirms: base models keep improving, fine-tunes lose capabilities and require constant re-curation, and "a big general model with a bit of context" tends to catch up. For the tuned model to win broadly, TML's efficiency edge would have to outrun frontier capability gains between now and Q1 2027 — the less likely outcome given the pace of closed-model releases. The realistic win is narrow: regulated or competitively sensitive workloads where sovereignty is non-negotiable.

Revisit by 2027-03-31: We're right if fine-tuned-Inkling wins remain limited to sovereignty-driven niche deployments and no public eval shows it beating a top-3 closed model plus RAG on a general task. We're wrong if an independent benchmark shows a tuned open-weight Inkling matching or beating frontier closed models on a broad enterprise workload.

Comments