Podcast episode
From Two Years to Two Weeks: Data Readiness
ai-in-adtech cloud-costs data-readiness engineering
On the Aperiam podcast, Datalinx AI founder Joe Luchs told hosts Corey Ferengul and Joe Zawadzki that the thing blocking every AI-in-marketing project is not the model. It's the data. Luchs claims over 90% of the work is data engineering and feature engineering (cleaning, structuring, and labeling data so a model can use it), and that Datalinx can compress what normally takes two years down to two weeks by running entirely inside your existing Snowflake, Databricks, or AWS environment.
Luchs cited a LiveRamp result: two weeks of work down to two minutes at 99.7% accuracy. One partner data point, inside a program Datalinx belongs to. The more durable signal is that Databricks invested and both Snowflake and AWS gave Datalinx accelerator and marketplace slots rather than building this themselves. The hyperscalers are choosing to be the toolkit, not the finished answer.
The disease is real: one large CPG Luchs mentioned burned $10 million in systems-integrator fees on a two-year data-readiness project before writing a line of model code. Budget for that problem. Treat any compression as upside, not the plan.
Full analysis
On the Aperiam podcast, Datalinx AI founder Joe Luchs told Corey Ferengul and Joe Zawadzki that the thing blocking every AI-in-marketing project is not the model. It is the data. Over 90% of the work is data engineering and feature engineering. The model is the last 10%. Datalinx claims to run inside your Snowflake, Databricks, or AWS environment, no data leaving the building, and to turn a two-year data-readiness slog into two weeks.
That claim is a vendor pitch. But the bottleneck it points at is real, and it changes how operators should read every "AI-powered" roadmap on their desk right now.
Here is what's actually being decided by operators reading this: not "should I buy Datalinx." It's "when I greenlight an AI yield, personalization, or measurement project in 2026, do I budget for the model, or for the two years of data plumbing underneath it?" That's a Type 1 decision. Get the diagnosis wrong and you burn a year and a systems-integrator invoice before you learn the model was never the problem.
The Skeptic. Luchs sells data-preparation software, so of course he says data prep is 90% of the work and slow and human-driven. His whole case rests on the premise that timelines "have barely shortened despite better tooling." That's contestable and self-serving. Databricks Genie and Anthropic's code tools, which Luchs himself name-checks, exist precisely to compress this work, and large enterprises are already using them. The "two years to two weeks" line has no third-party case study behind it. The LiveRamp result, two weeks to two minutes at 99.7% accuracy, is one partner data point inside a program Datalinx belongs to. In plain terms: the disease is real, the miracle cure is unproven.
The Market Analyst. The interesting structural signal isn't Datalinx. It's that Databricks put money in and Snowflake and AWS gave it accelerator and marketplace slots rather than building it themselves. The hyperscalers have decided the domain-specific ontologies for advertising and commerce, the rules that turn messy marketing data into deterministic outputs, are worth outsourcing to specialists. For an informed outsider: the platforms that store everyone's data are choosing to be the toolkit, not the finished answer, and they're seeding startups to fill the gap. That's a repeatable pattern. Identity, clean rooms, and measurement vendors that deploy inside the warehouse instead of demanding data export will keep clearing security review in days rather than months. That deployment architecture is becoming a moat, and vendors still asking customers to ship data out are on the wrong side of it.
The Operator. Tuesday morning, the person running this doesn't care about the ontology. They care that InfoSec signs off, that the project draws down the Snowflake commitment already on the books, and that the data actually maps to something usable. Warehouse-native deployment genuinely helps the first two. But Luchs said the quiet part: marketplace reps will not sell your product for you. Drawing down cloud spend is a procurement win, not a magic go-to-market. And "two weeks" assumes the customer's data is coherent enough to compress. If your event data is a swamp of inconsistent identifiers, no application fixes that in a fortnight. The 90-day surprise is that the tool worked and the org still couldn't agree what a "customer" was.
The CFO. The one number worth carrying out of this episode is the 10 million dollars in systems-integrator fees a single large CPG racked up on a two-year data-readiness effort, and that figure didn't even include back-end observability and monitoring. That's the benchmark to hold every AI vendor against. If a warehouse-native app credibly takes a slice out of that, the payback math is easy, because the alternative is a seven-figure SI engagement that runs for years. The trap is treating "two weeks" as the plan and budgeting accordingly. Budget for the SI-sized problem, and treat any compression as upside.
The Long-Term Thinker. Luchs floated the real endgame: "why do I really need a platform just to send emails based on rules?" If agents do the work, some marketing software layers stop earning their seat. Three years out, the companies that win are the ones sitting closest to clean, warehouse-resident data with the domain logic to act on it. That's why this matters beyond one pitch. The value migrates from the application layer down to whoever owns the data-readiness step. Zawadzki's InfoSum parallel lands here: the data-local "bunker" idea was right and too early, because deployment was too expensive. Now the warehouse makes it cheap. Same idea, better timing.
Where the council splits. The Skeptic says the two-year problem is already shrinking on its own, so any "two weeks" vendor is selling urgency that's expiring. The Long-Term Thinker and Market Analyst say the opposite: the platforms just voted with their balance sheets that this stays hard enough to need specialists. Both can't be right. The second disagreement is the Operator versus the CFO on where the money goes. The CFO wants to budget for the two-year problem to be safe. The Operator warns that the real failure isn't cost, it's organizational: nobody can agree what the data means, and no software vote settles that.
What it hinges on. One belief: has AI-assisted tooling actually shortened enterprise data-readiness timelines, or not? If Luchs is right that it's still one-to-two years of human-driven work, warehouse-native specialists with marketing ontologies are a real category and the hyperscalers' bets pay off. If the Skeptic is right and Genie-class tools are already collapsing the timeline, Datalinx and its peers are selling a window that's closing. Before committing, an operator should run its own bake-off: take one live data-readiness task, run it through the warehouse's native AI tooling and through a specialist app, and measure the delta honestly. Don't buy the tagline. Buy the test result.
Impact of this specific episode on the broad ad-tech audience is moderate and indirect. No financials, no M&A, no regulatory move. The durable takeaway is the diagnosis, not the vendor: your AI roadmap will stall on data engineering, not model quality, and the vendors deploying inside your warehouse are the ones to watch.
No high-conviction prediction this week.
The episode is a founder origin story, not a market event. The one testable claim here, "two years to two weeks," rests on a single partner data point and a vendor with every incentive to inflate it. There's no earnings call, no deal, and no regulatory anchor that would settle any Datalinx-specific bet in the next six months. Making a call that reaches past this topic to something else would mean predicting on something the episode barely touched, which is worse than no call at all.
Comments