Featuring a Q&A with Domo Chief AI and Analytics Officer Ben Schein
For years, cleaning up customer data was the project that kept sliding into next quarter. The cost of skipping it was real but diffuse — an analyst losing a Thursday to deduplication, two reports that disagreed about how many customers the company actually had. Irritating. Never urgent enough to fund against a campaign.
AI changed the accounting. Not the underlying problem, which is the same one it’s always been, but where the bill lands and how often. Ben Schein, Chief AI and Analytics Officer at Domo, argues that sloppiness now shows up on the invoice — metered, every time a model runs. Asked where climbing AI spend actually comes from in a marketing organization, he didn’t talk about model pricing. He talked about how much of that spend buys the same cleanup work, over and over, forever.
The Reconciliation Tax Nobody Itemizes
The shift Schein describes is from labor cost to compute cost, and compute cost doesn’t take weekends off.
“Bad data used to cost you in analyst hours. Now it costs you per token, and the meter never sleeps. When your customer data is duplicated, inconsistent, or contradictory, every AI run pays a tax you never see itemized. You pay for the model to read three versions of the same customer and guess which one is real. You pay for retries when it hits a conflict and produces something nobody trusts. You pay for it to reconcile, on every single run, definitions your team never agreed on: what counts as a lead, a conversion, an active customer.”
That last item is the one with marketing’s fingerprints on it. Sales counts a conversion at opportunity creation, the web team counts it at form fill, finance counts it when the invoice clears. Nobody ever resolved the disagreement because it lived in a slide footnote and cost nothing to leave there. Point an agent at all three systems and it re-litigates that argument on every run, and you’re paying per token for the privilege — what Schein calls “renting intelligence at retail rates to redo janitorial work nightly.” Getting three teams to agree on one definition used to be a political errand with no budget line. It’s now a cost-reduction project, and it’s defensible in front of a CFO.
Garbage In, Garbage Out Was the Generous Version
The old rule was harsh but at least it was fair. Fix the input, fix the output. Schein thinks probabilistic systems quietly voided that deal, and he’s blunt about what it means for anyone counting on cleanup alone to solve the problem.
“We used to comfort ourselves with garbage in, garbage out. It was a harsh rule, but it was a fair one: fix the input and you fixed the output. AI broke the symmetry. These are probabilistic systems, so now quality in can still produce garbage out: a confident summary that misreads clean numbers, an invented detail stitched seamlessly into governed data. That is not an argument for skipping the cleanup. Clean data is what makes those failures rare, detectable, and cheap to catch instead of common, invisible, and expensive. But it changes the posture: data quality buys down the risk, it does not retire it.”
Worth sitting with, because it changes what a cleanup budget actually purchases. Not immunity. Lower incidence and faster detection, which is a different thing to promise a board. Schein’s read is that you need governance on the input and observability on the output, and you should fund both — most organizations have a line item for the first and nothing for the second. Verification labor is where that gap gets expensive. Research from BetterUp Labs and Stanford’s Social Media Lab, published in Harvard Business Review in September 2025, put a number on it: workers surveyed spent close to two hours untangling each piece of low-quality AI output that landed on them. Schein has a name for the person doing that untangling. “Human in the slop,” he calls it — a job nobody applied for that’s becoming a real fraction of knowledge work.
One Is a Repair. The Other Is a Subscription.
So the question isn’t whether to clean. It’s when, and Schein’s framing here is the sharpest thing in his answers.
“Clean upstream and you fix a definition once, in a pipeline, deterministically, and every run afterward inherits it free. Clean downstream and you pay a probabilistic model to guess at the same mess on every run, forever. One is a repair. The other is a subscription to your own sloppiness.”
There’s a second edge here that most marketing leaders haven’t priced in. Schein contends that good context lets a smaller, cheaper model do the same job — that “a mid-tier model with governed data beats a frontier model with a messy CRM, at a fraction of the price.” If that holds in your stack, the vendor conversation inverts. You stop shopping for a more capable model to compensate for bad inputs and start asking what the cheapest model is that your data can support. But the compounding return isn’t cost, it’s trust. When inputs are governed, Schein says, outputs stop needing inspection, “and that is the moment AI shifts from a demo to an operation.” Everything currently stuck in pilot is stuck at that line.
Machine Speed or Meeting Speed
Say “govern your data” to a marketing team and most of them hear “more approvals.” Schein offers a test for telling the two apart.
“One test: does the control run at machine speed or at meeting speed? Governance that makes AI cheaper is plumbing. Row-level security, masked fields, shared definitions, lineage, and logging of every AI action all execute automatically, invisibly, in milliseconds. Nobody schedules a meeting for them… Governance that is only friction is posture. It shows up as approval steps where a human clicks yes without changing anything, and policies that govern the deck instead of the pipeline.”
The second test is easy to run and more uncomfortable: check your approval rate. If reviewers wave through ninety-nine percent of what crosses their desk, Schein’s verdict is that you haven’t built oversight, you’ve built a toll booth — and you’re paying the toll in cycle time while training people to rubber-stamp the one percent that mattered. That leads him to a vocabulary worth borrowing, since human-in-the-loop has become a phrase people use to mean any human involvement at all.
“Everyone knows human in the loop. I would add two more positions: human in the lead and human in the slop. In the lead, you sit upstream of the AI. You decide whether it runs at all, on what data, with what context, and you own the outcome… In the slop, you sit underneath the output, triaging generated volume nobody chose to review. The first two positions are designed. The third is inherited.”
Applied to a marketing org, that distinction is a staffing audit. Who’s upstream, setting the terms an agent operates under? Who sits at a checkpoint that genuinely changes an outcome? And who’s just cleaning up volume that exists because a tool made generating it easy? The first thing a leader has to give up, Schein says, is the review meeting as the control point — harder than it sounds, because staring at dashboards together feels like rigor. His line for it: “A dashboard is a place where a human goes to find work. An agent is a worker that comes to a human with a decision.”
Which brings the deferred cleanup project back around, this time with a deadline on it. Before signing the next AI contract, Schein’s advice is to audit your own house, because the vendor demo runs on their clean data and your production runs on yours. Pull one real customer and count how many versions of them exist across CRM, email, ad platforms, and web analytics — that’s your identity resolution problem, and if the answer is four, every run pays to reconcile four. Then count the segments that still require a human to export, massage, and re-upload. That’s your manual-intervention rate, and any vendor promise sitting on top of it inherits it.
None of that is glamorous work, and none of it will appear in a keynote. But Schein’s underlying claim is that the model is a commodity and roughly everyone is selling the same one, which means the differentiated asset was never the AI. It was your data, your definitions, and your customer history. Get those wrong and, as he puts it, “you are not buying AI capability. You are buying a very fast amplifier for your existing mess.”


