This article was based on the interview with Eikona CEO Nir Weingarten on the limitations of traditional A/B testing and how to make it better by Greg Kihlström, AI and MarTech keynote speaker for The Agile Brand with Greg Kihlström podcast. Listen to the original episode here:
For as long as most of us have been in marketing, A/B testing has been the bedrock of data-driven optimization. It’s the tool we reach for to settle debates, validate assumptions, and incrementally improve performance. We’ve built entire teams and processes around the familiar cycle of hypothesize, test, analyze, and implement. It provides a comforting sense of scientific rigor in a field that can often feel more like art. But what if that comfort is a liability? What if the “winning” variation from your latest test is actually the wrong choice for the vast majority of your audience? In the complex, diverse world of enterprise marketing, the notion of a single champion creative is becoming increasingly outdated.
The challenge we face as leaders is not a failure of the A/B testing concept itself, but a limitation of its manual, resource-intensive nature in an era that demands scale and personalization. The sheer effort required to run a statistically significant multivariate test often means we test too little, too slowly, and with inherent biases. This is where the conversation needs to shift—from a mindset of finding a single winner to one of continuous learning and adaptation. I recently spoke with Nir Weingarten, CEO and founder of Eikona, whose background in AI and reinforcement learning offers a compelling new paradigm. His perspective suggests that the natural evolution of A/B testing is already here, and it promises to move marketing from a state of periodic testing to one of perpetual optimization.
The Foundational Flaws of Our Favorite Tool
Before we can embrace a new approach, we must be clear-eyed about the limitations of the current one. While no one is suggesting we abandon the principles of testing, it’s crucial to acknowledge where the traditional A/B testing model breaks down, particularly at the enterprise level. The process is manual, slow, and fundamentally constrained. It forces marketers to make narrow choices about what to test, often based on intuition, which introduces bias from the very beginning. As Nir points out, the issues are systemic and threefold, revolving around scale, depth, and the very human biases we bring to the process.
“I think there’s three major problems with A/B testing, and first and foremost is scale. It’s a manual process that doesn’t scale… The second part is multi-variant testing. So, you know, you’d wanna ask yourself, ‘What am I going to change?’… You always change one thing, because otherwise you won’t know what caused the uplift… And I would say that the third thing is bias… you’re limited in your resources, so you have to test certain things. You can’t test everything. You’re gonna test stuff that you think that are going to work, and when you’re gonna look at the results, you’re gonna see the stuff that you would wanna see.” – Nir Weingarten
Nir’s breakdown resonates with the practical reality many of us experience. The sheer manpower required—from creative production to campaign setup to data analysis—creates a bottleneck that severely limits the volume and complexity of tests. We end up testing a new subject line or a button color, not because it’s the most impactful variable, but because it’s the easiest to isolate. This one-variable-at-a-time approach means we miss the powerful interplay between elements—the combination of an image, a headline, and a CTA that might unlock a new level of performance. And the bias he mentions is subtle but pervasive. We test our own pet theories and then, thanks to confirmation bias, find the evidence we were looking for. The result is a process that feels scientific but is often just a resource-intensive way of confirming what we already believed.
Reinforcement Learning: A More Intuitive Approach
If traditional testing is a series of isolated experiments, the alternative is a system that learns continuously from every single interaction. This is the core idea behind reinforcement learning, a branch of AI that powers everything from recommendation engines to the large language models we’re all becoming familiar with. Instead of a rigid “test and declare a winner” framework, reinforcement learning employs an “explore and exploit” model. It’s constantly trying new variations (exploring) while simultaneously serving more of what’s already proven to work (exploiting). Nir offers a remarkably intuitive analogy for how this works, stripping away the complex jargon to reveal a process that mirrors how we, as humans, learn.
“Imagine a baby, and the baby is born, and they walk around in the world exploring their world, and they see a piece of candy… they eat that piece of candy, and it’s sweet, and it’s tasty, and now they’ve learned, ’cause they were rewarded, right?.. But on the other day, the baby walked around, and then they saw a lemon… and they took a bite, and it was sour… So the baby learned by trial and error that candy is nice and lemons aren’t… Basically, a reinforcement learning is a type of machine learning that is based around these concepts.” – Nir Weingarten
This analogy is powerful for marketing leaders because it reframes optimization from a discrete task to an ongoing state. In our world, the “candy” is a click, a conversion, or a purchase. The “lemon” is an unsubscribe or a bounce. A system built on reinforcement learning can test hundreds of combinations of headlines, images, copy, and offers in a single campaign. It learns in real time which combinations are “candy” for specific audience segments and automatically allocates more traffic to them, while learning from the “lemons” and reducing their exposure. This moves us away from finding the single best message for everyone and toward finding the best message for each person, or each cohort, at that specific moment. It’s a model that embraces audience diversity rather than trying to average it away.
The New Role for Marketers: The 10/80/10 Rule
Naturally, the introduction of such a powerful automation layer raises questions about the role of the marketer. If a machine is handling the heavy lifting of testing and optimization, what’s left for the team to do? This is where many conversations about AI devolve into fear or hype. The reality, however, is more nuanced and, frankly, more interesting. It’s not about replacement; it’s about augmentation and elevating the strategic function of our teams. Nir introduces a framework for this new human-machine collaboration that he calls the “10/80/10 rule.”
“There’s this saying online today that AI is sort of, like, using AI is sort of like this, only you take the 20% and you split it into two tens… AI is very capable. It has a lot of muscle, but it doesn’t know what to do… The first 10% is creating that [initial brief or creative]. Then the 80% is all the heavy lifting of creating the variations… And then you need the last 10% of the marketer to come along and say, ‘Hey, you know, that’s off brand’ or, ‘I didn’t really like that,’ applying their judgment on the system.” – Nir Weingarten
This model is a practical guide for structuring workflows in an AI-driven marketing world. The first 10% is quintessentially human: strategy, brand knowledge, campaign objectives, and the initial creative spark. AI cannot invent the core marketing strategy from a vacuum. The machine then takes over for the 80%—the grunt work of generating countless on-brand variations, testing them, and analyzing performance data at a scale no human team could ever match. Finally, the crucial last 10% brings the human back in for curation, oversight, and judgment. This is where the brand manager ensures everything is on-brand, the copy chief polishes the tone, and the strategist extracts insights for the next campaign. This framework frees marketers from the tactical weeds of test setup and allows them to focus on higher-value work: strategy at the front end and critical analysis at the back end.
Beyond the Black Box: Measuring What Matters
One of the primary concerns for any leader evaluating a new technology, particularly an AI-driven one, is the “black box” problem. How do we measure its impact, and how do we gain insights if the machine is making decisions autonomously? A system that simply says “trust me, it’s working” is a non-starter. True enterprise-grade solutions must provide both transparent measurement and actionable insights. The good news is that the principles of sound measurement don’t disappear; they are simply applied more dynamically. A constant, randomized control group remains the gold standard for proving incremental lift.
Furthermore, the goal isn’t necessarily to understand the precise, inscrutable “why” behind every algorithmic decision. Instead, it’s to surface the characteristics of what’s winning so that human marketers can learn and apply those insights across other channels.
“When something does succeed… you’re gonna get a push notification for us saying, ‘Hey, we found a winner’… we’re not gonna say why, ’cause we don’t know… What we’re gonna do is we’re gonna say, ‘Hey, these are the buy intents that we’ve inferred from this creative… This is the type of messaging that was most dominant… It’s urgency, it’s novelty… and this is how it differs from the control.’ And now, you know, you’re the marketer, you take the judgment, you decide what you wanna do with it.” – Nir Weingarten
This approach strikes the right balance between automation and human intelligence. The system doesn’t offer false certainty by trying to explain the inner workings of a neural network. Instead, it provides tangible, actionable creative intelligence. By identifying winning themes (e.g., messaging focused on “novelty” outperforms messaging focused on “urgency” for a specific product launch), the platform equips marketers with insights that can inform strategy far beyond a single email campaign. A winning concept from an owned media channel can be ported over to paid social creative, landing page copy, or even future product positioning. The system becomes not just an optimization engine but a powerful, perpetual insight-generation machine.
Summary
The shift from traditional A/B testing to AI-powered reinforcement learning is not merely a technological upgrade; it represents a fundamental change in our philosophy of marketing optimization. We are moving away from a world where we conduct periodic, slow, and narrow experiments to find a single, temporary winner. Instead, we are entering an era of continuous, broad, and automated learning, where the goal is to adapt to every customer’s preference in real-time. This approach acknowledges the reality that our audiences are not monoliths and that a one-size-fits-all “champion” is an illusion.
For us as leaders, the mandate is clear. Our role is not to become prompt engineers or data scientists overnight. It is to understand these new capabilities and redesign our teams and processes to leverage them effectively. The 10/80/10 model provides a compelling blueprint, empowering our teams to focus on strategy and curation while handing the immense tactical workload to machines that are built for it. The future of marketing isn’t one where humans are obsolete; it’s one where they are freed from manual constraints to become true orchestrators of intelligent, adaptive customer experiences. The era of the single winner is over. Welcome to the age of perpetual, personalized adaptation.



