← Learning Hub

Featured

The Agentic Era Rewards Truth, Not Speed

95% of AI efforts failed in 2025: not from weak models, but weak fundamentals. Why product marketing built on customer truth is the moat in the agentic era

Jul 12, 2026, 11:40:37 PM 15 min read

Why roughly 95% of companies saw no measurable return on generative AI in 2025, and the two questions every CEO racing to go agentic should bring to their next board meeting. One of them is "Do we actually have the fundamentals?"


 

Grow to ten million euros in annual recurring revenue (ARR). Ten people. Maximum. That's what our business plan says.

Every time I say that out loud, a little voice in my head does the math and looks at me like I've forgotten how companies work. And that little voice is right. Ten humans do not run a €10M tech business. The headcount doesn't add up, and it's not supposed to. The rest of the company is an agentic workforce. Agents doing the work, making decisions, moving things forward, at a scale and speed ten people never could.

So my co-founders and I spent the last year staking our company on a single bet: that agents can carry the load.

And somewhere in that year we realised the bet isn't really about the agents at all.


The change nobody gets to opt out of

Let's start with what's actually happening, because it's bigger than my company or yours.

Every board on earth has "become agentic" somewhere on the 2026 agenda. This isn't a fad you can wait out. Gartner expects that by 2028, at least 15% of day-to-day work decisions will be made autonomously by AI agents, up from essentially zero in 2024, and that a third of enterprise software will have agentic capabilities baked in. The shift is happening independently of what any one of us decides. Agents are going to be making decisions inside your business. Many of them. Faster than your people, and in far greater numbers.

That's the change. And it comes with a stake attached, the way real changes always do: agents are already drafting the work your people rubber-stamp, and the autonomous share is rising. The only open question is what those decisions will be grounded on. Opting out doesn't remove the stake, either; it only changes how quickly the bill arrives.

The reason 95% of it failed

Start from something every strong company already knows: growth begins with a deep understanding of your customers. The best operators make almost every decision a customer-centric one, most of all in product and go-to-market, which happens to be exactly where most companies spend most of their money. So when we talk about making good decisions, whether it's decisions made by humans or agents, these need to be grounded in deep customer insights and truth.

Now the uncomfortable part.

In 2025, MIT's Project NANDA published The GenAI Divide: State of AI in Business: 52 structured interviews, 153 survey responses from senior leaders, and a review of 300+ real AI initiatives. The headline number went everywhere: roughly 95% of organizations got zero measurable return on their generative-AI investment. Billions spent, a rounding error back. (The authors themselves call it a directionally-accurate six-month snapshot, not a final verdict, and it's drawn its share of pushback, so take the exact figure with a grain of salt. The direction is harder to argue with.)

 

But don't stop at the number, because the instinct it triggers, the technology just isn't ready yet, is the wrong lesson. Read the why and a different story shows up. MIT found the failures weren't about model quality at all. They were about what the authors call the learning gap: tools that couldn't retain context, couldn't plug into real workflows, couldn't improve over time. Buyers who partnered with vendors and integrated deeply succeeded far more often than teams building disconnected tools of their own: internal builds succeeded only about half as often (roughly 33% vs 67%).

 

And don't write this off as a 2025 story that newer models have since solved. S&P Global's 2025 survey found firms abandoning most of their AI projects at more than double the prior year's rate, and through 2026, a year of fast model progress, the numbers barely moved: 2026 surveys from PwC and Foundry landed in the same place, with most enterprises still unable to show a return. Capability had stopped being the bottleneck: the frontier models converged, and the gap moved into the workflow, not the model. That's the tell. If a year of dramatically better models doesn't move the number, the number was never about the models.

 

Gartner tells a compatible story from the agentic side: it predicts over 40% of agentic AI projects will be cancelled by the end of 2027, naming escalating cost, unclear business value, and inadequate risk controls. Different words from MIT's learning gap, but they rhyme: projects die on integration, ownership, and value, not on how smart the model is.

 

Here's the through-line I read across both, and I'll own it as my read, not their finding. When a system isn't grounded (no reliable source to reach, no one owning the truth it acts on) it fills the gap with a guess, delivered with full confidence. A tool that cannot retain context has to re-guess that context on every call. MIT names the mechanism a learning gap; Gartner counts the wreckage in cancelled projects. I call the thing underneath both the confident guess. That, not AI hype and not slow models, is the failure mode.

 

Left to its own devices, an agent fills every gap it hits with whatever its training data suggests, a plausible-sounding assumption, or a deep search across whatever the open web happens to serve up that day. It doesn't hesitate. It doesn't flag the seam between what it knows and what it invented. It hands you fluent, self-assured, beautifully-formatted output, and some unknowable fraction of it is made up.

 

I want to be honest here: we nearly walked into this ourselves. Early on it's intoxicating to watch an agent produce a polished answer in seconds. You have to keep reminding yourself that polished and true are not the same word. Everyone chasing agentic is up against the same failure mode, and most don't know it yet.

 

Let me make that concrete, because I've watched it cost real money. A team we work with built an entire campaign around a message they were sure would land: a positioning call made from gut and a few confident voices in a room, not from anything their customers had actually said. The creative made by agents looked sharp at first sight. The launch was clean. And it produced almost zero pipeline, because the message answered a problem the buyers didn't really have. The copy turned out to be generic. The buyers saw right through it. That one ungrounded decision cost them well into six figures: a confident guess, delivered beautifully. And note that this was a human failure, made at human speed; the agents only executed the creative.

 

Why is that guess so much more dangerous now than it was a year ago? That was one team, one campaign, one bill. When an agent makes the same kind of call from the same kind of assumption, that flawed logic runs across thousands of decisions, automatically, downstream, at machine speed. The agentic era doesn't just scale your good execution. It scales your bad execution just as fast, and hands you the bill quarters later, when someone in a budget review asks what the project actually returned and the room goes quiet.

 

That silence is what a cancellation sounds like. It's what 95% sounds like.

What a grounded agent looks like

So picture the other side of that divide.

 

Picture an agent that, when it hits a gap, doesn't guess. It reaches into a repository of verified customer truth: real findings, anchored to real quotes from real customers, each carrying a confidence label that says how much weight it can bear. It answers. And every claim traces back to something a customer actually said.

 

A product marketer hands an agent the next launch. Instead of inventing benefit statements no buyer has ever uttered, it drafts positioning built from the exact phrases churned customers used on their way out, each line traceable to the interview it came from, each carrying a confidence label that says how many customers actually said it. Or a sales development (SDR) agent personalizing outreach: not a plausible-sounding pain scraped from a job title, but the real objection three lost deals raised last quarter, quoted. Win/loss, churn interviews, ideal customer profile (ICP) refinement, messaging, battlecards (the competitive one-pagers sales teams carry), willingness-to-pay: these are the surfaces where go-to-market teams already spend, and every one of them is a place an agent either grounds in what customers said or makes something up.

 

That's not a fantasy feature. It's how the technology was designed to work. Retrieval-augmented generation (RAG), grounding a model in an external, trusted source instead of its own training data, was shown years ago to produce measurably more factual, grounded output, with sources you can inspect and cite. The mechanism has existed the whole time.

 

If you've built with RAG, you're already objecting: we wired up retrieval and it still hallucinated. You're right, and it's the most important thing to be clear about. RAG the technique is commoditized, a retrieval step anyone can bolt on. Point it at an unvalidated dump of documents and it will still lie confidently, because it's grounding in noise. The differentiator was never the retrieval. It's the corpus: whether what the agent reaches is curated, quote-anchored, confidence-labeled, and traceable to a real source, or a pile of PDFs nobody validated. Naive RAG over a messy corpus is a confident guess with a citation stapled to it. The moat is the quality and structure of what you ground in, not the act of grounding.

 

On that side of the divide, your fundamentals stop being academic hygiene and become the thing that lets the business move fast without making things up. Confidence labels and quote-level traceability turn into a governance layer: everyone can see how much an insight can bear and trace it back to the customer's own words. That's what makes it safe to let an agent act. For those of us building in Europe, that traceability isn't a nice-to-have either. When an agent makes a claim about a person, or acts on customer-interview data, provenance is what lets you show where a decision came from and honor an access or erasure request: the difference, under GDPR, between an agent you can account for and one you can't.

 

And this is the flip, the counterintuitive heart of the whole thing:

 

Model limitations make your fundamentals more valuable, not less.

 

The more decisions you hand to agents, the more leverage a single well-grounded, validated insight carries, because it now shapes hundreds of automated decisions, not one person's Tuesday. The researcher who used to "produce reports people skimmed" becomes the owner of the customer-truth layer the whole business runs on, humans and agents alike. Rigor was never the boring part. In the agentic era, rigor is the moat.

The agent matters less than what you feed it

This is the part everyone gets backwards. They spend the budget on the agent: the model, the orchestration, the demo that dazzles the board. They spend almost nothing on what the agent stands on.

 

The general fix is grounding: give the agent a trusted external source to reach instead of its own training data, and make someone own it. That principle is well established, and it isn't mine. The bet I'm making is narrower: that for the decisions product and go-to-market teams hand to agents, the source worth grounding in is a customer-truth layer: insights stored atomically, each one a finding plus its evidence quote, its tags, its confidence label; connected, queryable, and wired into the tools where decisions actually happen. Grounding is the principle. Customer truth is where I'm putting my chips.

 

Connect that layer to the applications running your agents, and every insight they reach is quote-anchored and verified against the source interview, so they reason from evidence instead of filling gaps with guesses. The agent stops being a liability and becomes a conduit for verified customer truth. MIT's buy-versus-build finding points the same way: teams that integrated a partner's system succeeded about twice as often as teams that built their own. That's the system we built at Bubble, and it's the reason I can put that absurd number in our business plan.

 

There's a sharper edge here for anyone who builds software. Gartner now puts up to $234 billion of enterprise-application software spend at risk from what it calls agentic arbitrage by 2030, as agents complete the work across systems and the interface stops being the differentiator; the market repriced that in real time when a February 2026 selloff erased roughly $285 billion in SaaS value in about 48 hours. Gartner's survivors are the vendors who capture and keep their customers' knowledge, and the deepest version of that knowledge (this extension is my read, not Gartner's) is what your customers actually said.

The strongest version of the claim I'm entitled to make

Here's the most honest thing I can offer, and I want to be precise about what it is and isn't: we run on this ourselves. Not proof: the €10M number is a target, not a result, and I'd be doing the exact thing this article warns against if I dressed a projection up as evidence. What it is, is conviction with skin in the game.

 

The €10M-with-10-people bet only works if our agents aren't guessing, so we grounded them in the same customer-truth layer we sell. When an agent makes a call about a segment, a pain, a churn risk, it's reaching a real quote with a confidence label, not improvising from training data. If I'm wrong about grounding, I don't get to watch it fail in a slide deck. I watch it fail live, in my own company. That's the strongest version of the claim I'm entitled to make: not this works, but I've bet the company that it does.

 

That's also why I think this is the moment, and not a year from now. Three things had to line up: a genuine board-level shift (agentic is real, not hype), a market that's been underserved (customer research built for everyone except the product marketers and product managers who need it most), and a defensible, EU-native system built specifically for them. Remove any one of those and it doesn't hold: a clever tool with no urgency, or urgency with nothing underneath. Right now all three hold at once.

The two questions to bring back to your board

If you take one thing from this, don't take a product. Take a diagnostic. Walk into your next leadership meeting and put two questions on the table:

 

1. Do we actually have the fundamentals: decisions rooted in verified customer truth, with the evidence and confidence to back them? Not slides. Not a wiki nobody reads. A living, validated, traceable layer of what your customers actually said.

 

2. Do our agents have access to them? Because a truth layer your agents can't reach is a truth layer that doesn't exist as far as the decisions are concerned.

 

Sit with what each "no" costs you.

 

If the answer to the first question is no, the agentic shift will take your weak fundamentals and scale them: bad decisions and poor execution, automated, downstream, faster than any review cycle, straight into your growth and your margins. And sitting the transition out doesn't avoid that bill; it just arrives more slowly, in your valuation and your next raise.

 

There's no version of this where the fundamentals don't matter. The agentic era just raised the stakes on getting them right, and shortened the time you have to do it.

 

Agents are already drafting the decisions your people sign off on, and their share is only growing. The question left is whether you've given them the truth to decide on.




 

Sources referenced: MIT NANDA, "The GenAI Divide: State of AI in Business 2025"; Gartner press release, "Over 40% of Agentic AI Projects Will Be Canceled by End of 2027" (June 2025); Gartner press release, "$234 Billion in Enterprise Application Software Spend Is at Risk from Agentic AI" (July 2026); Lewis et al., "Retrieval-Augmented Generation" (2020). Return-on-AI figures draw on S&P Global Market Intelligence's 2025 Voice of the Enterprise survey, PwC's 2026 Global CEO Survey, and Foundry's 2026 State of the CIO. Market corroboration of the February 2026 SaaS selloff draws on contemporary reporting including CIO Dive.

 

Get your free PMM Research Agent

Your free PMM Research Agent will help you get there. Fast & accurately

Get Amy Now