AI

9 min read

How to Measure the ROI of AI in B2B Marketing (Beyond Pilots)

Instrument the work the AI does, not the tool it runs on. The 'useful work per dollar' framework, applied to enterprise B2B marketing.

How to Measure the ROI of AI in B2B Marketing (Beyond Pilots)

The AI ROI question tends to arrive after the pilot ends and the board asks what it produced. Somebody pulls the seat count. Somebody pulls a bill for tokens. Somebody points at a slide with a productivity claim on it. Nobody in the room believes the number.

That is the pattern behind the numbers boards have been reading. MIT NANDA put the enterprise AI pilot failure rate at 95%: "about 5% of AI pilot programs achieve rapid revenue acceleration; the vast majority stall, delivering little to no measurable impact on P&L" (MIT NANDA via Fortune, 2025). S&P Global's 2025 Voice of the Enterprise put the abandonment rate at 42%, up from 17% a year earlier (S&P Global / 451 Research, 2025). Two large samples, same shape: the money went in, the answer never came out.

OpenAI put a name on the fix in July: measure "useful work per dollar," tasks completed, time saved, decisions improved, and workflows ready to scale (OpenAI, 2026). It is the first CFO-legible ROI framing to come from a frontier lab, and it lands exactly as boards force the AI ROI conversation into the finance review.

The framing is right. The trap is that most teams cannot answer it, because they have instrumented the tool instead of the work.

How do you measure ROI of AI in B2B marketing?

Measure the work the AI does, not the tool it runs on. The right unit is "useful work per dollar": qualified opportunities created, hours reclaimed on a specific workflow, decisions improved against a baseline (OpenAI, 2026). Report the delta on the actual output, pipeline, cost per opportunity, cycle time, not the seat count or the token bill.

The default move is to measure inputs: seats, tokens, subscription lines. Inputs are easy to count. They are also what the board already saw on the invoice. Reporting them back as ROI is why the number does not land.

Useful work per dollar is a different question. It asks what shipped that would not have shipped, or what shipped faster or cleaner, or what got decided with a better read than the team would have had otherwise. Then it puts a number on that against the fully-loaded cost of getting it: model calls, tool subscriptions, and the human hours that stayed in the loop. That is the unit finance will engage with.

The reason "useful work per dollar" is not a slogan is that OpenAI wrote it from the inside of enterprise deployments. Their framework pairs it with a five-step portfolio approach: visibility, model selection by full-cost-to-result, maturity-staged funding, portfolio approach, capacity planning (OpenAI, 2026). That is a CFO reading list, not a marketing one, which is exactly why it works.

The pilot-to-ROI failure has a shape once you look for it.

Why do most AI pilots fail to show ROI?

The 95% figure and the 42% abandonment figure are not tool problems. They are measurement problems. Teams instrument the tool, logins, prompts, tokens, instead of the workflow the tool was supposed to change. The pilot ends, the data does not describe an outcome, and the board reads a productivity slide instead of a P&L line (MIT NANDA, 2025).

We have sat in enough post-mortems to know the shape. The tool got adopted. Sessions grew. Somebody ran a survey and 70% of users said it helped. The board asked what changed on pipeline, cycle time, or cost per opportunity, and the room went quiet.

The mismatch has a structural source. Enterprise AI is being sold at the tool tier and consumed at the workflow tier. The buyer signs for licenses; the value has to show up in a report someone else builds. If the reporting layer was not designed to catch a workflow delta, it will not catch a workflow delta. The pilot did not fail. The measurement never ran.

The abandonment rate is doing the same thing from the other side. When a team cannot show what changed, they stop defending the spend, and the pilot gets rolled up (S&P Global / 451 Research, 2025). That looks like the tool underperformed. Usually the tool did fine, and no one built the scoreboard.

Fixing the scoreboard is a specific job with a specific shape.

What is "useful work per dollar" and how do you actually calculate it?

Pick one workflow: a proposal draft, a QBR report, an ad concept round, a lead qualification pass. Measure two things before AI: unit cost (fully loaded) and quality signal (win rate, response rate, conversion). Run the AI-assisted version over 30 to 60 days. Measure the same two things. Report the delta as $/unit-of-useful-work.

The temptation is to boil the ocean, instrument every workflow, dashboard every model call, wire everything up before reporting anything. That is the multi-quarter version of never reporting anything. Pick one workflow and prove it out.

The unit-of-useful-work is what finance can price. For a proposal drafter, it is a shipped proposal at the same or better win rate. For a QBR report, it is a delivered report at the same or better retention signal. For a lead-qualification pass, it is a qualified opportunity at a lower or equal cost per opportunity. Each unit has a fully-loaded cost (model calls, subscriptions, human review) and each has a quality signal that can be measured against the pre-AI baseline.

Here is the difference laid out.

Dimension

Tool-first ROI (why it fails)

Work-first ROI (what to bring to the board)

The unit

Seats, tokens, subscription lines

Qualified opportunities, shipped units, hours reclaimed on a workflow

The comparison

Cost of AI vs cost of last year's software

Cost per unit of useful work with AI vs without

The quality signal

Adoption, session count, satisfaction survey

Win rate, response rate, cycle time, cost per opportunity

Whether it survives scrutiny

Falls apart at "so what changed in pipeline?"

Holds, because it names a workflow and prices its output

Once the unit is right, the next question is whether the tool doing the work is actually doing the work.

What is an AI agent in B2B marketing, and how is it different from a chatbot with a nice wrapper?

An AI agent is a system that plans and executes a multi-step workflow with tools and memory, then delivers an outcome a human would otherwise have delivered. A wrapper is a single-prompt chatbot dressed up. VentureBeat's June 2026 enterprise wave found 71% of organizations say a quarter or fewer of their deployed "agents" are true orchestrated workflows (VentureBeat, 2026).

The word "agent" is being applied to two different things, and buyers are absorbing the cost of the confusion. A wrapper answers one prompt and stops. An agent runs a workflow, calls tools, remembers state, and hands back an artifact: a draft, a decision, a set of trafficked ads. The ROI math is different for each, because the work is different.

The diagnostic is simple. Ask what the AI produces at the end of a run: a message, or a deliverable. A message is a wrapper output; a deliverable is agent output. Then ask how many discrete steps it took, and whether any state was maintained between them. If the answer is one step and no memory, the ROI ceiling is the labor cost of the single task it displaced. If the answer is a multi-step workflow with memory, the ROI ceiling is the labor cost of the entire workflow, usually an order of magnitude higher.

This distinction matters for buyers evaluating partners too. Publicis Groupe now reports 87% of net revenue from "AI-powered marketing services," while Publicis Sapient, the group's transformation arm at 13% of net revenue, posted a mid-single-digit decline (Adweek, 2026). "AI-powered" is a holdco reporting construct, not something a client can inspect. If your partner cannot show the workflow, the memory, and the artifact, you are paying for a wrapper at agent prices.

The same discipline decides what AI actually changes about how B2B media gets planned and bought.

How is AI changing B2B media planning and buying?

AI reshapes the workflows with clear inputs, structured outputs, and a scoreboard first: audience segmentation, creative iteration, bid optimization, and campaign QA. Personalization at scale is narrower than the pitch. It works where the CRM signal is clean and fails everywhere else. The through-line: AI compresses cycle time on operational work and exposes measurement gaps upstream.

We have watched this play out across accounts. Segmentation gets faster and more granular, which means the old media plan can be rebuilt against a tighter audience in hours instead of weeks. Creative iteration goes from three variants a quarter to twenty a month, which changes what "creative testing" even means. Bid strategies that used to require analyst attention now run against a policy the platform executes overnight.

Personalization at scale is the one to be honest about. It works when the underlying signal (firmographic, intent, account context) is clean and up to date. It falls apart when the CRM is a graveyard of half-filled records, which is most CRMs. AI personalization does not fix a broken data layer; it just publishes at higher volume on top of it.

The compression cuts both ways. Faster cycles surface measurement holes the team could paper over when the cadence was slower. A creative that used to run for a quarter now runs for a week, so the fatigue signal has to be read weekly, not monthly. A segmentation that used to change annually now changes per campaign, so the attribution model has to hold up under that churn. Teams that stood up AI on top of a shaky measurement architecture are the ones showing up in the 95% and 42% figures. The tool did the work. The measurement did not survive the speed.

Which is why the ROI question always turns into a measurement-forensics question one layer down.

What should you audit before you report AI ROI?

Audit the source data, the stage integrity, and the opportunity ownership before you report a number. AI systems optimize toward whatever they are told to optimize toward, so if the underlying attribution is wrong, the AI will produce the wrong result more efficiently and the report will look clean. Verify the scoreboard is measuring the outcome you meant to measure.

The pattern we keep seeing: an AI-assisted campaign reports strong new-customer acquisition, and the underlying platform has been quietly harvesting brand demand. Across three anonymized accounts in different industries, we found the same shape. In one, 69% of "new customer" purchases flagged by an AI-driven prospecting campaign turned out to be brand searches. The campaign was set up to prospect and was harvesting existing demand. The report looked like it was working. The measurement was wrong at the source.

That is the honest version of the AI ROI question. Every AI pilot ROI question is a measurement-forensics question one layer down. The AI compounds whatever it is given. Clean data compounds into a real result. Dirty data compounds into a confident wrong answer at the same speed. The board reads either way as "the AI produced this," and no one goes looking until the number stops matching the pipeline.

Verifying the scoreboard is a specific piece of work. Source data: is the input feed capturing what you think it is capturing? Stage integrity: do "MQL," "SQL," "opportunity," and "closed-won" have the same definition across marketing, sales, and finance? Ownership: is there one team accountable for each metric, or does the number live in three tools with three different owners? Getting those three right is the measurement architecture underneath the ROI number. Standing that up is the unglamorous core of Moving Parade's Foundation and performance-modeling work.

One move: Before your next AI pilot post-mortem, pick one workflow the tool was supposed to change. Write the unit of useful work in one sentence, a qualified opportunity at $X, a shipped proposal at Y% win rate, an hour reclaimed on Z. If you cannot finish the sentence, the pilot did not fail. The scoreboard did.

Frequently asked questions

Is it possible to get real ROI from AI marketing pilots? Yes, and about 5% of enterprise pilots do, the ones that measure a workflow delta, not a tool adoption curve (MIT NANDA, 2025). The teams in that 5% picked one workflow, priced its output, and reported the delta against a fully-loaded baseline. Every team we have seen exit the failure pattern did that first.

How is AI marketing ROI different from other AI ROI? Marketing has a live scoreboard, pipeline, cost per opportunity, cycle time, that most other functions do not. That should make AI ROI easier to prove and harder to fake. In practice, it does both: the layer that scores human work is the same layer that scores AI work, so if attribution is broken, AI ROI is broken by inheritance.

What is the difference between "AI-powered" services and truly agentic work? "AI-powered" is a reporting construct on the seller's income statement. It says the seller uses AI somewhere in their stack. Agentic work is a workflow the buyer can inspect: multi-step, tool-using, artifact-producing, with the memory and audit trail to prove it (VentureBeat, 2026). Ask to see the workflow, the state, and the artifact.

Should we buy AI tools or build with AI agents? Both, and the ROI math is different. Tools give you a shared UI on top of a model, priced per seat: the ceiling is the labor cost of the single task the tool displaces. Agents run a workflow, priced per unit of work: the ceiling is the labor cost of the whole workflow. Decide by the output, not the box.

Ready to build pipeline?

Tell us where you are.
We'll tell you what we can do.

Ready to build pipeline?

Tell us where you are.
We'll tell you what we can do.

Ready to build pipeline?

Tell us where you are.
We'll tell you what we can do.

Ready to build pipeline?

Tell us where you are.
We'll tell you what we can do.