All notes
Marketing Engineering 2 August 2026 15 min read

When marketing runs on agents, the job becomes management

81% of martech leaders are piloting AI agents. No business function has scaled them past 10%. The gap is supervision: agents fail differently every run, and reviewing them is a management discipline most marketers have never been trained in. What the technical floor actually is, how many agents one person can hold, and who trains the next generation of supervisors.

Gartner surveyed 413 martech leaders in mid-2025 and found 81% piloting or implementing AI agents. McKinsey surveyed 1,993 organisations across 105 countries over the same summer and found that no business function anywhere had scaled agents past 10%, marketing and sales included.

Between those two numbers sits every marketing team finding out that buying an agent and running one are separate projects. Of the martech leaders who did get agents into pilots or production, 45% say the vendor’s capability fell short of what was promised. Gartner’s forecast for the wider category: over 40% of agentic AI projects cancelled by the end of 2027, on costs, unclear value, and thin risk controls. The same release estimates only around 130 of the thousands of vendors selling agentic AI are selling something agentic.

Vendor overclaiming explains part of the gap. The rest is a job nobody has hired for. An unsupervised agent produces work faster than any marketing team can read it, in a quality band that shifts from run to run, and the person holding that problem is a marketer rather than a data scientist or an ops admin. You end up applying marketing judgment at review time instead of production time, which is the shape of a management job.

Why do agent pilots stall in marketing?

The hours come back somewhere else.

Workday surveyed 3,200 employees at $100M+ companies, all of them active AI users, and found nearly 40% of AI time savings lost to rework: correcting errors, rewriting output, verifying results. 77% of them review AI-generated work as carefully as human work or more carefully. Glean’s Work AI Index, drawn from 6,000 digital workers, put a name on the residue: 6.4 hours a week of “botsitting”, giving agents context, checking their work, cleaning up after them, against roughly 11 hours saved. 36% of AI sessions in that sample failed outright, needing a full restart or substantial rework. Glean sells enterprise AI search, so read the framing with that in mind. The direction matches everything else.

Then there is the work that looks finished. BetterUp Labs and Stanford’s Social Media Lab surveyed 1,150 full-time US employees and, writing up the results in HBR, named the output workslop: content that “masquerades as good work, but lacks the substance to meaningfully advance a given task.” 40% had received some in the previous month. Each instance cost the recipient an average of one hour fifty-six minutes, which the researchers priced at $186 per employee per month.

For a marketing team the workslop problem is worse than for most functions, because marketing output is judged on taste rather than on whether it compiles. A campaign brief that is subtly wrong about the buyer looks exactly like a campaign brief that is right. You find out at the pipeline meeting eight weeks later.

None of this argues against running agents. It argues that the supervision cost is real, unbudgeted, and sitting on the same people who were supposed to be freed up. If you have not decided which work should stay manual, the agent will make that decision for you.

What makes an agent harder to manage than a workflow?

Marketers already run automation. A HubSpot workflow, an n8n branch, a Clay enrichment table. Those systems are deterministic: same input, same output, and when they break they break loudly and identically until somebody fixes them. You QA one once and trust it until the schema changes. Error handling in n8n is a solved discipline for that reason.

Agents behave differently, and Anthropic’s engineering team is blunt about the cost. From their guidance on building agents: “the autonomous nature of agents means higher costs, and the potential for compounding errors.” From their guidance on evaluating them: “Regardless of agent type, agent behavior varies between runs, which makes evaluation results harder to interpret than they first appear.” The unit you inspect stops being the output and becomes the transcript, meaning the full record of what the agent called, in what order, and why.

Compounding is the part that catches marketing teams. A deterministic workflow with a bad enrichment field routes one lead wrong. An agent with a bad enrichment field reasons from it, writes a segment definition around it, drafts messaging for that segment, and hands you a campaign that is internally consistent and pointed at the wrong company. Every downstream step inherits the first error and adds confidence to it.

Salesforce’s own research lab put numbers on this for CRM work. Their CRMArena-Pro benchmark measured agents on enterprise CRM tasks at roughly 58% success on single-turn requests, falling to around 35% once the task ran multi-turn. That is the shape of the problem: a worker who is competent on short tasks and unreliable on long ones, and who never tells you which kind of day it is having.

The management response is evals. Build a set of cases you know the right answer to, run the agent against them, grade the results, and re-run the set every time you change the prompt, the model, or the data source. Marketers who have run a testing programme already have the instinct. The difference is that an A/B test measures the market’s response and an eval measures whether your worker is still competent.

You cannot fully automate the grading either. A June 2026 paper from Berkeley’s School of Information tested 21 LLM judge models and found exact-match scoring overstates chance-corrected agreement by 33 to 41 percentage points on MT-Bench, with judge rankings shifting by as many as 14 positions across benchmarks. Reproducible does not mean valid. Anthropic’s own advice is that LLM-based rubrics “should be frequently calibrated against expert human judgment,” and in a marketing context you are the expert.

How many agents can one marketer actually run?

Gartner’s projection is the number that reframes the org chart: by 2028 the average Fortune 500 enterprise will run over 150,000 agents, up from fewer than 15 in 2025. In the same release, 13% of organisations believe they have adequate agent governance. Max Goss, the analyst on it, describes what most of the rest have as “an ungoverned sprawl of agents.”

McKinsey’s estimate for the supervision ratio: “a human team of two to five people can already supervise an agent factory of 50 to 100 specialized agents running an end-to-end process.” The same paper names the role that results, an “M-shaped supervisor,” broad enough to orchestrate agents across domains rather than deep in one.

Practitioners running agents today report lower numbers. Tomasz Tunguz manages four concurrently and puts engineers he considers very productive at 10 to 15, with roughly half the output discarded. His observation about why the ratio stays low: agents “interpret instructions. They improvise. They occasionally ignore directions entirely.”

There is also a ceiling on the reviewer. A modelling paper from June 2026 argues that oversight capacity is finite and degrades with volume: “By the three-hundredth approval of a routine, benign action, a human is fatigued and primed to keep clicking Approve.” Ask any marketer who has approved 200 personalised email drafts in a sitting how carefully they read number 180. Piling more approval gates onto an agent can lower the quality of every gate.

Hold those two numbers next to each other and you have the job. McKinsey’s ratio works out to somewhere between 10 and 50 agents per person depending where you land in the range. What practitioners actually run, while reading every output, is closer to four. You close that distance by deciding which outputs get read.

What does an autonomy ladder look like in GTM?

Two frameworks published a year apart converge on the same shape. Gartner grades agents by what they may do: Observe, Advise, Act with Approval, Act Autonomously. Researchers at the Knight First Amendment Institute grade the same spectrum by what the human becomes: Operator, Collaborator, Consultant, Approver, Observer. Their framing is the more useful one for a marketing team, because it tells you what your calendar looks like.

Merging the two against the GTM stack:

AutonomyThe agentYou areGTM exampleWhat you check
ObserveReads and reports, changes nothingOperatorWeekly account signal digest into SlackWhether it read the right sources
AdviseDrafts and recommends, you executeCollaboratorNext-best-action per open opportunityWhether you would have acted on the advice
Act with approvalQueues the action, waits for releaseApproverOutbound sequence built and heldA sample, graded against your rubric
Act autonomouslyExecutes inside defined guardrailsObserverInbound enrichment and routing in real timeEscalations, plus a drift report on the eval set

Gartner’s warning is that enterprises treat the ladder as a switch. Shiva Varma, the analyst on that release: “Enterprises are treating AI agent governance as binary, either locked down or fully trusted, and that is the root cause of failure.” The prediction attached to it is that 40% of enterprises will demote or decommission autonomous agents by 2027, after finding the governance gap in production.

Grade the task, not the agent. The same agent can sit at Act Autonomously for enrichment, where the cost of an error is a wasted API call, and at Advise for anything a prospect will read. Inbound routing belongs at the bottom of the ladder because errors are cheap, recoverable, and measurable. A CEO-facing 1:1 ABM message stays at Approver permanently.

What does “technical” actually mean for this job?

Job postings that ask for “familiarity with AI tools” describe a user. The seat described here needs five things, and none of them require a computer science degree.

Write a spec an agent can execute. Most agent failures start as brief failures. The agent inherits whatever context you gave it and improvises the rest, so the ICP definition, the disqualifiers, the tone rules, and the source hierarchy all have to be written down rather than held in the head of whoever runs the campaign. Context engineering is the marketing version of the skill.

Build the eval set before the agent. Twenty to fifty cases where you know the right answer, half of them edge cases you have already seen in production. Grade every version against it. Without this you are managing on vibes, and vibes are how a model swap degrades your outbound for six weeks before anyone notices.

Grade autonomy per task. The ladder above, applied task by task, with the expensive stuff kept at Approver even when the agent looks ready.

Instrument the output. Baseline, metric, delta, same as proving ROI on any GTM engineering work. An agent with no measured baseline is a cost with a story attached.

Design the escalation. OpenAI’s practical guide names the two triggers worth building first: failure thresholds, where exceeding a retry or action limit routes to a human, and high-risk actions, where anything sensitive or irreversible stops for approval until the agent has earned the trust. Their framing of guardrails as layered rather than singular is the right one: “a single one is unlikely to provide sufficient protection.”

Marketers are pricing this in ahead of their leaders. Indeed’s Hiring Lab tracked AI mentions in US marketing job postings rising from 8.4% to 14.9% across 2025, inside a hiring market barely above its pre-pandemic baseline. Lightcast, working from 1.3 billion postings, found AI skills carrying a 28% salary premium and 51% of AI-skill postings sitting outside IT.

The leadership layer is further behind than the job market. Gartner’s survey of 402 senior marketing leaders found only 15% of CEOs believe their marketing leaders are AI-savvy, while 32% of marketers think their own skills need a significant update. 65% of CMOs agree AI will fundamentally alter marketing and a fifth say their personal skills need no change at all. Gartner expects AI literacy to become a top-three reason large-enterprise CMOs get replaced by 2027.

Who trains the next generation of supervisors?

Supervising an agent requires knowing what good looks like. You learn that by producing bad work under someone who tells you why it is bad, which is what junior marketing jobs are for.

Those jobs are thinning. Stanford’s Digital Economy Lab, using ADP payroll data, found a 16% relative employment decline for 22 to 25 year olds in the most AI-exposed occupations. Brynjolfsson and his co-authors went back in February 2026 to test the obvious alternative explanation, that interest rates were doing the work, and found little support for it. In marketing specifically, the Content Marketing Institute’s 2026 survey of 644 marketers found one in three companies reducing entry-level hiring, a net score of minus 19.8, and job search times stretching from 3.1 months in 2024 to 5.2 months in 2026.

The same survey shows where the work went. 76% of marketers report doing the work of more than one job and half took on new responsibilities without a pay rise, while only 11% say their company replaced workers with AI. Almost nobody is running visible AI layoffs. A junior leaves, the role goes unbackfilled, the work redistributes across people who are already at capacity, and the seat where taste used to get built disappears without anyone deciding to remove it.

A team that reviews agent output all day, staffed by people who never did the work themselves, will approve confident nonsense at scale. Whoever runs marketing in 2031 is learning to tell good from plausible right now, in a job that is getting cut this quarter. Take that into the headcount plan before the cut looks free.

Does this reduce marketing headcount?

ICONIQ Growth surveyed 150+ B2B software GTM leaders in January 2026. At $25M to $100M in revenue, the high AI adopters run 45 GTM FTEs against 65 for everyone else, and generate roughly twice the net new ARR per GTM head: $640K against $370K. Sales and post-sales headcount keeps growing 10 to 20%. Marketing and RevOps stay flat or close to it.

Read that as a compositional shift rather than a cull. Robert Half still has 65% of marketing leaders planning to expand permanent headcount in the first half of 2026. The seats moving are junior and generalist. The seats appearing are technical and senior, and Gartner’s 2026 CMO Spend Survey shows the budget behind them: 15.3% of marketing budget going to AI, with only 30% of CMOs judging their internal processes mature enough to scale it. Ewan McIntyre, Gartner’s chief of research for marketing, on that gap: “CMOs recognize AI’s potential as a force multiplier for growth, efficiency and transformation, but most marketing organizations are not yet built to capture that value.”

Being built to capture it means somebody owns the specs, the evals, the autonomy grades, and the escalation paths. That is one seat, and it looks like the marketing engineer with a fleet to run.

Where this leaves the marketer

The seat that survives this owns both halves: the system the agents run inside, and the number the agents are supposed to move. Splitting them puts the person who understands agent failure modes three meetings away from the person who finds out the campaign underperformed. That intersection is Revenue Engineering in its marketing-led form, and agent supervision is the newest thing it has to hold.

I run that seat. On the engine side, rebuilding inbound around real-time firmographic scoring moved routing accuracy from 55% to 88% and shifted 40 hours a week of manual ops into automation. On the programme side, ABM running on that infrastructure influenced $6.7M in pipeline. The stalled deal spotter and the self-improving revenue system are the agent-shaped versions of the same argument: build the loop, then supervise it, then let it tune itself where the errors are cheap.

The hiring question that sorts candidates for this: show me an agent you put into production, and tell me how you knew when it got worse. Anyone who has run one has an answer involving a set of test cases and a bad week. Anyone who has only used one talks about the prompt.

FAQ

What does it mean for a marketer to manage AI agents?

Managing AI agents means owning their specification, evaluation, autonomy level, and escalation paths rather than their outputs one at a time. The marketer writes the context the agent works from, maintains a set of test cases to detect quality drift, decides which tasks the agent may execute without approval, and defines what stops for a human. Microsoft’s 2025 Work Trend Index calls the resulting role an “agent boss,” a human manager of one or more agents, and found 36% of leaders expecting to be managing agents themselves within five years.

How is managing an agent different from managing a marketing automation workflow?

A workflow is deterministic: the same input produces the same output, and failures repeat identically until fixed. Agents vary between runs, improvise when instructions are incomplete, and compound early errors into confident downstream work. Anthropic’s engineering guidance warns of “the potential for compounding errors” and notes that “agent behavior varies between runs, which makes evaluation results harder to interpret than they first appear.” Workflows need QA once. Agents need continuous evaluation.

How many AI agents can one person supervise?

McKinsey estimates a team of two to five people can supervise 50 to 100 specialised agents running an end-to-end process, which works out to somewhere between 10 and 50 agents per person. Practitioners running agents today report lower working numbers, around four concurrent agents for a knowledge worker and 10 to 15 for productive engineers, with roughly half the output discarded. The gap comes down to how much output a human reads. Grading autonomy by task, so cheap and recoverable actions run unsupervised, is what moves the practical ratio toward the theoretical one.

Do marketers need to be technical to work with AI agents in 2026?

The floor has moved above tool familiarity. The working skills are writing specifications an agent can execute against, building and maintaining eval sets, grading autonomy per task, instrumenting outcomes against a baseline, and designing escalation triggers. AI mentions in US marketing job postings rose from 8.4% to 14.9% during 2025 according to Indeed’s Hiring Lab, and Lightcast puts the salary premium on AI skills at 28%.

What is an eval, and why do marketers need one?

An eval is a fixed set of test cases with known correct answers, run against an agent to measure whether it still performs. Marketers need them because agent quality drifts without anyone noticing when a model, prompt, or data source changes, and marketing output is judged on taste rather than on whether it errors. Twenty to fifty cases, half of them edge cases already seen in production, re-run on every change, will catch degradation that spot-checking misses.

Are AI agents reducing marketing headcount?

The evidence points to compositional change rather than straight reduction. ICONIQ found high AI adopters at $25M to $100M revenue running 45 GTM FTEs against 65 for lower adopters, with roughly twice the net new ARR per head, while sales headcount kept growing and marketing stayed flat. The Content Marketing Institute found one in three companies cutting entry-level marketing hiring and 76% of marketers doing the work of more than one job, though only 11% report workers being replaced by AI directly. Roles are going unbackfilled rather than being eliminated.

Why do most AI agent projects in marketing fail?

Gartner predicts over 40% of agentic AI projects will be cancelled by the end of 2027 due to escalating costs, unclear business value, and inadequate risk controls, and estimates only around 130 of the thousands of vendors selling agentic AI have genuine agentic capability. On the buyer side, 45% of martech leaders with agents in pilots or production say vendor capabilities fell short of what was promised. The unbudgeted cost is supervision: Workday found nearly 40% of AI time savings lost to rework, and Glean measured 6.4 hours a week spent managing agents against roughly 11 hours saved.

Keep reading