Skip to content
AI in Marketing

AI Marketing Operations: How a Modern Agency Runs Funnels with Claude, Agents and Automation

Most teams use AI; few have scaled it. The SOPs, approval points, tool lanes, costs and data rules we use to run funnels with Claude and agents.

Ray GillespieRay GillespieCo-Founder & COO

Published 20 min read

Five workflow lanes run from inputs through a gold human approval gate; four lanes ship and one is stopped at the gate
On this page

Key takeaways

  • Most organisations use AI; few have scaled it. In 2026, 47% of organisations were still piloting and only 25% had reached scaling.[1]
  • The gap is operations, not tools. McKinsey's AI high performers are nearly three times as likely to have redesigned their workflows, and they define when outputs need human validation.[2]
  • Pick chat, a fixed workflow or an agent by how much judgment the job needs, not by hype. Give each tool one lane.
  • Every workflow we run has five parts: a source of truth, defined inputs, an owner, an approval point before anything is sent or written, and a fallback for when the run fails.
  • MCP is plumbing with permissions. Start every connector read-only, and add writes only behind an approval step.
  • Measure against a baseline. In our experience, AI saves our ops team 15 to 20 hours a week, but only because we counted the hours before we started.

AI in marketing is no longer the hard part. Almost everyone has it. Running it well is the hard part.

Running it well means treating AI like any other part of operations. Every recurring job gets a source of truth, defined inputs, an owner, a human approval point and a plan for when it breaks. Chat, workflows and agents each get the jobs that fit them. Each tool gets one lane.

This guide is how we do that at Victory across funnels, ads, sales floors and client builds.

AI is everywhere. Scaled AI isn't.

The adoption numbers are high. HubSpot's 2026 survey of 1,500+ marketers found 86.4% of marketing teams using AI in at least some areas.[3] Salesforce found 75% of marketers had adopted it.[4] McKinsey found 88% of organisations using AI regularly in at least one function.[2]

The scaling numbers are not. In SmarterX's 2026 survey of 2,100+ professionals, 47% said their organisation was still piloting AI, 28% were still in an "understanding" phase and only 25% had reached scaling.[1] People are ahead of their companies: 53% put themselves personally in the integration or transformation phase.[1]

Impact lags further. Only 39% of McKinsey's respondents attributed any EBIT impact to AI, and most of those said it was under 5% of EBIT.[2]

47% vs 25%

Share of organisations still piloting AI vs the share that has reached scaling, 2026

[1] SmarterX / Marketing AI Institute, 2026-05-182,100+ professionals across functions, fieldwork February to April 2026. Not marketing-only.

There's a second reason this matters to anyone who hires an agency. In Gartner's 2025 CMO survey, 22% of CMOs said generative AI had reduced their reliance on external agencies for creativity and strategy, and 39% planned to cut agency budgets.[5] Agencies that use AI as a toy will get cut by clients who use it as a system.

What separates the 25% from the 47%? McKinsey's answer matches ours. Its high performers were nearly three times as likely to have fundamentally redesigned individual workflows, and they more often defined when model outputs need human validation.[2] That's operations work. The rest of this guide is what it looks like in a funnel business.

Chat, workflow or agent: pick by how much judgment the job needs

Most bad AI projects start with the wrong format: an "agent" for a job with five fixed steps, or a 40-step job pasted into a chat window.

Three formats, defined

Anthropic draws the line clearly. In a workflow, the model and its tools run through predefined code paths. In an agent, the model directs its own process and decides which tools to use.[6] Chat is the third format: you drive every step and the model answers.

The trade-off is cost. Anthropic's own guidance says agentic systems trade latency and cost for better task performance, that more autonomy means higher costs and the potential for compounding errors, and that you should start with the simplest solution that works.[6]

The decision rule we use

  • Known steps, known data: a fixed workflow. A new registrant gets a tag, a confirmation, three reminders and a pipeline stage. No judgment needed. This lives in GoHighLevel or n8n, not in a model.
  • Judgment on known data: chat with context. Turning a discovery call into a phased plan. Scoring a sales call. Writing ad hooks from a client's transcripts. A person starts the job, the model does the heavy thinking, a person approves.
  • Open-ended, multi-step research: an agent with checkpoints. Auditing a competitor's entire ad library. Monitoring an ad account overnight and flagging anomalies. The agent decides how to get there. A person checks what it found before anything changes.

If you can write the steps down, it's a workflow. If you can't write the steps but you can check the answer, it's an agent job. If you can't check the answer, it isn't an AI job yet.

What agents still can't do

Agents are improving fast, and they still fail in predictable ways. In METR's 2025 study, frontier agents were close to 100% successful on software tasks that take a skilled human under four minutes, and under 10% successful on tasks that take more than about four hours.[7] Those were early-2025 models on software tasks, and METR found the task length agents can handle was doubling roughly every seven months, so the exact figures age quickly.[7] The shape doesn't change: the longer the unattended run, the more likely it goes wrong.

The market is also full of noise. Gartner forecasts that over 40% of agentic AI projects will be cancelled by the end of 2027, and estimates that only about 130 of the thousands of vendors claiming agentic products are the real thing.[8] That's a forecast, not a measurement, and it points the same way as the adoption data. McKinsey found 23% of organisations scaling an agentic system somewhere and 39% experimenting, but no more than 10% scaling agents in any single function.[2]

Our rule follows from that. We run one Manus monitoring task per active ad account, and a person reviews every flagged change before it's applied. Nothing an agent finds touches a budget, a bid or a live ad without a human decision. We go deeper in our evaluation of Manus, Hermes and Claude agents.

The lane map: one job per tool

The fastest way to waste money on AI is six subscriptions and no rule for which one does what. Context scatters, and nobody knows where the latest plan lives.

We give each tool one lane.

Our AI lane map
LaneToolWhat it ownsWhat it doesn't own
Context-heavy thinkingClaude (Projects, Claude Code)Call notes, phased plans, briefs, copy drafts, analysis, task draftsSending anything, writing to systems without approval
Long, bounded tasksManusFull ad-library audits, per-account monitoring, campaign downloadsBudget or creative changes
ResearchPerplexitySourced first-pass researchFinal facts (we open every source)
Video and tonalityGeminiReviewing video and voiceStrategy
Creative volumeHiggsfieldStatics, GIFs, short avatar clipsTestimonials, proof, founder credibility
Site first draftsGPT-6 Astra (OpenAI)First rough draft of a homepageFinal build, which is refined in v0 and shipped on Vercel or GHL
Contact-level automationGoHighLevelNurture, reminders, pipeline stagesCross-system logic
Cross-system bridgesn8nMoving data between toolsContact-level nurture
On-page behaviourMicrosoft ClarityHeatmaps, recordings, friction signalsAttribution
Agents we've testedHermes AgentNothing in productionClient work

Lanes are how we use each tool, not a ranking of the tools. Hermes Agent appears only as tested. GPT-6 Astra is OpenAI's model; our part is the design system and prompting method around it.

Two rules keep the map honest. First, one system of record per kind of data: contacts live in the CRM, tasks live in ClickUp, call history lives in the call recorder. AI reads from those systems and drafts into them. It doesn't become a shadow database. Second, a tool earns a lane only if it makes a recurring job faster or easier while a person still owns the judgment.

That second rule is Hormozi's value equation applied to tooling. In his framing, value rises as the time delay and effort it takes to reach an outcome fall. AI mostly attacks those two terms. It rarely changes the outcome itself. A tool that saves an hour but adds forty minutes of review and cleanup has barely moved the equation, and it doesn't get a lane.

A mechanic is only as good as his tools, and the tools are only as good as the mechanic.

Ray Gillespie, Co-Founder & COO, Victory Sales Agency

The full map, with how we score each tool, is in which AI tools a marketing team should use. For the head-to-head between the two main chat assistants, see Claude vs ChatGPT for marketing teams.

Context is the product: the per-client running thread

The biggest change in how we use AI wasn't a new model. It was giving the model everything it needed to know about a client, in one place, every time.

A clever prompt on an empty context produces generic work. A plain question on a full context produces work that sounds like it came from someone who sat in every meeting.

So every client gets one maintained context:

  • Instructions: who the client is, the offer, the audience, the voice, what we're building and what we're not allowed to say.
  • Approved background: offer docs, past campaign results the client has shared with us, brand guidelines, the current funnel map.
  • A dated call log: every meeting transcript, synced in order, so the model can see what was decided, what changed and who said it.

We keep that context per client, not per topic. A "copywriting" project that mixes five clients produces copy that blends five voices, and it moves one client's data into another's work. One client, one context.

The sync is a step we run, not magic. Fathom's MCP integration works with Claude and can only reach meetings the user already has permission to view.[9] That's a feature. It means the model's access is never wider than the person running it.

The same idea scales up in Claude Code. A folder with an instructions file, a set of skills and connections to the call recorder and the task manager becomes a "brain" for one area of the business. Ask it what happened in yesterday's meeting, and it reads the transcript, drafts the action items and shows them to you before it writes anything. We walk through the setup in Claude for marketing operations.

Hormozi's Document, Demonstrate, Duplicate applies here too. He describes training people by writing down exactly how a task is done, doing it in front of them, then having them do it while you fix the checklist. A well-kept context folder is the "document" step, written once and read by both people and models.

Five SOPs we run, with inputs, approvals and failure handling

These are the workflows that do the most work in our week. Each one has the same five parts. The approval point is the part most teams skip.

1. Call → phased build plan → ClickUp tasks

A discovery or strategy call is the richest input a project gets, and most of it evaporates by the next morning. This SOP turns it into a plan the client can read and tasks the team can work.

SOP: call to phased plan to tasks
PartHow we run it
InputsCall transcript, the client's running context, the current funnel map, any files shared on the call
System of recordCall recorder for the transcript; ClickUp for tasks
AI taskDraft a phased build plan with owners, dependencies and open questions; then draft the tasks
Human approvalA strategist reviews the plan before the client sees it; task writes are previewed and approved before anything lands in ClickUp
Failure handlingIf the transcript is missing or partial, stop and flag it; never plan from memory
MeasureTime from call end to approved plan; tasks bounced back for missing detail

In our experience, we get a first-draft phased plan from one discovery call in under 30 minutes, including the human review. The closest public reference is a vendor anecdote: Anthropic says its own marketing team cut case-study drafts from 2.5 hours to 30 minutes.[10] Neither number is a controlled study. Ours is the target we hold.

ClickUp's official MCP server can execute actions as well as read data, with daily call limits that run from 100 calls on the Free plan to 2,500 on Business Plus without the AI add-on.[11] That's why we preview before writing. A model that can create tasks can also create 200 duplicates.

2. Sales-call audit and objection bank

Every recorded sales call contains two things most teams waste: coaching data for the rep and market research for marketing.

SOP: sales-call audit and objection bank
PartHow we run it
InputsRecorded calls, the scorecard categories, the offer and the script
System of recordCall recorder; a shared objection bank
AI taskScore each call by category out of 10; pull every objection, verbatim, into the bank; draft a weekly scorecard per rep
Human approvalThe sales manager reviews scores before they reach the rep; marketing reviews new objections before they become ad angles
Failure handlingLow-confidence scores are flagged for a human listen, not averaged in
MeasureShare of calls scored; time to score; close rate by rep over time

Our target is 100% of recorded sales calls scored within 24 hours, with a weekly rep scorecard. We haven't found an industry benchmark for audit coverage, so treat that as our standard, not a norm. The value is in the questions you can now ask: how is the new rep handling price, which objection showed up most this week, which close lines actually worked. The full build is in using AI to audit every sales call.

3. Funnel friction review with Clarity and Claude

Microsoft Clarity records where visitors struggle on a page. Claude turns those signals into a ranked fix list tied to your funnel's goal.

SOP: funnel friction review
PartHow we run it
InputsClarity dashboard metrics for each funnel step, notes from a handful of recordings, the funnel goal and page map, stage conversion rates
System of recordClarity for behaviour; the CRM and attribution tool for conversions
AI taskRank fixes by stage drop-off × share of sessions showing friction; list the recordings that support each item
Human approvalA person watches the recordings behind the top three items before anything ships; fixes ship as tests
Failure handlingIf the data pull fails, keep the last stored export rather than reasoning from nothing
MeasureShare of sessions on the fixed step showing friction; step conversion

Know the limits before you start. Clarity's official MCP server returns aggregate metrics, not recordings, and each project allows up to 10 requests a day, covering at most 3 days of data with up to 3 dimensions per request.[12] Clarity's own Copilot can summarise sessions and heatmaps, and Microsoft warns that generative AI can misinterpret or produce incorrect information.[13]

Contentsquare's 2026 benchmark found 35.2% of sessions affected by at least one frustration signal, across more than 99 billion sessions on 6,500 sites.[14] After a fix pass, our target is to bring the share of sessions on the fixed step showing any friction signal down to 25% to 30%. The two use different detection rules, so read that as a direction, not a like-for-like comparison. The method is in Microsoft Clarity + Claude.

4. Creative volume with Claude and Higgsfield

Meta rewards creative volume. Video budgets don't. This SOP closes the gap without letting AI pretend to be a customer.

SOP: creative volume
PartHow we run it
InputsCompetitor ads, the client's call transcripts, our seven creative variables filled in for the client, brand assets
System of recordAd account; a creative library with every asset's prompt
AI taskClaude writes hooks and Higgsfield prompts; Higgsfield produces statics, GIFs and short avatar clips
Human approvalA media buyer picks what launches; every asset gets an ad-policy and AI-disclosure check
Failure handlingRejected renders are logged with their prompt so the next batch improves
MeasureCost per lead and cost per acquisition against a human-footage control

The seven variables are Victory's creative brief: avatar, problem, outcome, niche, the common advice we'll challenge, a sourced surprising stat and identity markers. They work as a prompt schema, so every hook Claude writes starts from the same facts.

Budget for credits, not "unlimited". Higgsfield's Plus plan listed at $59 a month for 1,200 credits at the time of writing.[15] Through its MCP connector, every generation deducts credits on any plan; unlimited models and free generations are web-only.[16] In our experience, an active client creative program runs $300 to $2,000 a month in AI video credits, with statics and GIFs making up most of the output.

The compliance pass isn't optional. Meta labels ads made or significantly edited with third-party AI tools with "AI info".[17] An AI avatar can't pose as a real customer giving a testimonial. The full workflow and the policy check are in Higgsfield AI for ad creative.

5. Cross-system bridges in n8n and GoHighLevel

Most funnel failures we're called in to fix aren't strategy failures. They're data that didn't move: a webinar attendee who never got tagged as attended, a sales disposition that never reached the attribution tool.

SOP: cross-system bridges
PartHow we run it
InputsWebhooks and API events from webinar, checkout, CRM and attribution tools
System of recordGoHighLevel for contacts and pipeline; the source tool for its own events
AI taskClaude drafts and documents the workflow logic; the run itself is deterministic, not a model deciding each time
Human approvalA developer reviews every new bridge against test records before it goes live
Failure handlingRetries, dedupe keys, an error workflow that alerts a named person, credentials owned by the agency account
MeasureFailed runs per week; time to fix

Our rule of thumb is that a day-one coaching or event build in GoHighLevel has 12 to 20 workflows. Every one of them is a place data can stop moving. GHL owns contact-level nurture and pipeline stages; n8n owns the bridges between systems. The split is in GoHighLevel workflows vs n8n, and the CRM side is in our GoHighLevel guide.

How the plumbing connects: MCP without the hype

Most of the SOPs above depend on one piece of plumbing: the Model Context Protocol.

Anthropic open-sourced MCP on 2024-11-25 as an open standard for connecting AI tools to the systems where data lives.[18] It's now supported by Claude, ChatGPT, VS Code, Cursor and others, and its own docs describe it as a USB-C port for AI applications.[19] That's a good analogy, and it's worth finishing. A USB-C port tells you a cable will fit. It says nothing about what's allowed to travel down it.

Keep three things separate:

  1. Connection. MCP lets Claude talk to Fathom, ClickUp, Clarity, Higgsfield or your CRM.
  2. Permission. Claude's connectors inherit the user's permissions in the source system. On Team and Enterprise plans, owners can set each tool to Always allow, Needs approval or Blocked.[20]
  3. Consent. Whether the client agreed to have their calls recorded, processed and stored. No protocol settles that for you.

Anthropic also warns that custom connectors can be targeted by prompt injection, where instructions hidden in a document or web page try to steer the model.[20] That's the practical reason for our setup rule.

How we set up a new connector

  • Start read-only. Let it read for a week before it writes anything.
  • Set every write action to "Needs approval".
  • Block delete actions outright unless the job truly requires them.
  • Connect it from the agency's account, not a contractor's, so access survives staff changes.
  • Note the vendor's rate limits before you depend on it. Clarity allows 10 requests a day per project; ClickUp's limits depend on the plan.[12][11]
  • Write down who owns it and who gets told when it breaks.

What it costs to run and maintain

The subscription is the smallest line. Plan for five costs:

  • Seats. Claude's Team plan listed at $25 per standard seat monthly, or $20 billed annually, at the time of writing.[21]
  • Credits. Agents and generators bill by usage. Manus Pro starts at $20 a month for 4,000 credits.[22] Higgsfield bills credits per generation, and the burn varies by model and resolution.
  • API calls. Daily caps on connectors are a real constraint. Hit one mid-week and the workflow stops until it resets.
  • Review time. Every approval point costs minutes. That time belongs in the business case.
  • Broken integrations. Vendors change APIs, tokens expire, a field gets renamed. Someone owns the fix.

Then measure honestly. Self-reported ROI is everywhere. In a SAS and Coleman Parkes survey, 93% of CMOs and 83% of marketing teams using generative AI said they saw a clear ROI.[23] The release doesn't say how ROI was measured, and SAS sells marketing AI software. Gartner's CMOs said genAI ROI shows up mostly as time efficiency (49%), cost efficiency (40%) and capacity to do more (27%).[5] HubSpot found about a third of teams saying AI saves them 10 to 14 hours a week and another third saying 15 or more, also self-reported.[3]

In our experience, AI saves our ops team 15 to 20 hours a week across call notes, task plans and reporting. We trust that figure because we wrote down what those jobs took before we changed them. Do the same. Before any pilot, record hours per task, turnaround time and the error rate for the jobs you plan to hand over. Measure again after, and include the review time.

The most rigorous evidence is older and outside marketing, but it's useful. An NBER study of 5,179 customer support agents found that access to an AI assistant raised issues resolved per hour by 14% on average and by 34% for novice and lower-skilled workers.[24] It dates from 2023. The pattern we see matches it: AI lifts the floor of a team more than its ceiling.

Data-safety rules for lead and transcript data

Funnel businesses handle sensitive data: names, phone numbers, call recordings, sometimes payment details and health or financial context. AI doesn't change the rules. It adds new places for data to leak.

The rules we run:

  1. Business tiers only for client data. Anthropic says inputs and outputs from its commercial products, including Claude for Work and the API, aren't used to train models by default.[25] API inputs and outputs are deleted from its back end within 30 days, with stated exceptions.[26] Consumer plans have different defaults, so check the plan before the transcript.
  2. One context per client. Data never crosses accounts. No shared "all clients" project.
  3. Redact before upload. Payment details, government IDs and sensitive health or financial fields come out first.
  4. Scope every connector. Read-only by default, approval on writes, deletes blocked.
  5. Never publish a prospect's words. Call transcripts feed our thinking. They don't get quoted in ads or content without permission.
  6. Keep regulated-category data out of shared tools. If the category is sensitive, the data stays in the system built for it.

Governance is rare. Only 13% of organisations in SmarterX's 2026 survey had all four foundations in place: an AI roadmap, an AI council, a generative AI policy and an AI ethics policy.[1] The top barriers it found were lack of education and training (38%) and lack of awareness or understanding (35%).[1] A one-page data policy and an hour of training puts a team ahead of most of the market. Our version is the AI data-safety checklist for agencies.

Where human taste still wins

AI has habits. In our experience, Claude left alone writes short staccato sentences in runs of two and three words, and scatters em dashes through copy. We edit those out, and we write our instructions to say what we want ("write in flowing paragraphs") rather than what we don't. Every model has some version of this. Someone with taste has to catch it.

Some things shouldn't be handed over at all:

  • The offer. AI can pressure-test an offer. It can't tell you what your market will pay for.
  • What launches. A model can rank twenty hooks. A media buyer who has watched the account for months picks the three that run.
  • Proof. Testimonials, results and founder stories have to be real. AI can edit them. It can't stand in for them.
  • Internal tools. A tool built in an afternoon by prompting a model looks finished and often isn't. Someone has to maintain it when the person who built it is gone. If a tool will run client work, it gets the same engineering review as anything else.

Some of our best operators barely use AI, and their work is still the standard. AI raises the floor of a team. People with judgment set the ceiling.

That's the thread through this whole guide, and it's what McKinsey's high performers did: redesign the workflow, and decide in advance where a person checks the output.[2]

Where to start

Don't start with a tool. Start with one recurring job that eats hours every week, such as call notes, weekly reporting or sales-call reviews. Write down how long it takes today. Write the SOP: inputs, owner, AI task, approval point, failure handling. Run it for a month. Measure it against the baseline. Then pick the next job.

If you want to see how we'd set this up inside your funnel, sales floor and CRM, see how we build it.

Frequently asked questions

Sources

  1. 1.The 2026 State of AI for Business. SmarterX / Marketing AI Institute, 2026-05-18.
  2. 2.The state of AI in 2025: Agents, innovation, and transformation. McKinsey, 2025-11-05.
  3. 3.2026 State of Marketing. HubSpot, 2026-04-10 (updated).
  4. 4.State of Marketing, 10th edition. Salesforce, 2026-02-19.
  5. 5.Gartner 2025 CMO Spend Survey reveals marketing budgets have flatlined at seven percent of overall company revenue. Gartner, 2025-05-12.
  6. 6.Building effective agents. Anthropic Engineering, 2024-12-19.
  7. 7.Measuring AI ability to complete long tasks. METR, 2025-03-19.
  8. 8.Gartner predicts over 40 percent of agentic AI projects will be canceled by end of 2027. Gartner, 2025-06-25.
  9. 9.Fathom MCP integration. Fathom Help Center, undated, viewed 2026-10-04.
  10. 10.How Anthropic uses Claude in marketing. Anthropic (Claude blog), 2026-01-26.
  11. 11.What is ClickUp MCP?. ClickUp Help Center, viewed 2026-10-04.
  12. 12.Microsoft Clarity MCP server. Microsoft Learn, 2026-06-24.
  13. 13.Copilot in Clarity overview. Microsoft Learn, 2025-05-12 (updated 2025-12-05).
  14. 14.2026 Digital Experience Benchmark: frustration. Contentsquare, 2026-04-01.
  15. 15.Higgsfield pricing. Higgsfield, viewed 2026-10-04.
  16. 16.How do I connect Higgsfield to an AI agent?. Higgsfield Help Center, 2026-09-24.
  17. 17.AI info labels on ads. Meta Help Center, updated 2026, viewed 2026-10-04.
  18. 18.Introducing the Model Context Protocol. Anthropic, 2024-11-25.
  19. 19.What is the Model Context Protocol?. Model Context Protocol, 2026 (spec version 2026-07-28).
  20. 20.Use connectors to extend Claude's capabilities. Anthropic (Claude Help Center), updated week of 2026-10-04.
  21. 21.Claude plans and pricing. Anthropic, viewed 2026-10-04.
  22. 22.What is the current membership pricing for Manus?. Manus Help Center, updated week of 2026-10-04.
  23. 23.New study: GenAI hype is over as 93% of CMOs see strong ROI. SAS / Coleman Parkes, 2025-09-26.
  24. 24.Generative AI at Work (Working Paper 31161). NBER (Brynjolfsson, Li, Raymond), 2023-04 (rev. 2023-11).
  25. 25.Is my data used for model training?. Anthropic Privacy Center, 2026-08-18 (updated).
  26. 26.How long do you store my organization's data?. Anthropic Privacy Center, 2026-07-01 (updated).
Ray Gillespie

Written by

Ray Gillespie

Co-Founder & COO

Ray runs day-to-day operations across every Victory engagement, building the systems, automations and AI-powered workflows that hold the machine together. He has overseen operations behind more than $120M in revenue.

Strategy call

Want us to run the numbers on your funnel?

Book a call with Ray and Devin. Bring your show rates, CPLs and close rates. You leave with the one constraint we would fix first.

Free Revenue Leak Diagnostic

Where is your revenue leaking?

Pick the areas you suspect

No pitch, no pressure. Just a prioritized action plan.

More in AI