Skip to content
AI in Marketing

AI Agents for Marketing: Manus, Hermes and Claude Agents Evaluated (What They Can and Can't Do Yet)

An operator's evaluation of Manus, Hermes Agent and Claude agents: where each earns its keep, where it fails, and the human checkpoints every run needs.

Ray GillespieRay GillespieCo-Founder & COO

Published 10 min read

A long agent run drawn as a track of steps under a falling reliability curve, with three gold human checkpoints before send, spend and write
On this page

Key takeaways

  • Agents earn their keep on long, bounded, read-heavy jobs with checkable output, such as a full competitor ad-library audit or one standing monitoring task per ad account.
  • Benchmarks separate answering from finishing. The best agent answered 76.9% of real advertising-analytics requests correctly, but the top agent on paid freelance projects delivered 2.5% of them to a client-acceptable standard.[1][2]
  • Manus is capable and carries vendor risk. After its Meta deal was unwound, affected accounts lost data created since late December 2025 unless they backed it up in August 2026.[3]
  • Hermes Agent is open-source, self-hosted and fast-moving. We've tested it. We don't run client work on it.
  • Put a human checkpoint before anything is sent, spent, written or deleted, and keep credentials out of the agent's reach.[4]

Use an agent when the job is long, bounded and easy to check, and its judgment saves more time than reviewing it costs. Use chat or a fixed workflow when you know the steps. Either way, put a person between the agent and anything that sends, spends, writes or deletes.

That's our evaluation of Manus, Hermes Agent and Claude's agent modes in brief. It isn't a sales page for any of them.

Chat, workflow, agent: the 60-second version

  • Chat: you ask, it answers, you decide. Best for single steps you want to steer.
  • Workflow: fixed steps in a fixed order, in a tool like n8n or GoHighLevel. Best when the steps are known and repeat (see GoHighLevel workflows vs n8n).
  • Agent: the model picks its own steps and tools until the job is done. Best when the path can't be scripted.

Anthropic's engineering guidance frames the trade-off: agentic systems often trade latency and cost for better performance, and autonomy brings the potential for compounding errors. Start with the simplest thing that works, and test agents in a sandbox with guardrails.[5]

Most marketing work that seems to need an agent is a workflow nobody has written down. Hormozi's 3Ds (document, demonstrate, duplicate) were built for training people, and they apply to agents unchanged: if you can't document the task, an agent won't do it reliably.

How we evaluated them

We judged each tool by job type, not feature list: a research audit, ongoing monitoring, a build, and outreach.

Framework

Five questions before any agent run

  1. Can a person check the output quickly? A spot-checkable audit is a good agent job. A strategy you'd have to redo to verify is not.
  2. How long is the run? Reliability falls as tasks get longer.
  3. Does it write to anything? Reading a library is low risk. Changing a budget, CRM record or send list is not.
  4. Can you recover the work? If the vendor changes, can you export outputs and task history?
  5. What does a finished job cost? Credits, concurrency and review time, not the plan price.

Victory's evaluation criteria. The document-first rule is credited to Alex Hormozi's 3Ds; the time-and-effort test to Hormozi's value equation.

The tie-breaker is Hormozi's value equation: run an agent only when it cuts time and effort by more than the review adds back. Benchmarks show the gap between answering and finishing.

Answering vs finishing: three agent benchmarks
BenchmarkWhat it measuresBest result
AD-Bench (2026)225 real advertising-analytics requests; answers only, no account changes76.9% right on the first try, 61.4% on the hardest tier[1]
TheAgentCompany (2025)175 tasks in a simulated company, fully completed30.3% (Claude 3.7 Sonnet: 26.3%)[6]
Remote Labor Index (2025)240 paid freelance projects, delivered at a client-acceptable standard2.5%, by Manus, the top agent[2]

Different tasks, models and dates, so read the pattern, not a ranking. TheAgentCompany and Remote Labor Index results use 2025-era models.

Agents already answer questions about data well. Finishing a deliverable a client would accept is a different job. On the Remote Labor Index, Manus earned $1,720 against $143,991 paid to the humans who did the same projects.[2]

Marketing evidence is thin. Besides AD-Bench, the only marketing benchmark we found is xbench's influencer-matching track (50 advertiser briefs, 836 candidates).[7] Nobody has published a Manus, Hermes and Claude head-to-head on marketing tasks.

Adoption is early. In Salesforce's latest State of Marketing, 13% of marketers had adopted agentic AI, against 75% using AI at all.[8] Gartner forecasts that more than 40% of agentic AI projects will be cancelled by the end of 2027, and counts only about 130 genuine agentic vendors among thousands claiming the label.[9] A forecast, but a fair warning about "agent washing".

Manus: where it earns its keep

Ray's line is that Manus is "for more complex tasks that I'm not going to vibe code with Claude": long, multi-step jobs you'd otherwise babysit in a chat.

The full ad-library audit

Our best use is a full audit of a competitor's Meta ad library. Manus records each ad's hook, format, offer and running time in a structured file. A person spot-checks a sample against the live library, and Claude turns the file into hook gaps for the next brief in our Claude-to-Higgsfield creative workflow. Read-only, long and checkable: the shape agents handle well.

MIT Technology Review's March 2025 hands-on found Manus promising on bounded research but prone to crashes and incomplete lists. After about three hours on one job it had produced 3 of 50 requested profiles.[10] An early version, but the lesson holds: scope tightly.

One task per ad account

Cost and concurrency

Manus's free plan gives 300 daily credits, one concurrent task and two scheduled automations. Pro starts at $20 a month for 4,000 credits or $40 for 8,000, with 20 concurrent tasks and 20 scheduled automations. Team starts at $20 a seat and adds SSO and a data-training opt-out.[12]

For monitoring, concurrency is usually the limit: the free plan covers one account. And since the training opt-out is a Team feature, check the training setting on lower plans before uploading client data.

Manus: where it fails

No task-data backup

Meta agreed to buy Manus in December 2025. In April 2026, China's NDRC ordered the deal withdrawn.[13] By August, the roughly $2 billion deal had been unwound and Manus said it would run independently again.[14]

Some users lost work. Manus's notice gave a backup window of August 11 to 23, 2026, then deleted affected accounts' data generated since December 29, 2025, on August 23 to 25.[3] Not every user was affected, but those who skipped the backup lost their task history. Backups came as separate account and task-data archives.[15]

Retention cuts the other way: after account deletion, Manus keeps personal data for no fixed period, "as long as necessary" for legal, accounting and fraud purposes.[16]

Our rule: treat any agent's workspace as temporary and export every output to your own drive.

Drift on long runs

Reliability falls with task length. METR's 2025 study found frontier agents succeeded almost every time on tasks taking a skilled human under 4 minutes, and under 10% of the time on tasks over about 4 hours.[17] The gap is closing fast. In METR's 2026 assessment, the best model it tested had a 50% time horizon of 16 to 20 hours on its software tasks, with the public frontier around 12 hours. At 80% reliability, the horizon was 3 to 4 hours.[18] Horizons have doubled roughly every 131 days since 2023.[19]

Caution: these are software tasks, and a 50% horizon means half of runs that long fail. Plan client work around the 80% figure, in checkable stages.

Vendor status

Manus went from independent to sold and back inside a year. That's no verdict on the product, but any vendor can change ownership or terms, so score portability first. Also: Manus's widely repeated launch benchmark scores trace only to the vendor's own March 2025 post, so we don't use them.

Hermes Agent: what testing showed

"Hermes" here means Hermes Agent by Nous Research, not Nous's Hermes language models. It's an MIT-licensed, self-hosted agent that writes and refines its own skills, keeps persistent memory and searches its past sessions. It has 40+ tools with MCP support, messaging gateways such as Slack and Telegram, a scheduler and subagents, and runs on any model provider.[20]

It moves fast. The repository was created in July 2025, tagged its first release in March 2026 and reached v0.21.5 on September 24, 2026, with about 251,000 GitHub stars.[20] It's still pre-1.0.

Strengths: you control hosting and data, you choose the model, and the learning loop smooths repeated tasks.

Costs: you run, patch and secure it. One independent audit of v0.8.0, filed as a GitHub issue, reported 4 critical and 9 high findings in the default configuration, including an unrestricted local shell and a mode that disables all checks.[21] That's one reporter's classification of an old version, and we haven't confirmed what's been fixed.

Supply chain is the wider risk. Ars Technica reported in August 2026 that coding agents including Hermes, Claude and Codex installed unowned packages listed in websites' llms.txt files: 227 install commands across 120 sites.[22] Any agent that installs what it reads can be steered.

We tested Hermes as a learning exercise and don't run client work on it. For an agency, the deciding factor isn't capability. It's that every patch, credential and failure is yours.

Claude as an agent: Cowork, Claude Code and connectors

Claude has agent modes too. In Claude Cowork, Manual mode (the default) asks before acting. Auto approves read-only actions and lets Claude decide on writes and deletes. Skip approves everything except blocked tools.[23] On anything that writes to a client system, we keep Manual.

For agents built on the Claude Agent SDK, the docs are clear that an allowedTools list on its own isn't a lockdown. Deny rules and PreToolUse hooks apply even in bypass mode, so that's where hard limits belong.[24]

How Claude compares on agency tasks is in Claude vs ChatGPT for marketing; where each tool sits is in our AI tools lane map.

The human checkpoints every agent run needs

Stated trust runs ahead of safe delegation. In Salesforce's survey, 81% of marketers said they would trust AI to respond to customers.[25] Anthropic's red team shows why that needs a boundary: in a controlled exercise, Claude completed a pasted credential-exfiltration prompt in 24 of 25 retries. The fix was architectural: credentials never enter the sandbox.[4]

The ad platforms won't supply the checkpoint. Meta's Marketing API caps ad-set budget changes at 4 an hour and spending-limit changes at 10 a day: rate limits, not review.[26] Google makes advertisers review campaigns and assets its own features generate, which says nothing about your agent.[27]

OWASP calls the risk Excessive Agency: damage from manipulated or unexpected output when a system has too much permission or autonomy.[28] Its agentic top 10 runs from goal hijack to rogue agents.[29] In NIST's red-team competition, more than 400 participants made over 250,000 attack attempts, and at least one succeeded against every frontier model tested.[30]

Checkpoints on every agent run

  • Before send: a person approves every email, text, DM and call script.
  • Before spend: the agent proposes budget and bid changes; a media buyer applies them.
  • Before write: CRM, ad account and site changes get reviewed.
  • Before delete: agents don't delete anything.
  • Credentials outside the sandbox: scoped, read-first tokens only, never passwords or admin keys.
  • No unreviewed installs: no packages or skills found on web pages.
  • Stop on conflict: if the data disagrees with the brief, the agent stops and reports.
  • Export on finish: every output lands in your own drive.

The verdict

Our verdict by job type
JobManusHermes AgentClaude (Cowork, Claude Code)
Research auditStrong on long, read-only jobsCapable if well hostedStrong when it fits in context
MonitoringOne task per account, human approvesScheduler; uptime is yoursScheduled tasks, writes on Manual
BuildsLong multi-step buildsSuits technical teamsStrong with Claude Code
OutreachDrafts; a person sendsDrafts; a person sendsDrafts; a person sends
Recovering workWeak: affected accounts lost dataYours, if backed upExport outputs
How we use itResearch and monitoringTested, not used for client workDaily

Agents are improving fast. The checkpoint is yours to build. How agents fit the rest of our operation is in the AI marketing operations guide.

Want it set up with checkpoints? See how we build it.

Frequently asked questions

Sources

  1. 1.AD-Bench (arXiv 2602.14257v2). AD-Bench authors, arXiv, 2026-06-22 (v2).
  2. 2.Remote Labor Index. Scale AI and Center for AI Safety, 2025-10-29.
  3. 3.Service change overview: what's happening and am I affected?. Manus Help Center, 2026-08-21.
  4. 4.How we contain Claude. Anthropic Engineering, 2026-05-25.
  5. 5.Building effective agents. Anthropic, 2024-12-19.
  6. 6.TheAgentCompany (arXiv 2412.14161). Carnegie Mellon University et al., arXiv, 2025-09-10 (v3).
  7. 7.xbench (arXiv 2506.13651). xbench authors, arXiv, 2025-06-16, updated 2026-05-13.
  8. 8.Tenth State of Marketing report. Salesforce, 2026-05-28.
  9. 9.Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027. Gartner, 2025-06-25.
  10. 10.Manus AI review (hands-on test). MIT Technology Review, 2025-03-11.
  11. 11.The state of AI in 2025. McKinsey, 2025-11-05.
  12. 12.What is the current membership pricing for Manus?. Manus Help Center, updated week of 2026-10-04.
  13. 13.Meta's purchase of AI startup Manus halted by China. PYMNTS, 2026-04-27.
  14. 14.Manus, China and Meta's unwound acquisition. CNBC, 2026-08-11.
  15. 15.Service change overview: how to back up your data. Manus Help Center, 2026-08.
  16. 16.Will my personal data be deleted after I delete my account?. Manus Help Center, updated week of 2026-10-04.
  17. 17.Measuring AI ability to complete long tasks. METR, 2025-03-19.
  18. 18.Frontier risk report. METR, 2026-05-19.
  19. 19.Time horizon 1.1. METR, 2026-01-29.
  20. 20.Hermes Agent (README and release history). Nous Research, GitHub, viewed 2026-10-04.
  21. 21.Security audit of Hermes Agent v0.8.0 (issue #7826). GitHub issue on NousResearch/hermes-agent, 2026-04-11.
  22. 22.Claude, Codex and Hermes installed unowned code inside corporate networks. Ars Technica, 2026-08-27.
  23. 23.Get started with Claude Cowork. Claude Help Center, updated week of 2026-10-04.
  24. 24.Agent SDK: permissions. Anthropic (Claude Code docs), viewed 2026-10-04.
  25. 25.State of Marketing 2026. Salesforce, 2026-02-19.
  26. 26.Marketing API rate limiting. Meta for Developers, 2026-05-06 (updated).
  27. 27.Google Ads Terms: review of automatically generated campaigns and assets. Google Ads Help, 2026-10-03 (updated).
  28. 28.LLM06:2025 Excessive Agency. OWASP Gen AI Security Project, 2025.
  29. 29.OWASP Top 10 for Agentic Applications. OWASP Gen AI Security Project, 2025-12-09.
  30. 30.Insights on AI agent security from a large-scale red-teaming competition. NIST Center for AI Standards and Innovation, 2026-03-23.
Ray Gillespie

Written by

Ray Gillespie

Co-Founder & COO

Ray runs day-to-day operations across every Victory engagement, building the systems, automations and AI-powered workflows that hold the machine together. He has overseen operations behind more than $120M in revenue.

Part of the guide: AI Marketing Operations: How a Modern Agency Runs Funnels with Claude, Agents and Automation

Strategy call

Want us to run the numbers on your funnel?

Book a call with Ray and Devin. Bring your show rates, CPLs and close rates. You leave with the one constraint we would fix first.

Free Revenue Leak Diagnostic

Where is your revenue leaking?

Pick the areas you suspect

No pitch, no pressure. Just a prioritized action plan.

More in AI