
- Who it's for
- Technical and marketing leads who want to run an AI agent in production, not just try one out.
- What you'll be able to do
- Plan an agent along the reference architecture: orchestration, tools with minimal permissions, guardrails, monitoring. And measure what it is worth.
- As of
- February 2026
In 2026, agents only work in production if tool access, data flows, guardrails, and evaluation are designed properly. If you say “agent” but really mean a chat box, you’re in for nasty surprises once it runs in production.
What actually counts as an “agent” in 2026
For me, a system only becomes an agent once it understands an objective and plans the steps towards it — plan first, then execute. It has to use tools: APIs, databases, the CMS, the ads manager. It has to check its results through guardrails and evals. And it has to leave an audit trail behind: logs, decisions, inputs and outputs.
Rule of thumb: The model is rarely the bottleneck. Integration + control is.
Reference architecture (robust, not flashy)
- 01OrchestrationRouting and state: which step runs when. Plus timeouts, retries, abort criteria and a cost limit per run.
- 02Tools with minimum permissionsCMS, analytics, ads, CRM, Search Console. The agent may create drafts, not publish them.
- 03KnowledgeBrand voice with examples and boundaries, your services and ideal customer profile, case studies with proof points, and the compliance rules you have to follow.
- 04GuardrailsPII filters, link checks, the claims policy and formatting rules.
- 05Evaluation and observabilityLogging of tool calls, cost, latency and error types. Quality through fact, style and SEO checks — and offline evals before the agent goes live.
1) Orchestration (agent runtime)
This is where routing and state live — which step runs when. It also holds timeouts, retries and abort criteria, plus a cost limit per run.
Rule: Reasoning ≠ execution. Secrets don’t belong in prompts.
2) Tools (minimum permissions)
Typical marketing/content tools:
- CMS (headless / git / WordPress)
- Analytics (Plausible/GA4)
- Ads (Meta/Google)
- CRM (HubSpot/Pipedrive)
- Search Console
Least privilege: drafts yes, auto-publish no.
3) Knowledge (RAG / guidelines)
The agent needs the brand voice with examples and do’s and don’ts, your services and offers along with the ideal customer profile, case studies with proof points — and the compliance rules the company has to follow.
4) Guardrails (policy engine)
That means PII filters, link checks, the claims policy further down, and formatting rules — when a list beats a wall of text, for instance.
5) Evaluation & observability (not optional)
If you don’t measure, you’re guessing. Logging covers tool calls, cost, latency and error types. Quality is measured through fact checks, style checks and SEO checks. And before the agent goes live, it runs against offline evals: real briefs with gold outputs.
Failure modes I see all the time (and the fixes)
“Tool spam” instead of outcomes
The symptom is twelve API calls and no decision. The fix is a step budget with a hard maximum, hard stop criteria, and a definition of done written as a checklist.
Hallucinated features / sources
The only fix is a hard rule: concrete claims only with primary sources, everything else framed as recommendations.
Uncontrolled publishing
The fix is human-in-the-loop approvals, plus separate roles and service accounts.
ROI framework: make it measurable
I track four layers:
- Efficiency: minutes per asset, cost per asset
- Quality: error rate, review loops
- Performance: CTR/CVR/rankings/funnel metrics
- Learning: how fast learnings update guidelines
- agent can only create drafts (no auto-publish)
- minimal tool scopes (least privilege)
- sources required for concrete claims
- logging + cost limits per run
- evaluation test set (10–30 real briefs)
- review flow: who approves what?
Claims policy (for Triple A Digital)
Product features, benchmarks, pricing and legal claims only stand with a primary source — otherwise they go. Best practices are labelled as recommendations, and numbers only appear with a source and a date.
Next step
If you want, I can build an agent in 7–14 days that ships a weekly structured draft in German and English: the outline and keywords, the draft itself as MDX, a fact-check list, and short texts for social networks.
Share your industry, target customers, and tooling stack — I’ll propose the right architecture.
