Skip to content

AI Agents in Practice (2026): What Actually Works (and What’s Still Hype)

A practical guide: agent architectures, tool use, guardrails, evaluations, and why most agent projects fail on data ownership and monitoring.

Attila Arndt
Attila Arndt

Triple A Digital, Cologne · · 3 min read

Who it's for
Technical and marketing leads who want to run an AI agent in production, not just try one out.
What you'll be able to do
Plan an agent along the reference architecture: orchestration, tools with minimal permissions, guardrails, monitoring. And measure what it is worth.
As of
February 2026
TL;DR

In 2026, agents only work in production if tool access, data flows, guardrails, and evaluation are designed properly. If you say “agent” but really mean a chat box, you’re in for nasty surprises once it runs in production.

What actually counts as an “agent” in 2026

For me, a system only becomes an agent once it understands an objective and plans the steps towards it — plan first, then execute. It has to use tools: APIs, databases, the CMS, the ads manager. It has to check its results through guardrails and evals. And it has to leave an audit trail behind: logs, decisions, inputs and outputs.

Rule of thumb: The model is rarely the bottleneck. Integration + control is.

Reference architecture (robust, not flashy)

Five building blocks, top to bottom
  1. 01
    OrchestrationRouting and state: which step runs when. Plus timeouts, retries, abort criteria and a cost limit per run.
  2. 02
    Tools with minimum permissionsCMS, analytics, ads, CRM, Search Console. The agent may create drafts, not publish them.
  3. 03
    KnowledgeBrand voice with examples and boundaries, your services and ideal customer profile, case studies with proof points, and the compliance rules you have to follow.
  4. 04
    GuardrailsPII filters, link checks, the claims policy and formatting rules.
  5. 05
    Evaluation and observabilityLogging of tool calls, cost, latency and error types. Quality through fact, style and SEO checks — and offline evals before the agent goes live.
Each section below explains one of them in detail.

1) Orchestration (agent runtime)

This is where routing and state live — which step runs when. It also holds timeouts, retries and abort criteria, plus a cost limit per run.

Rule: Reasoning ≠ execution. Secrets don’t belong in prompts.

2) Tools (minimum permissions)

Typical marketing/content tools:

  • CMS (headless / git / WordPress)
  • Analytics (Plausible/GA4)
  • Ads (Meta/Google)
  • CRM (HubSpot/Pipedrive)
  • Search Console

Least privilege: drafts yes, auto-publish no.

3) Knowledge (RAG / guidelines)

The agent needs the brand voice with examples and do’s and don’ts, your services and offers along with the ideal customer profile, case studies with proof points — and the compliance rules the company has to follow.

4) Guardrails (policy engine)

That means PII filters, link checks, the claims policy further down, and formatting rules — when a list beats a wall of text, for instance.

5) Evaluation & observability (not optional)

If you don’t measure, you’re guessing. Logging covers tool calls, cost, latency and error types. Quality is measured through fact checks, style checks and SEO checks. And before the agent goes live, it runs against offline evals: real briefs with gold outputs.

Failure modes I see all the time (and the fixes)

Three failure modes and their fix
Failure modeTwelve API calls and no decision.
The fixA step budget with a hard maximum, hard stop criteria, and a definition of done written as a checklist.
Failure modeHallucinated features and sources.
The fixConcrete claims only with primary sources — everything else framed as recommendations.
Failure modeUncontrolled publishing.
The fixHuman-in-the-loop approvals, plus separate roles and service accounts.
The sections below explain each failure mode on its own.

“Tool spam” instead of outcomes

The symptom is twelve API calls and no decision. The fix is a step budget with a hard maximum, hard stop criteria, and a definition of done written as a checklist.

Hallucinated features / sources

The only fix is a hard rule: concrete claims only with primary sources, everything else framed as recommendations.

Uncontrolled publishing

The fix is human-in-the-loop approvals, plus separate roles and service accounts.

ROI framework: make it measurable

I track four layers:

  1. Efficiency: minutes per asset, cost per asset
  2. Quality: error rate, review loops
  3. Performance: CTR/CVR/rankings/funnel metrics
  4. Learning: how fast learnings update guidelines
Production checklist (copy/paste)
  • agent can only create drafts (no auto-publish)
  • minimal tool scopes (least privilege)
  • sources required for concrete claims
  • logging + cost limits per run
  • evaluation test set (10–30 real briefs)
  • review flow: who approves what?

Claims policy (for Triple A Digital)

Product features, benchmarks, pricing and legal claims only stand with a primary source — otherwise they go. Best practices are labelled as recommendations, and numbers only appear with a source and a date.

Next step

If you want, I can build an agent in 7–14 days that ships a weekly structured draft in German and English: the outline and keywords, the draft itself as MDX, a fact-check list, and short texts for social networks.

Contact

Share your industry, target customers, and tooling stack — I’ll propose the right architecture.

Questions

Answered in brief

When is a system really an AI agent and not just a chat box?

Once it understands an objective and plans the steps towards it — plan first, then execute. It has to use tools, check its results through guardrails and evals, and leave an audit trail behind: logs, decisions, inputs and outputs.

May an agent publish on its own?

No. The agent may create drafts, not publish them — with minimum permissions on every tool. On top of that come human-in-the-loop approvals, separate roles and separate service accounts.

Where do agent projects fail most often?

Not on the model. The agent spams tools instead of deciding — the fix is a step budget with hard stop criteria. It hallucinates features and sources — the fix is requiring primary sources for concrete claims. And it publishes without control — the fix is human-in-the-loop approval.

Related

What I do in this area


Keep reading

Attila Arndt

Attila Arndt · Triple A Digital, Cologne

Is there a process like this in your company?

Pick a time that suits you. In the intro call, we'll work out which process is worth tackling first — and whether I'm the right person for it.

Free intro call (opens in a new tab)

Or email me directly: hello@tripleadigital.de · I reply within 48 hours.