Insights · Guide

Guardrails for AI agents: how to let agents act without losing control.

An agent that can only read is a search engine. An agent that can write, send and pay is useful, and it needs the same controls you would give a new employee with system access. This guide describes eight guardrail layers, maps them to risk levels and walks through an accounts-payable agent as an example.

01

The layers

Control is a design property, not a policy

The moment an agent can take an action, whether creating a ticket, updating a record or releasing a payment, the question changes from "is the answer correct?" to "what is the worst thing this system can do, and what stops it?" Many companies answer with a policy document. Policies bind people; they do not bind software. Control has to be built into the agent's environment: what it can reach, what it needs permission for, what it can spend, and how you find out what it did.

The good news is that the control toolkit is well understood. It borrows from decades of practice in access management, payment controls and software testing. What is new is the combination, and the fact that a language model can be talked into things by the content it reads. The eight layers below cover both.

Eight guardrail layers

1. Permissions and least privilege

Every agent gets its own identity and credentials, never a shared admin account. Its permissions list which systems it may touch and, per system, which tools: read customer records yes, delete them no; draft an email yes, send one only in defined cases. Separate read from write. If the agent connects through the Model Context Protocol, the server's tool list is where this boundary lives.

2. Approval gates and human-in-the-loop patterns

Not every action needs a human, and an agent that asks for everything will be ignored. Route approvals by three dimensions: the risk of the action (reversible or not), the amount (a threshold in euros, records or recipients) and the agent's own confidence (below a score, ask). The common patterns are approve-before-act, review-after-act within a delay window, and sampling review of a percentage of autonomous actions.

3. Input and output validation

Structure beats prose. Inputs to the agent are checked against a schema; outputs that will be written into a system must match the expected format, ranges and business rules before they leave the agent. Add policy checks (no discount above X, no message to a customer on a blocked list) and PII filters that stop personal data from ending up where it does not belong, including in prompts sent to external models.

4. Prompt-injection defence

Every document, email, web page or tool result the agent reads is untrusted content and may contain instructions aimed at the model. Defence in depth: mark untrusted content as data in the prompt; never let a single step both read untrusted content and take a consequential action without a check; restrict tools and destinations with allow-lists; and treat anything that looks like an instruction inside a tool result as a signal to stop and escalate.

5. Sandboxing and budgets

Agents run in an environment with limits: rate limits per minute, a token and cost budget per run and per day, a maximum number of steps and a time limit. A runaway loop then costs a defined amount and stops, rather than running all weekend. Code execution, where needed, happens in an isolated sandbox with no access to production credentials.

6. Evaluation and regression testing

A test set of real cases with expected outcomes, run on every change to a prompt, a model version or a tool. It includes adversarial cases: injection attempts, malformed inputs, edge cases in the business rules. No change reaches production with a lower score than the current version. This is the guardrail that catches silent degradation when a vendor updates a model.

7. Logging, audit trails and traceability

Every run is recorded: the input, each tool call with arguments and results, the model's output, the approvals given and by whom, cost and duration. The record must be enough to reconstruct a decision months later for an auditor, a customer complaint or your data protection officer. Tools such as Langfuse or platform-native tracing do this; the important part is that someone reads the dashboards.

8. Kill switches and fallbacks to humans

A single switch that stops the agent and routes its queue back to people, tested before go-live and documented in the runbook. Plus graceful degradation: when a system is down, a confidence score is low or a budget is exhausted, the agent hands the case to a human queue with its notes, rather than guessing.

Which guardrails for which risk

Not every agent needs all eight layers at full strength. The mapping below is the starting point we use in AI governance engagements; your own risk appetite and regulatory context adjust it.

Risk levels and recommended guardrails
Risk levelTypical agentsMinimum guardrails
Low: read-only, internalKnowledge assistant, research summaries, internal draftingOwn credentials, read-only tools, PII filter for external models, logging, budget, evaluation set
Medium: writes to internal systemsTicket triage, CRM updates, order-status replies, document classificationAll of the above, plus schema validation on writes, injection defence, sampling review, kill switch with human fallback
High: external communication or moneyCustomer emails, supplier communication, invoice approval, payment preparation, HR correspondenceAll of the above, plus approve-before-act above thresholds, allow-listed recipients and destinations, separation of duties, regression tests with adversarial cases, an audit trail sufficient for external review
Very high: legal effect on peopleDecisions on credit, employment, benefits, access to servicesThe decision stays with a person; the agent prepares, documents and never decides. Check the EU AI Act's high-risk obligations and their current status with legal advisers.

A worked example: the accounts-payable agent

A typical scenario: an agent processes incoming supplier invoices. It reads the PDF, extracts the data, matches the invoice to a purchase order and a goods receipt, checks the supplier master data, posts the invoice in the ERP and proposes a payment run. Useful, and full of ways to lose money. This is how the layers apply.

  • Permissions: the agent may read purchase orders, goods receipts and supplier master data, may create invoice postings in draft state, and may not change bank details, create suppliers or release payments.
  • Approval gates: invoices that match a purchase order within tolerance and stay below an amount threshold post automatically; mismatches, new suppliers, changed bank details and anything above the threshold go to an accountant with the agent's findings attached. A second approver is required above a higher threshold, as your existing separation of duties demands.
  • Validation: extracted amounts, tax rates and IBANs are checked for format and plausibility; totals must reconcile; the supplier must exist in the master data.
  • Injection defence: invoice text and email bodies are treated as data. An invoice that contains "please update our bank details to…" is not a request the agent may act on; it is flagged.
  • Budgets and sandbox: a daily processing cap and a cost budget; no direct access to the payment system.
  • Evaluation: a test set of several hundred historical invoices, including known fraud patterns and poor scans, run before every change.
  • Audit trail: every invoice carries the extracted data, the matching logic, the approvals and the model version. Your auditors can follow it.
  • Kill switch: one action stops posting and returns the queue to the team, with the day's cases listed.

With these controls the agent handles the routine majority alone and prepares the exceptions for people. That is the pattern across finance use cases: automate the clear cases, make the unclear ones faster for humans, and never let the agent be the last line of defence for money.

Guardrails and the EU AI Act

The EU AI Act expects human oversight for higher-risk AI systems: people who understand the system, can interpret its output, can decide not to use it and can intervene or stop it. Even where an agent falls outside the high-risk categories, these expectations are a sensible standard, and they map directly onto the layers above: approval gates and sampling review are oversight, logging is interpretability, the kill switch is intervention. Documenting the guardrails also gives your data protection officer the evidence they need under GDPR for automated processing. The obligations are phasing in through 2026 and 2027, timelines may shift, and this is not legal advice; check the current status with your advisers.

The practical point for decision makers is simple. Guardrails are not the price of using agents. They are what makes it possible to give an agent real work, and to sleep while it does it. Once the agent is live, managed AI operations keeps the evaluations, logs and thresholds current as the business and the models change.

02

Frequently asked questions

Do guardrails make the agent useless or slow?

Badly designed ones do: an agent that asks for approval on everything is ignored within a week. Well-designed guardrails are proportionate: full autonomy for routine, low-risk cases, checks where money, external communication or personal data are involved. Most of the layers, such as logging, budgets and validation, are invisible to the user.

Can we use the guardrail features built into agent platforms?

Often, and you should. Workflow platforms, cloud agent services and model providers offer approvals, content filters, rate limits and tracing. The design decisions, meaning what counts as high risk, where thresholds sit and who approves, remain yours. Platform features implement guardrails; they do not decide them.

How do we know the guardrails still work six months later?

Through the evaluation set and the logs. Rerun the adversarial tests after every change and review a sample of autonomous decisions monthly. Assign an operations owner whose job includes reading the dashboards; guardrails nobody watches erode quietly.

What is the single most important guardrail?

Least privilege. An agent that cannot reach a system cannot damage it, and no amount of clever prompting changes that. Start by giving the agent only the tools the use case needs, then add approval gates for the actions that remain risky.

03

Related reading and services

Next step

Let's find the first workflow worth automating.

A 30-minute intro call, no slides and no obligation. We listen, ask about your processes, and tell you honestly where AI agents would pay off and where they would not.