Service · monthly retainer

Agents that still work six months after launch.

An agent that performed well in March can quietly degrade by June: the provider updated the model, someone edited a prompt, a data source changed shape. Managed AI operations is the discipline of noticing before your customers do. We monitor, evaluate, upgrade, control costs and handle incidents for your agents and LLM applications, and we review the results with you every month.

01

Why operations is the hard part

Shipping an agent is a project. Keeping it accurate, safe and affordable is a job, and most companies have not given that job to anyone.

Traditional software behaves tomorrow as it does today unless someone changes the code. AI systems do not. Model providers retire and replace models on short notice, update them in place, change pricing and rate limits and adjust safety behaviour. Your prompts, retrieval sources and tools change too. Any of these can shift output quality without a single error in the logs.

The consequences are rarely dramatic at first. Answers get slightly longer, a classification drifts, a tool call fails silently and the agent works around it. Nobody notices until a customer complains, a monthly invoice doubles or an auditor asks for evidence of how decisions were made.

AgentOps, as we practise it, borrows from what site reliability engineering did for web systems: observe everything, test every change, define what good looks like, respond to incidents with a plan and review regularly. The difference is that the thing observed is language rather than latency, which makes evaluation the centre of the work.

  • FormatMonthly retainer, scoped per system
  • Service levelsDefined per contract
  • ForCompanies with agents or LLM applications in production
  • OutcomeStable quality, controlled costs, audit-ready documentation
Need something built first? See agent development →
02

What is included

Scoped per system at the start, then run as a monthly service. The parts you already cover internally are simply left out.

i.

Monitoring and tracing

Every request, tool call, retrieval and model response traced end to end, with dashboards for quality signals, latency, error rates and cost. We work with tools such as Langfuse, LangSmith and OpenTelemetry-based stacks, and with what you already run.

ii.

Evaluation suites and regression tests

A curated set of real cases with expected outcomes for each agent, run automatically whenever a prompt, model, tool or data source changes. Nothing reaches production without passing the same tests it passed at launch.

iii.

Model and version upgrades

When a provider deprecates a model or releases a better one, we test the candidate against your evaluation suite, compare cost and quality, migrate in a controlled window and keep a rollback path open.

iv.

Cost and token budget control

Budgets per agent and per team, alerts before they are exceeded, and regular optimisation: caching, routing simpler requests to cheaper models, trimming context. Cost reports are part of the monthly review.

v.

Incident handling and fallbacks

Defined severities, response within the times set in your contract, fallback paths that keep the process running when a model or tool is down, and a written post-incident review for anything that affected users.

vi.

Abuse monitoring and audit documentation

Detection of prompt-injection attempts, misuse and unusual patterns, plus the logs, change records and evaluation results that a GDPR or EU AI Act review will ask for. See AI governance.

03

How it works

  1. Onboarding and baseline

    We inventory the systems in scope, connect tracing, build or extend the evaluation suite from real traffic and record a quality and cost baseline. Existing agents get a health check; issues found are fixed or documented.

    Weeks 1–3
  2. Runbooks and service levels

    Together we define severities, response times, escalation paths, fallback behaviour and who may approve a model change. Service levels are agreed per contract and written down.

    Weeks 2–4
  3. Steady-state operation

    Continuous monitoring, weekly quality and drift review, automated regression tests on every change, cost tracking and incident response. Changes to prompts and models go through the same controlled process every time.

    Ongoing
  4. Monthly review with the business owner

    One hour with the person who owns the process: quality trends, incidents, costs, upcoming model changes, improvement backlog. Decisions are recorded, and the report doubles as audit documentation.

    Monthly
  5. Quarterly deep review

    A wider look at each agent: is it still solving the right problem, has the process around it changed, should its scope grow or shrink, what would a newer model make possible. Recommendations go to the roadmap.

    Quarterly
04

Who this is for

  • Companies whose agents we built

    Most clients of our agent development and workflow automation services move straight into managed operations. The evaluation suite already exists; we keep it running.

  • Companies with agents built by others

    In-house or by a previous partner, and now the person who built it has moved on. We start with a health check, add tracing and evaluation, and take over operation without a rewrite.

  • Regulated businesses that need evidence

    If auditors, regulators or enterprise customers ask how your AI behaves and how you know, the logs, tests and change records from operations are the answer. See financial services.

  • Software companies with AI features in the product

    Product teams that want quality and cost under control without building an internal platform team first. See software & SaaS.

05

How we run operations

  • Evaluate before you change

    No model swap, prompt edit or new tool goes live without passing the regression suite. Boring by design.

  • Humans see what matters

    Alerts are tuned to the signals that predict user impact, not to everything that can be measured. The monthly review is with a business owner, not only with IT.

  • Vendor-neutral operations

    We run agents on OpenAI, Anthropic, Google, Mistral, Azure, AWS Bedrock, Vertex AI or EU-hosted open-weight models. If a provider change makes sense, we recommend it, migrate and stay neutral about the destination.

  • Documentation is a by-product, not a project

    Traces, test results and change records are generated by the process itself. When a review comes, nothing has to be reconstructed after the fact.

06

Tools and platforms we operate with

  • Langfuse
  • LangSmith
  • OpenTelemetry
  • Grafana
  • LangGraph
  • n8n
  • OpenAI
  • Anthropic Claude
  • Google Gemini
  • Mistral
  • Azure OpenAI
  • AWS Bedrock
  • Google Vertex AI

We name these because we work with them. We hold no partnerships or reseller agreements and choose tools per system.

07

Frequently asked questions

What response times do you offer?

They are defined per contract and per severity, because a customer-facing agent and an internal summarisation tool need different levels. Typically the highest severity carries a short response window with an on-call path, and lower severities are handled within business hours. The levels are written down during onboarding and reported against every month.

Do you need access to our production systems?

Yes, but scoped: tracing endpoints, model provider dashboards, the deployment pipeline and, where agreed, the relevant logs. Access follows least privilege, is documented and can be revoked at any time. A data processing agreement (Auftragsverarbeitungsvertrag) under the GDPR is standard, and where you host in the EU, nothing leaves your environment.

Can you operate agents we built ourselves or with another partner?

Yes. We start with a health check of the system, its prompts, tools and data sources, then add tracing and an evaluation suite if they are missing. Framework and platform matter little. We sometimes recommend changes, but taking over operation does not require a rebuild.

What does the monthly review cover?

Quality trends against the baseline, incidents and what was learned, cost per agent and per team against budget, upcoming model deprecations with their migration plan, and a prioritised list of improvements. About an hour with the business owner; the written version serves as audit documentation.

How do you handle a model being deprecated?

We track provider announcements, so a deprecation rarely comes as a surprise. The candidate model is run against your evaluation suite, cost and quality are compared, and if the results hold, we migrate in an agreed window with a rollback path, usually well before the provider's deadline.

Is this the same as MLOps?

They overlap but are not the same. MLOps is about training, deploying and monitoring your own models. AgentOps is about LLM-based systems that call external models and tools, where evaluation, prompt versioning, tool reliability and cost control matter more than training pipelines. If you train your own models, data foundations is the place to start.

08

Related services

Next step

Let's find the first workflow worth automating.

A 30-minute intro call, no slides and no obligation. We listen, ask about your processes, and tell you honestly where AI agents would pay off and where they would not.