Service · monthly retainer
Agents that still work six months after launch.
An agent that performed well in March can quietly degrade by June: the provider updated the model, someone edited a prompt, a data source changed shape. Managed AI operations is the discipline of noticing before your customers do. We monitor, evaluate, upgrade, control costs and handle incidents for your agents and LLM applications, and we review the results with you every month.
Why operations is the hard part
Shipping an agent is a project. Keeping it accurate, safe and affordable is a job, and most companies have not given that job to anyone.
Traditional software behaves tomorrow as it does today unless someone changes the code. AI systems do not. Model providers retire and replace models on short notice, update them in place, change pricing and rate limits and adjust safety behaviour. Your prompts, retrieval sources and tools change too. Any of these can shift output quality without a single error in the logs.
The consequences are rarely dramatic at first. Answers get slightly longer, a classification drifts, a tool call fails silently and the agent works around it. Nobody notices until a customer complains, a monthly invoice doubles or an auditor asks for evidence of how decisions were made.
AgentOps, as we practise it, borrows from what site reliability engineering did for web systems: observe everything, test every change, define what good looks like, respond to incidents with a plan and review regularly. The difference is that the thing observed is language rather than latency, which makes evaluation the centre of the work.
- FormatMonthly retainer, scoped per system
- Service levelsDefined per contract
- ForCompanies with agents or LLM applications in production
- OutcomeStable quality, controlled costs, audit-ready documentation
What is included
Scoped per system at the start, then run as a monthly service. The parts you already cover internally are simply left out.
Monitoring and tracing
Every request, tool call, retrieval and model response traced end to end, with dashboards for quality signals, latency, error rates and cost. We work with tools such as Langfuse, LangSmith and OpenTelemetry-based stacks, and with what you already run.
Evaluation suites and regression tests
A curated set of real cases with expected outcomes for each agent, run automatically whenever a prompt, model, tool or data source changes. Nothing reaches production without passing the same tests it passed at launch.
Model and version upgrades
When a provider deprecates a model or releases a better one, we test the candidate against your evaluation suite, compare cost and quality, migrate in a controlled window and keep a rollback path open.
Cost and token budget control
Budgets per agent and per team, alerts before they are exceeded, and regular optimisation: caching, routing simpler requests to cheaper models, trimming context. Cost reports are part of the monthly review.
Incident handling and fallbacks
Defined severities, response within the times set in your contract, fallback paths that keep the process running when a model or tool is down, and a written post-incident review for anything that affected users.
Abuse monitoring and audit documentation
Detection of prompt-injection attempts, misuse and unusual patterns, plus the logs, change records and evaluation results that a GDPR or EU AI Act review will ask for. See AI governance.
How it works
Onboarding and baseline
We inventory the systems in scope, connect tracing, build or extend the evaluation suite from real traffic and record a quality and cost baseline. Existing agents get a health check; issues found are fixed or documented.
Runbooks and service levels
Together we define severities, response times, escalation paths, fallback behaviour and who may approve a model change. Service levels are agreed per contract and written down.
Steady-state operation
Continuous monitoring, weekly quality and drift review, automated regression tests on every change, cost tracking and incident response. Changes to prompts and models go through the same controlled process every time.
Monthly review with the business owner
One hour with the person who owns the process: quality trends, incidents, costs, upcoming model changes, improvement backlog. Decisions are recorded, and the report doubles as audit documentation.
Quarterly deep review
A wider look at each agent: is it still solving the right problem, has the process around it changed, should its scope grow or shrink, what would a newer model make possible. Recommendations go to the roadmap.
Who this is for
Companies whose agents we built
Most clients of our agent development and workflow automation services move straight into managed operations. The evaluation suite already exists; we keep it running.
Companies with agents built by others
In-house or by a previous partner, and now the person who built it has moved on. We start with a health check, add tracing and evaluation, and take over operation without a rewrite.
Regulated businesses that need evidence
If auditors, regulators or enterprise customers ask how your AI behaves and how you know, the logs, tests and change records from operations are the answer. See financial services.
Software companies with AI features in the product
Product teams that want quality and cost under control without building an internal platform team first. See software & SaaS.
How we run operations
Evaluate before you change
No model swap, prompt edit or new tool goes live without passing the regression suite. Boring by design.
Humans see what matters
Alerts are tuned to the signals that predict user impact, not to everything that can be measured. The monthly review is with a business owner, not only with IT.
Vendor-neutral operations
We run agents on OpenAI, Anthropic, Google, Mistral, Azure, AWS Bedrock, Vertex AI or EU-hosted open-weight models. If a provider change makes sense, we recommend it, migrate and stay neutral about the destination.
Documentation is a by-product, not a project
Traces, test results and change records are generated by the process itself. When a review comes, nothing has to be reconstructed after the fact.
Tools and platforms we operate with
We name these because we work with them. We hold no partnerships or reseller agreements and choose tools per system.
Frequently asked questions
What response times do you offer?
They are defined per contract and per severity, because a customer-facing agent and an internal summarisation tool need different levels. Typically the highest severity carries a short response window with an on-call path, and lower severities are handled within business hours. The levels are written down during onboarding and reported against every month.
Do you need access to our production systems?
Yes, but scoped: tracing endpoints, model provider dashboards, the deployment pipeline and, where agreed, the relevant logs. Access follows least privilege, is documented and can be revoked at any time. A data processing agreement (Auftragsverarbeitungsvertrag) under the GDPR is standard, and where you host in the EU, nothing leaves your environment.
Can you operate agents we built ourselves or with another partner?
Yes. We start with a health check of the system, its prompts, tools and data sources, then add tracing and an evaluation suite if they are missing. Framework and platform matter little. We sometimes recommend changes, but taking over operation does not require a rebuild.
What does the monthly review cover?
Quality trends against the baseline, incidents and what was learned, cost per agent and per team against budget, upcoming model deprecations with their migration plan, and a prioritised list of improvements. About an hour with the business owner; the written version serves as audit documentation.
How do you handle a model being deprecated?
We track provider announcements, so a deprecation rarely comes as a surprise. The candidate model is run against your evaluation suite, cost and quality are compared, and if the results hold, we migrate in an agreed window with a rollback path, usually well before the provider's deadline.
Is this the same as MLOps?
They overlap but are not the same. MLOps is about training, deploying and monitoring your own models. AgentOps is about LLM-based systems that call external models and tools, where evaluation, prompt versioning, tool reliability and cost control matter more than training pipelines. If you train your own models, data foundations is the place to start.
Related services
Custom AI agent development
Single- and multi-agent systems designed and built for production: tool use, memory, MCP servers, evaluation suites, guardrails and EU deployment.
Agentic workflow automation
Multi-step business processes automated end to end by LLM-powered agents across ERP, CRM, ticketing and email, with human approval steps built in.
AI governance, EU AI Act & GDPR
An AI register, risk classification under the EU AI Act, GDPR-aligned processes and a usage policy your teams will actually follow, built together with your lawyers and your data protection officer.
Guardrails for AI agents
Eight layers of control that let an agent act in your systems, a risk-to-guardrail mapping and a worked accounts-payable example.
Let's find the first workflow worth automating.
A 30-minute intro call, no slides and no obligation. We listen, ask about your processes, and tell you honestly where AI agents would pay off and where they would not.