Insights · Guide
Why AI pilots fail, and how to get an agent into production.
Most AI pilots never reach daily operation, and the reasons are rarely technical. This guide names the ten causes we see most often, defines what production-ready actually means, and lays out a phased path from pilot to a live agent with a clear owner.
The pattern
The pilot that never ends
Most companies do not lack AI pilots. They lack AI systems that do real work every day. The pattern is familiar: a demo impresses the leadership team, a pilot is approved, a small group builds something that works on the examples in the slide deck, and then the project enters a long twilight of "almost ready". Six months later the budget is gone, the sponsor has moved on, and the pilot is quietly retired with a lessons-learned document nobody reads.
This is not a technology problem. Today's models are good enough for a wide range of business tasks. Pilots fail for organisational and engineering reasons that are visible in the first two weeks, if you know where to look. This article names them, defines what production-ready actually means, and describes a path that gets an agent into daily operation with an owner.
Ten reasons pilots stall
In our experience, failing pilots share the same handful of causes. They cluster around three moments: before the pilot starts, during the build, and around the pilot in the organisation.
Before the pilot starts
No business owner. The pilot belongs to IT, an innovation team or an external vendor, but no department head has put their name on the outcome. Nobody fights for the agent when it inconveniences the process it changes.
No baseline and no success metric. Nobody measured how long the task took before, how many cases arrive per week or what an error costs. Without a baseline the pilot cannot succeed, because success was never defined.
The wrong first use case. Teams pick the most impressive scenario rather than the one with high volume, clear rules and tolerable error costs. The demo is spectacular and the path to production is impossible. Our guide to choosing a first agent use case covers the selection in detail.
During the build
A demo built on hand-picked data. The pilot works on twenty clean examples. Real inboxes contain scanned PDFs, missing fields, two languages and a customer who writes in capital letters. The first contact with unfiltered data is often the end of the pilot.
No evaluation suite. Every prompt change is judged by looking at a few outputs. Nobody can say whether version four is better than version three, so the team argues instead of measuring.
Integration and permissions left for later. The agent reads from an export file instead of the ERP and writes into a spreadsheet instead of the ticketing system. Connecting it properly, with credentials, roles and audit trails, turns out to be the actual project.
No plan for exceptions and errors. What happens when the agent is unsure, a system is down or the input is malformed? If the answer is "someone will notice", the agent cannot be left alone.
Around the pilot
Governance and the works council involved too late. Data protection, the EU AI Act and co-determination questions arrive in week ten as blockers instead of week one as design inputs. In German companies this alone can delay a launch by a quarter.
No operations owner after go-live. The consultants leave, the prompts drift, a vendor changes a model version, and there is nobody whose job it is to notice.
Change management ignored. The people whose work changes were informed, not involved. They route around the agent, and the usage numbers then prove that it "did not work".
What production-ready actually means
"Production-ready" is used loosely, so it helps to turn it into a checklist. An agent is ready when every row in the following table is true, and a pilot is a project whose purpose is to make them true.
| Area | Production-ready means | What a stalled pilot usually has |
|---|---|---|
| Evaluations | A test set of real cases with expected outcomes, run automatically on every change, with an agreed accuracy threshold | A few examples checked by eye |
| Observability | Every run is traced: inputs, tool calls, outputs, cost and latency, with alerts on failure rates | Console logs on a developer's laptop |
| Permissions | Its own credentials, least privilege per system, write actions gated by risk | A shared admin account |
| Fallbacks | Defined behaviour for low confidence, outages and malformed input, with a queue for humans | The agent guesses or stops |
| Cost controls | Budgets per run and per day, rate limits, a known cost per processed case | Nobody has looked at the invoice |
| Documentation | What the agent does, what it must not do, how to change it and how to switch it off | The prompt is the documentation |
| Owner | A named business owner for the outcome and a named operations owner for the system | The project team |
| Training | The affected team knows how to work with the agent, escalate and give feedback | An announcement email |
From pilot to production in four phases
The path below is the one we use in our agent pilots, which typically reach a production-grade state in six to eight weeks. The phases matter more than the calendar.
Phase 1: Frame the case
One to two weeks. Confirm the business owner, measure the baseline, define the success metric and the error budget, and write down the exception policy: which cases the agent handles alone, which it prepares for a human, which it must not touch. Involve data protection and, where relevant, the works council now, with a one-page description of data, purpose and human oversight.
Phase 2: Build on real data
Two to three weeks. Connect the real systems from the first day, even if only read-only. Assemble the evaluation set from real historical cases, including the ugly ones. Build the agent against that set and let the evaluation, not opinion, decide between versions.
Phase 3: Shadow and supervised operation
Two to three weeks. The agent runs on live work, but a human reviews every output before it takes effect. This produces the accuracy numbers, the exception rate and the cost per case you need for the go-live decision, and it trains the team on how the agent behaves.
Phase 4: Go-live and handover
Switch the agent to autonomous operation for the low-risk case types first, keep approval gates for the rest, and hand the system to an operations owner with monitoring, a runbook and a change process. If you do not have that owner internally, managed AI operations provides one. Our approach page describes how these phases fit into a broader programme.
How to talk to the board about it
Boards do not reject agents. They reject uncertainty. A pilot report that earns a go-live decision fits on one page and answers five questions:
- What did the task cost before, and what does it cost now per case, including model and platform costs?
- How accurate is the agent on real work, measured against the evaluation set, and what happens to the cases it cannot handle?
- What can go wrong, how likely is it, and which control catches it? A short risk register beats a reassurance.
- Who owns the outcome, and who runs the system after go-live?
- Which decision are you asking for: extend the pilot, go live for specific case types, or stop?
Take the stop option seriously. A board that sees you would kill a weak pilot trusts your recommendation to scale a strong one. And be honest about the number that matters most: not what the demo could do, but what the agent did with last month's real cases.
Frequently asked questions
How long should an AI pilot take?
Long enough to run on real, unfiltered work for several weeks under supervision, and short enough that the sponsor is still in the room. For most agent use cases that is six to eight weeks from framing to a production-grade state. A pilot that has not touched real data after a month is a prototype, not a pilot.
Should we start with a pilot or with a strategy?
If you already have a candidate use case with a business owner and a measurable outcome, a pilot teaches you more than any document. If you have twenty ideas and no owner, a short readiness assessment or strategy sprint saves you from piloting the wrong thing.
What if the pilot shows the agent is not accurate enough?
Then the pilot was a success: it answered the question early and cheaply. Often the fix is a narrower scope, better data or a human checkpoint for one case type. Sometimes the honest answer is to stop. The evaluation set tells you which, which is why it is worth building in week two.
Who should own the agent after go-live?
Two people. The business owner is responsible for the outcome and the process; the operations owner keeps the system healthy, monitors quality and manages changes. Smaller companies combine both roles in one department head, with external support for the technical side.
Related reading and services
Custom AI agent development
Single- and multi-agent systems designed and built for production: tool use, memory, MCP servers, evaluation suites, guardrails and EU deployment.
Managed AI operations (AgentOps)
Monitoring, evaluation, model upgrades, cost control and incident handling for agents and LLM applications after go-live, with a monthly review and documentation your auditors can read.
How we work
Five phases from assessment to operations, six principles we do not bend, and the engagement formats we offer, with typical durations.
Choosing your first AI agent use case
A scoring framework, the traits of good and bad first use cases, eight concrete candidates and a pilot design that proves something.
Let's find the first workflow worth automating.
A 30-minute intro call, no slides and no obligation. We listen, ask about your processes, and tell you honestly where AI agents would pay off and where they would not.