Insights · Guide

How to choose your first AI agent use case.

The first agent you put into production decides whether your company believes in the next ten. This guide gives you a scoring framework, the characteristics of good and bad first use cases, eight concrete candidates across departments, and a pilot design that produces evidence rather than a demo.

01

Why it matters

Why the first use case matters more than the tenth

Every company that has tried AI has a story, and it is usually about the first attempt. A visible failure, whether a chatbot that embarrassed the service team or a pilot that never left the sandbox, buys eighteen months of “we tried that, it doesn't work here”. A visible success does the opposite: department heads start asking for their own agent, IT gets budget for the plumbing, and governance discussions become practical instead of theoretical.

That is why the first use case should be neither the most ambitious one nor the one with the largest theoretical saving. It should be the one most likely to reach production with a measurable result inside a quarter. Ambition comes in the second wave, on the credibility the first one built.

A scoring framework for candidate use cases

We score every candidate on seven criteria. The numbers matter less than the conversation they force between the process owner, IT and leadership. Score each criterion from one to five and be honest about the low scores; they are the ones that decide.

Scoring criteria for a first AI agent use case
CriterionThe question to askScores high when
Business valueWhat changes for the company if this works: hours, cycle time, errors, revenue, risk?The value is large enough that leadership notices, and it arrives within months, not years
Frequency and volumeHow often does the process run, and how many cases per week?Daily or continuous, with hundreds of cases a month or more
FeasibilityCan today's models do this reliably, and can we reach the systems involved?The task is reading, classifying, extracting, drafting or matching; the systems have APIs or exports
Data availabilityDo we have the documents, records and examples the agent needs, and may we use them?Historical cases with known outcomes exist and are accessible without a data project
Risk and reversibilityWhat happens when the agent is wrong, and can it be undone?Errors are caught by a review step and cost little; no legally sensitive decisions about people
MeasurabilityCan we state the current baseline and the target in one sentence?A baseline exists or can be measured in two weeks; success is a number
SponsorWho wants this, owns the process and will defend the change?A named department head with budget authority who will spend time on the pilot

Three of these are effectively kill criteria. Without a sponsor, the pilot dies at the first integration hurdle. Without a baseline, you cannot prove it worked. Without reversibility, the first mistake becomes the story.

02

Good and bad candidates

What good first use cases have in common

  • High volume, low glamour. The work is repetitive enough that nobody defends it, and frequent enough that improvements show up in weeks.
  • Clear rules with judgement in the middle. The process is documented, but an experienced person still applies rules of thumb. That is exactly the gap a language model fills and RPA cannot.
  • Unstructured inputs. Emails, PDFs, tickets, free-text forms. If the input were structured, a simpler integration would already exist.
  • A person already checks the output. The review step exists today, so the agent slots into a level of autonomy the organisation already understands.
  • Measurable. Cycle time, backlog, error rate or hours per case can be measured before and after.
  • Bounded. One workflow, one team, two or three systems. Not “customer service”, but “first-response drafts for order-status tickets”.

What to avoid the first time

  • Customer-facing, high stakes. An agent that speaks to customers without review, in a situation where a wrong answer costs a contract, is a second-wave project.
  • Legally sensitive decisions. Anything touching hiring, promotion, credit or access to essential services is high-risk under the EU AI Act and carries oversight and documentation duties you do not want to learn on your first pilot. Our EU AI Act guide lists the categories.
  • No baseline. If nobody can say how long the process takes or how often it fails today, measure that first or pick another process.
  • A process that is being redesigned anyway. Automating a moving target produces an agent for a workflow that will not exist in six months.
  • Value only at full autonomy. If the business case works only when the agent acts without any review, you are betting the first pilot on level four. Choose something that already pays at level two or three.

Eight concrete candidates

These are typical first use cases we see across departments. Each is described in more depth on the linked solution page.

  1. Customer service: ticket triage and first-response drafts. The agent classifies incoming tickets, pulls order and contract data and drafts a reply that a service agent approves. Baseline: first-response time, handling time. See customer service.
  2. Finance: invoice matching and exception explanation. Incoming invoices are matched to purchase orders and goods receipts; discrepancies are explained with evidence and routed. Baseline: days to post, share of manual exceptions. See finance.
  3. Sales: inbound lead qualification and CRM enrichment. Leads are researched, scored against your criteria and handed over with a suggested first reply. Baseline: time to first contact, data completeness. See sales.
  4. Operations: order entry from emails and PDFs. Orders are read, validated against price lists and stock and entered into the ERP, with a person confirming exceptions. Baseline: hours per order, entry errors. See operations.
  5. HR: employee helpdesk on policies and processes. Questions are answered from the handbook with sources, and routine requests such as certificates are prepared. Baseline: HR tickets per month, response time. See HR.
  6. Procurement: supplier document intake and comparison. Quotes and certificates are extracted, normalised and compared against requirements; a buyer decides. Baseline: time per tender evaluation. See procurement.
  7. Legal: first-pass contract review against a playbook. Standard agreements are checked against your positions, deviations flagged with reasons. Baseline: review time per contract, backlog. See legal and compliance.
  8. IT: incident enrichment and runbook execution with approval. Incidents are enriched with logs and history; known fixes are executed after a click. Baseline: time to resolve, tickets per engineer. See IT and engineering.
03

The pilot

Designing the pilot

A first pilot should be scoped to reach a production-grade agent for one workflow in typically six to eight weeks. That means real data, real systems and real users from the start, not a synthetic demo. Four design decisions matter most.

  • Baseline first. Measure the current process for two weeks before the agent touches it: cases, time per case, error rate, backlog. Without this number the pilot cannot succeed, only impress.
  • Two or three success metrics, written down. One for value (hours or cycle time), one for quality (error rate or rework), one for adoption (the share of cases the team lets the agent handle). Define the go/no-go threshold before the pilot starts.
  • Human checkpoints by design. Decide per action what the agent may do alone, what it proposes and what it prepares for approval. Start conservative and loosen based on evaluation data, not enthusiasm.
  • An evaluation set. Fifty to a few hundred historical cases with known correct outcomes. Every prompt or tool change is tested against it, so improvements are proven and regressions are caught.

Our AI agent development engagements follow this structure, and the readiness assessment that often precedes them produces the shortlist and the baselines.

Common mistakes

  • Starting from the technology. “We have Copilot licences, what can we do with them?” is a procurement question, not a use-case selection.
  • No process owner. IT-led pilots without a department that wants the result rarely reach production.
  • Measuring activity instead of outcome. Number of prompts, number of users and satisfaction surveys are not a business case.
  • Too much autonomy too early. The first visible error at level four costs more credibility than a slow start at level two ever will.
  • Skipping evaluation. Prompt changes made on feel produce an agent that is good on Tuesday and bad on Thursday.
  • No plan for operations. An agent needs an owner, monitoring and a change process after the pilot. If nobody has thought about managed AI operations, the pilot is the end of the story.
04

Frequently asked questions

How long does a first agent pilot take?

Typically six to eight weeks from kick-off to a production-grade agent for one workflow, provided the baseline data exists and the systems are reachable. Two weeks of baseline measurement before that are well spent.

Do we need clean data before we start?

You need the documents and records the process already uses, and a set of past cases with known outcomes. You do not need a data warehouse. Where data quality genuinely blocks a candidate, score it low on data availability and pick another one; our data foundations work can run in parallel.

Should the first use case be in IT, where the technical team is?

Only if IT also owns a high-volume process with a business metric. Usually the better first candidate is in service, finance or operations, with IT as the partner that provides access to systems. What matters is a sponsor who owns the outcome.

What if the first pilot fails?

A pilot with a baseline, defined metrics and an evaluation set does not fail silently; it tells you why. Often the fix is scope: the agent handles a narrower set of cases at a lower autonomy level and still delivers. Our article on why AI pilots fail covers the usual causes.

05

Related services and reading

Next step

Let's find the first workflow worth automating.

A 30-minute intro call, no slides and no obligation. We listen, ask about your processes, and tell you honestly where AI agents would pay off and where they would not.