Reference

The vocabulary of AI agents, without the hype.

Every vendor conversation uses these words. Here is what they mean in practice, what they imply for your project and where the catch is. Grouped by topic, written for people who make decisions rather than build models.

Foundations

Large language model (LLM)
A neural network trained on very large amounts of text to predict and generate language. In business use it reads, summarises, classifies, extracts, drafts and, with tools, acts. Models differ in capability, cost, speed, context size and where they can be hosted; the model is rarely the deciding factor in a project's success.
Token
The unit in which language models read and write text, roughly three-quarters of a word in English and somewhat less in German. Providers charge per token, and models have limits on how many tokens they can consider at once. Token counts drive both cost and what fits into a request.
Context window
The maximum amount of text, measured in tokens, a model can take into account in one request: instructions, documents, conversation history and its own output. Larger windows allow more material per request but cost more and do not replace good retrieval.
Prompt
The instructions and material given to a model for a task. In production systems prompts are versioned, tested and treated like code, because small wording changes can change behaviour.
Prompt engineering
Designing instructions, examples and structure so a model reliably produces the wanted output. Useful, but a smaller part of a production system than evaluation, integration and guardrails.
Hallucination
A confident, fluent output that is wrong or invented: a citation that does not exist, a number that was never in the source. Mitigated by grounding answers in retrieved documents, requiring citations, validating outputs and keeping humans in the loop for consequential decisions.
Fine-tuning
Further training of a model on your own examples to change its style, format or domain behaviour. Needed less often than assumed; retrieval and good prompting solve most business cases. Worth it for narrow, high-volume tasks with stable requirements.
Small language model
A compact model that runs cheaply, quickly and often on your own hardware. Good enough for many classification, extraction and routing tasks, and a common choice when data must not leave your environment. See sovereign AI.
Open-weight model
A model whose trained parameters are published so anyone can run it on their own infrastructure, such as the Llama, Mistral, Qwen or Gemma families. Gives full control over data and versions at the cost of operating the model yourself.
Multimodal model
A model that processes more than text, typically images and increasingly audio and video. In business terms: it can read a scanned invoice, a photo of a damaged part or a screenshot, which opens many document and service workflows.

Agents

AI agent
Software that uses a language model to pursue a goal over several steps: reading context, deciding what to do next, calling tools such as your CRM or email, and handing over to a person when a decision needs one. Unlike a chatbot it completes tasks. See What is agentic AI?
Agentic AI
The general term for systems built from AI agents, with varying degrees of autonomy. Used loosely by vendors; the practical questions are what the agent may do without approval, what it may access and how its work is evaluated.
Agentic workflow
A business process automated end to end by one or more agents orchestrated across systems, with defined human checkpoints. Example: an invoice arrives by email, is extracted, matched to a purchase order, coded and routed for approval. See agentic workflow automation.
Tool use / function calling
The mechanism by which a model calls software functions, such as "look up order 4711" or "create a ticket", and receives structured results. This is what turns a text generator into something that can act inside your systems.
Orchestration
The logic that runs an agent's loop: which step comes next, when to call a tool, when to retry, when to ask a person, when to stop. Provided by frameworks and platforms; the quality of orchestration determines reliability more than the model does.
Multi-agent system
Several specialised agents cooperating on a task, for example one that researches, one that drafts and one that checks. Useful for complex work; adds coordination cost and failure modes, so a single well-scoped agent is usually the right first step.
Memory
What an agent retains across steps or sessions: the current task context, facts about a customer, decisions made earlier. Short-term memory is the context window; long-term memory is a database the agent reads and writes. Both need governance about what may be stored.
Planning
An agent's ability to break a goal into steps and adjust when a step fails. Modern models plan reasonably well; production systems constrain plans to allowed actions and require approval for consequential ones.
Human-in-the-loop
A design in which a person reviews, approves or corrects an agent's work at defined points, such as before a payment or a customer message is released. The main lever for making agents approvable and for catching errors before they cost money.
Computer-use agent
An agent that operates software through its screen, keyboard and mouse like a person would, useful where no API exists. Powerful and slow, and it needs strict sandboxing; APIs and MCP servers are preferred where available.
Voice agent
An agent that speaks and listens on phone or voice channels, handling calls such as order status or appointment booking. Emerging in customer service; requires careful design for transparency, escalation and latency.
Copilot vs. agent
A copilot assists a person inside a tool, suggesting and drafting while the person does the task. An agent does the task itself within limits and reports back. Most companies need both: copilots for knowledge work, agents for process work.

Integration

Retrieval-augmented generation (RAG)
A pattern in which the system first retrieves relevant documents from your own sources and then lets the model answer using them, with citations. The standard way to build knowledge assistants that know your company rather than the internet. See enterprise knowledge assistants.
Vector database
A database that stores text as numerical representations and finds passages by meaning rather than exact words. The retrieval component of most RAG systems; often combined with classic keyword search for precision.
Embeddings
Numerical representations of text (or images) that place similar meanings close together. Produced by a model and stored in a vector database so that "invoice dispute" finds "contested bill".
Knowledge graph
A structured representation of entities and their relationships: customers, contracts, products, people. Used alongside RAG when questions require following relationships rather than finding passages.
Model Context Protocol (MCP)
An open standard that defines how AI applications connect to tools and data sources through reusable servers. An MCP server for your ERP lets any compatible agent use it under defined permissions. See MCP explained.
Agent2Agent protocol (A2A)
A protocol for agents to discover and communicate with each other across vendors and platforms. Addresses agent-to-agent cooperation, whereas MCP addresses agent-to-tool connections. Early in adoption; relevant when agents from different systems must collaborate.
API
An application programming interface: the defined way software systems exchange data and trigger actions. Agents act through APIs, which is why API access to your core systems is a prerequisite for most automation.
Low-code automation
Platforms such as n8n, Make, Zapier or Microsoft Power Automate that let teams build workflows visually. Combined with language models they orchestrate many agentic workflows; the trade-off is maintainability at scale. See build vs. buy.
Intelligent document processing (IDP)
Extracting structured data from documents such as invoices, delivery notes and contracts, using OCR and language models. The entry point of many agentic workflows in finance, operations and procurement.

Quality and operations

Evals (evaluations)
Test sets and metrics that measure whether an agent does its task correctly: accuracy of extractions, quality of drafts, correct routing decisions. Run before go-live and after every change to prompts, tools or models. The single most underused practice in AI projects.
Guardrails
The controls that keep an agent within limits: permissions, approval gates, input and output validation, budgets, rate limits, kill switches. See guardrails for AI agents.
Observability / tracing
Recording every step an agent takes, including inputs, tool calls, outputs, latency and cost, so behaviour can be debugged, audited and improved. Tools such as Langfuse or LangSmith provide this.
LLMOps / AgentOps
The operational discipline of running language-model applications and agents in production: monitoring, evaluation, model updates, cost control, incident handling. See managed AI operations.
Prompt injection
An attack in which malicious instructions hidden in content the agent reads, such as an email or a web page, try to make it act against its instructions. Mitigated by treating all external content as untrusted, limiting permissions and requiring approval for consequential actions.
Drift
A gradual change in an agent's behaviour or quality over time, caused by model updates from the provider, changing inputs or accumulated prompt changes. Detected by running evaluations continuously.
Latency
The time between a request and the model's response. Matters for customer-facing and interactive agents, less for batch processing. Influences model choice and architecture.
Token costs
The running cost of language-model usage, billed per token processed. Driven by volume, context size and model choice. Budgets and monitoring per use case prevent surprises; the cost is usually small next to integration and operations effort.

Governance

EU AI Act
The European Union's regulation on artificial intelligence, in force since August 2024 with obligations applying in phases through 2026 and 2027. It classifies AI systems by risk and imposes duties on providers and deployers. See our EU AI Act guide. Not legal advice.
Risk classes
The EU AI Act's tiers: prohibited practices, high-risk systems (for example in employment, credit and critical infrastructure) with extensive obligations, limited-risk systems with transparency duties (such as chatbots), and minimal-risk systems. Most business agents fall into the last two, but the classification must be done and documented.
General-purpose AI (GPAI)
Models trained for broad use, such as the large language models from major providers, which carry their own obligations under the EU AI Act for the model provider. Companies using them are deployers and inherit transparency and literacy duties.
AI literacy
The EU AI Act's requirement, applicable since February 2025, that organisations ensure staff who use or oversee AI systems have sufficient understanding of them. Typically met through role-specific training. See AI training and workshops.
GDPR / DSGVO
The EU's data-protection regulation, applicable to any personal data processed by AI systems: legal basis, purpose limitation, data minimisation, processor agreements, transfer rules and data-subject rights. Applies regardless of the AI Act.
Data protection impact assessment (DPIA)
A structured assessment required under the GDPR for processing likely to pose high risks to individuals, which many AI deployments involving personal data trigger. Prepared with the data protection officer before go-live.
AI register
An inventory of the AI systems an organisation uses or provides, with purpose, data, risk class, owner and controls. The foundation of AI governance and the first thing auditors and regulators ask for. See AI governance.
ISO/IEC 42001
The international standard for AI management systems, describing how organisations govern AI responsibly. Useful as a reference framework for policies and processes; certification is optional and done by accredited bodies.
Data residency
Where data is stored and processed geographically and under which jurisdiction. A central question when choosing model providers and hosting, driven by the GDPR, sector rules and client contracts.
Sovereign AI
AI capabilities a company or country can run and control without dependence on foreign providers: EU-hosted models, European providers or self-hosted open-weight models. A spectrum of options rather than a single choice. See sovereign AI options.
02

Go deeper

Next step

Let's find the first workflow worth automating.

A 30-minute intro call, no slides and no obligation. We listen, ask about your processes, and tell you honestly where AI agents would pay off and where they would not.