Service · incremental, use-case-driven

The data your first agent needs, and nothing more.

The advice to fix your data first has stalled more AI initiatives than bad models ever did. Agents need something narrower: the right documents readable, the right systems reachable, the right definitions agreed and the right people permitted. We build exactly that, one use case at a time, and let the platform grow from there.

01

No big-bang data warehouse before the first agent

Language models read documents, emails and records as they are. That removes the old excuse, and most of the old prerequisites.

For a decade the standard advice was to centralise, clean and model all your data before doing anything intelligent with it. Agents change the economics: they work on unstructured material directly, they call systems where the data already lives, and they need a narrow slice of your data done properly rather than all of it done eventually.

So we turn the order around. The use case comes first, from a readiness assessment or an AI strategy. Then we build what that use case needs: an inventory of what exists, a pipeline that turns documents into usable text, an integration layer that gives the agent controlled access, search where search is needed, and definitions the business agrees on.

Each piece is reusable. The second agent inherits the pipelines, APIs and access rules of the first. After three or four use cases you have a minimal data platform that grew from real demand, with no year of ingestion work in front of the first result.

  • FormatInventory, pipelines, integration, search, governance, per use case
  • DurationTypically delivered alongside the first agent pilot
  • ForCIOs, heads of IT, data leads, COOs
  • OutcomeUsable data for the first agent and a reusable base for the next
See the agent pilots this feeds →
02

What we build

Six building blocks. A first use case rarely needs all of them; a third one usually reuses most.

i.

Data inventory and access map

What data exists, where it lives, who owns it, how current it is, what personal data it contains and how an agent would reach it. A short, honest document that replaces guesswork in every later decision.

ii.

Document pipelines

OCR and intelligent document processing for scans, PDFs, emails and forms: layout-aware parsing, table extraction, classification, metadata and version tracking. The raw material of most agent use cases, made machine-readable.

iii.

Integration layer and APIs for agents

Controlled ways for agents to read from and write to your ERP, CRM, DMS and ticketing systems: APIs, connectors and MCP servers with scoped permissions, so every agent uses the same safe door.

iv.

Vector stores and search

Semantic and keyword search over your documents and records, sized for your volume: from pgvector in your existing database to a dedicated search service. The retrieval backbone of knowledge assistants.

v.

Semantic layer and definitions

What active customer, order value or on-time delivery mean, written down once and used by every agent and report. Small in scope, large in effect: many wrong answers come from undefined terms, not from the model.

vi.

Data quality, access control and protection by design

Quality checks where they change outcomes, not everywhere. Role-based access, pseudonymisation where agents do not need identities, retention rules and logging, designed with your data-protection officer.

03

Mittelstand or enterprise: the right size of foundation

  • A minimal viable data platform for the Mittelstand

    An owner-managed company with an ERP, a file server and a shared mailbox does not need a lakehouse. It needs document pipelines, a few APIs, a search index and clear ownership, often on infrastructure it already pays for. See Mittelstand, Germany's mid-sized businesses.

  • Enterprise: make the existing platform agent-ready

    Corporates usually have a lakehouse or data warehouse already. The work is to make it usable for agents: unstructured data, permission-aware access, MCP servers in front of core systems, and definitions that hold across departments. See enterprise.

  • Startups: build it in from the beginning

    Young companies can design their data model, event logging and document handling with agents in mind, before legacy accumulates. Cheap now, expensive later. See startups.

04

How we work

  1. Start from the use case

    We take the first agent use case and list precisely which data, documents and systems it touches. Everything outside that list waits.

    Week 1
  2. Inventory and gap analysis

    We locate the data, check access, quality and personal-data content, and mark the gaps that would block the agent. Usually there are two or three, not twenty.

    Weeks 1–2
  3. Build the minimum

    Pipelines, APIs, search and definitions for exactly this use case, built to be reused: documented, tested, permission-aware, on infrastructure that fits your data-protection requirements.

    Weeks 2–4
  4. Prove it with the agent

    The first agent runs on the new foundations. Retrieval accuracy, data freshness and access checks are part of its evaluation suite, so data problems surface as measurable failures, not opinions.

    Weeks 4–6
  5. Grow with the roadmap

    Each further use case adds what it needs and reuses the rest. When the pattern is clear, we help you decide whether a lakehouse, a data catalogue or a dedicated data team is now worth it.

    Ongoing
05

Tools and platforms we work with

Chosen per project, vendor-neutral. Naming a product means we work with it or evaluate it, not that we are partnered with or certified by its vendor.

  • Azure AI Document Intelligence
  • Google Document AI
  • AWS Textract
  • Docling
  • PostgreSQL and pgvector
  • Qdrant
  • Weaviate
  • Azure AI Search
  • Elasticsearch / OpenSearch
  • Microsoft Fabric
  • Databricks
  • Snowflake
  • dbt
  • Airbyte
  • n8n
  • Model Context Protocol (MCP)
06

Frequently asked questions

Do we really not need a data warehouse first?

For most first use cases, no. Agents read documents, emails and records where they are and call systems through APIs. A warehouse becomes worthwhile when several use cases need the same consolidated, historical data, or when reporting is the use case. We say so when that point comes; the pipelines built until then carry over.

Our data is messy. Is that a blocker?

Rarely for the whole use case, often for one part of it. The inventory shows exactly which field, document type or system is the problem, and we fix that, not everything. Agents also surface inconsistencies as they run, so quality improves where it matters.

How do you handle personal data and GDPR?

By design: we map which personal data a use case touches, minimise it, pseudonymise where the agent does not need identities, restrict access by role, log what is accessed and agree retention with your data-protection officer. These are engineering principles, not legal advice; our AI governance service covers the wider framework.

Which cloud or hosting do you recommend?

The one that fits your data-protection requirements, your existing licences and your team, from EU regions of Azure, AWS and Google Cloud to EU providers and on-premises. We are vendor-neutral and build so that pipelines can move. See sovereign AI and EU hosting.

Can our IT team maintain this afterwards?

That is the design goal. We use tools your team already knows where possible, document every pipeline and API, and hand over with training. If capacity is the issue, managed AI operations can run the foundations together with the agents.

07

Related services

Next step

Let's find the first workflow worth automating.

A 30-minute intro call, no slides and no obligation. We listen, ask about your processes, and tell you honestly where AI agents would pay off and where they would not.