AI adoption

Find the AI worth building — and the AI that is not

There is pressure to do something with AI, and very little agreement on what would count as success. The useful discipline is unglamorous: pick one measurable use case, decide the acceptance threshold before you start, build the smallest real version, and evaluate it honestly.

You might recognise one of these

  • “The board asked what our AI strategy is and we improvised.”
  • “We built a demo. It impressed everyone and then quietly went nowhere.”
  • “Our team is already pasting client data into chatbots and we have no policy.”

What we do about it

Most AI initiatives fail in one of two ways. Some never leave the slide deck, because no one was ever accountable for building the thing. Others ship a demo that impresses in a meeting and then quietly fails in production — because nobody defined, in advance, what accuracy would be good enough.

Both failures share a cause: starting with the technology rather than with a workflow and a measurable threshold.

So the first question we ask is not which model to use. It is whether this problem needs AI at all. A surprising number of "AI opportunities" are process problems, integration problems, or reporting problems wearing a fashionable label — and those have cheaper, more reliable, more explainable solutions. We will tell you when that is the case. It is usually the most valuable thing we say.

When AI is the right tool, we define what success means before building: the evaluation set, the accuracy threshold, the acceptable cost per transaction, and what the system does when it is unsure. Then we build the smallest version that runs on real data in the real workflow, and we measure it against that threshold honestly — including when the answer is that it did not clear the bar.

Alongside that, the unglamorous governance work: who can access what, what data leaves your boundary, how output is reviewed, how generated content is identified, and what your developers are permitted to paste into which tools. Most organisations discover they needed that policy about six months after they needed it.

For technical readers

Most business use cases we see are retrieval and extraction problems rather than generation problems, which is good news: they are far more evaluable. A retrieval system can be measured against a labelled set; a chat interface that "feels helpful" cannot.

We build the evaluation harness before the feature. It is a modest amount of work and it is the difference between knowing your extraction is 94% accurate on the cases that matter and believing it is good because the demo went well. It also makes model changes a measurement rather than an argument.

How we decide whether AI is the right tool

Is AI actually necessary for this workflow?
If no

Deterministic software or a fixed process solves it better, cheaper and more explainably. We say so — it is often the most valuable answer we give.

If yes

Define a measurable use case and the acceptance bar — before building anything.

Does a small prototype clear the accuracy bar, at acceptable cost and risk?
If no

Revise the approach, or stop. A demo that impresses but fails in production helps no one.

If yes

Controlled production rollout, human accountable, with a defined fallback when the model is unavailable or unsure.

Capabilities

What this covers

Document extraction and classification

Turning unstructured paperwork into structured, checkable data with a human in the loop.

Search and retrieval

Letting people find and ask questions of the knowledge the organisation already holds.

Summarization and case preparation

Condensing long material into something a professional can verify quickly.

Internal copilots

Assistance inside the workflow people already use, rather than another tab to remember.

Prioritization and anomaly detection

Surfacing the cases that need a human first, and the ones that look wrong.

Governance and safe adoption

Data handling, access control, evaluation before production, and a defined fallback when the model is unavailable or unsure.

Deliverables

What you end up holding

Tangible artefacts, not a verbal summary. You keep all of it.

  • Business use-case inventory, scored and ranked
  • Data-readiness evaluation
  • Security, privacy and confidentiality assessment
  • A working prototype on real data and the real workflow
  • Evaluation against acceptance criteria agreed in advance
  • Cost model at realistic production volume
  • Build-versus-buy recommendation
  • Internal AI usage and data-handling policy for developers
  • A prioritized next pilot — or a documented reason to stop

How the engagement runs

Typical shape and duration

  1. Inventory — week 1

    Workshop to surface candidates, scored on business value, data readiness, risk and feasibility. We also name the ones that need ordinary automation instead.

  2. Set the bar — week 1

    Measurable acceptance criteria, an evaluation set, privacy constraints, and the fallback behaviour. Agreed before anything is built.

  3. Prototype — week 2

    The smallest version that exercises the real workflow on real data, with human review in the loop.

  4. Evaluate — week 3

    Accuracy, cost, latency, failure modes and usability against the threshold. Then: proceed, revise, or stop.

Where it usually starts

AI Adoption Sprint

Identify, prototype and evaluate one high-value AI use case while establishing responsible adoption guidelines — including an honest recommendation when AI is not the right answer.

Duration
3 weeks
Format
Remote, workshop plus working prototype

See what it includes

Questions we are usually asked

What if AI turns out to be the wrong tool?

Then that is the deliverable, with the reasoning and the cheaper alternative. Roughly a third of the workflows people bring us are better solved by deterministic software or by fixing the process. Saying so in week three is worth considerably more than the sprint costs.

Is our data safe?

Data handling, retention and vendor terms are assessed in week one, and we design for data minimization and access control from the start. If a use case cannot be done within your confidentiality obligations, we will tell you before building it, not after.

How accurate does it need to be?

That is a business decision, and it is the first thing we pin down. The right threshold for triage suggestions that a human reviews is very different from the right threshold for something that posts to your ledger unattended.

Do you help teams use AI coding tools, rather than build AI features?

Yes, and for many companies that is the higher-return engagement. It sits under Development Team Services: tool adoption, review practices, test generation with human oversight, and secure data-handling guidelines.

Which model or vendor do you use?

Whichever fits the accuracy, cost, latency and data-residency constraints of your use case. We will show you the comparison rather than assert a preference, and we design so the choice can be changed later.

What about hallucination and liability?

Human accountability is a design constraint, not a disclaimer. That means evaluation before production use, traceability where it matters, clear identification of generated content, and defined fallback behaviour when the system is uncertain.

Tell us what you are trying to build or fix.

A 30-minute call is enough to tell you whether we can help, what it would take, and what the sensible first step is. No obligation, no pitch deck.