AI consulting
Assessment, feasibility, and architecture from a team that will still be here when it is time to implement — including telling you when the answer is not to build anything.
AI consulting services help an organisation decide where AI is worth applying, prove it is feasible against real data, and design the architecture and evaluation approach before committing to a build. TrueLeaf Tech delivers this as an engineering studio rather than a pure advisory firm, so recommendations come from the people who would implement them.
The work
We are engineers who advise, not advisors who subcontract. Every engagement below is designed to leave you with something running.
A structured pass over your workflows to find where AI pays for itself — and, more usefully, where it does not. You get a ranked shortlist with effort and cost estimates against each.
A time-boxed build against your real data to answer one question: can this reach the accuracy the business needs? Cheaper than a programme that discovers the answer in month six.
Retrieval design, model selection and routing, cost modelling, and the abstraction that keeps you portable across providers.
Independent review of an AI system you already run — what it gets wrong, how you would know, and what the fix costs.
Honest comparison against off-the-shelf tools. We have told clients to buy a product rather than hire us, and we would rather do that than build something you regret.
Bringing your engineers up to speed on retrieval, evaluation, and agent design so the capability stays in-house.
Why most AI programmes stall
The most common reason an AI project drifts is that nobody wrote down the accuracy the business actually needs. Without a number, every result is arguable, no release is clearly safe, and the project becomes a matter of opinion. Fixing this costs a week and saves months.
Prototypes built in notebooks, with hand-picked examples and no error handling, do not become production systems — they get rebuilt. If the pilot is meant to survive, it needs real data, real failure cases, and real observability from the start, which is a different budget and a different team.
Teams choose a vector database in week one and spend the next four months discovering that chunking, metadata, and query strategy decide answer quality far more than the store does. Retrieval is where most quality problems live, and most of it is unglamorous data work.
Per-request cost that looks trivial in testing becomes the dominant line item at production volume. Cost modelling belongs in the architecture phase, alongside routing rules that send easy requests to cheaper models.
How we engage
A focused assessment with a written recommendation: what to build, what to buy, what to leave alone, and what it costs. Useful before committing budget.
Our engineers work inside your team for a defined period, building alongside yours and transferring the practice rather than delivering over a wall.
The assessment leads into a build, with the same people. Most of our work looks like this, because the team that scoped it is the team that ships it.
In practice
The value of an assessment is often in what it removes from the roadmap.
A document-heavy operations workflow with a clear accuracy target and a human already reviewing every output. Retrieval was feasible against existing content, the review step meant errors were caught before they mattered, and the volume justified the build. This became a production system.
A client planning a custom transcription and summarisation pipeline. Off-the-shelf tooling already met the accuracy requirement at a fraction of the build cost, and the differentiating work was downstream of transcription anyway. We scoped the integration instead of the build.
An ambitious autonomous agent proposal where the underlying data was inconsistent enough that no retrieval strategy would have reached the required accuracy. The honest sequencing was to fix the data model first. Six weeks of unglamorous work beat six months of an AI project that would have failed on inputs.
Common questions
At minimum: finding where AI is worth applying, proving feasibility against your real data, designing the architecture, and defining how quality will be measured. Ours also includes building the thing — we are an engineering studio, so recommendations come from people who will have to implement them.
Scale and staffing. A large consultancy can field a hundred people and a formal methodology; we field a small senior team that writes the code. If you need organisational change management across a multinational, they are a better fit. If you need a working system and a straight answer, we are.
Rarely as a precondition. Waiting for a tidy data estate is how AI programmes never begin. Most useful first projects need one workflow and the data already attached to it — and building that project usually teaches you more about what your data strategy should be than a strategy exercise would.
Yes, and it is a common request. We review retrieval quality, evaluation coverage, cost per request, guardrails, and operational readiness, then report what is wrong, how you would detect it in production, and what remediation costs.
An advisory sprint delivers a recommendation in one to two weeks. A feasibility prototype against real data typically takes three to four. A production slice runs six to ten weeks. We sequence deliberately so there is something concrete to judge at each step.
No. We hold no reseller relationships and take no vendor commissions, which is precisely why the model recommendation can be made on measured cost and accuracy alone.
Related
The engineering practice our consulting feeds into.
When the recommendation is an agent, this is how we build it.
Why we hand over the test suite as the primary artefact.
The checklist we run during a quality audit.
Let's build
Whether you're testing a hypothesis or scaling an established product, we'd be glad to spend a half-hour helping you think through the next step — no pitch deck required.