Private AI for business
Your teams want ChatGPT-level help with contracts, customer records and code. Your security team does not want any of it leaving the building. Private AI gives you both: capable language models running inside infrastructure you control.
Private AI means running AI models — usually open-weight large language models — inside infrastructure your organisation controls: your own servers, a dedicated account in your cloud, or an air-gapped network. Prompts, documents and answers stay inside that boundary instead of going to a third-party AI provider. TrueLeaf Tech designs, builds and hands over these systems, and has shipped on on-prem open-weight model stacks as well as Claude, GPT and Gemini.
Why now
People are already using AI at work. When the approved option is missing, they use whatever is at hand — and your data goes with it.
organisations reported a breach caused by shadow AI — AI tools used without IT oversight.
IBM Cost of a Data Breach 2025higher breach costs for organisations with high levels of shadow AI than for those with little or none.
IBM Cost of a Data Breach 2025of organisations that reported a breach of AI models or apps had no AI access controls in place.
IBM Cost of a Data Breach 2025of organisations say their privacy programmes have expanded because of AI.
Cisco Data & Privacy Benchmark 2026Banning AI tools tends to push usage into personal accounts. A private deployment gives people a sanctioned assistant that is good enough to actually use — with access control and an audit trail built in from the start.
When the model runs inside your network, your data doesn't travel.
That is the whole idea. Everything else — the architecture, the models, the rollout — exists to make that practical.
Public AI vs private AI
Both approaches can use capable models. What changes is who runs the model, where your prompts and files are processed, and who decides what happens to them.
| Public AI API | Private AI | |
|---|---|---|
| Where prompts and files are processed | On the provider's infrastructure | Inside infrastructure you control |
| Which models | The provider's own models | Open-weight models you choose and can swap |
| Data location | The provider's regions; some offer regional residency | Wherever you run it, including a single country |
| How cost behaves | Pay per use; grows with every request | Infrastructure you run; flatter as usage grows |
| Customisation | Prompting, some fine-tuning | Full control, including tuning on your own data |
| Internet connection | Required | Optional — can run fully offline |
| Best for | Low-sensitivity work and a fast start | Confidential or regulated data, and steady high volume |
Enterprise plans have improved: OpenAI, for example, states that it does not train on business data by default and offers data residency for eligible customers. Your prompts and files are still processed on its infrastructure — and for many regulated teams, that is exactly the line.
Deployment options
The right option depends on how sensitive the data is, what you already run, and how much change your team can take on.
Architecture
Five layers, all inside your boundary. The model is one of them — and rarely the part that decides whether the system is trusted.
Swipe to see the whole diagram →
A private chat assistant, plus AI inside the tools and systems your teams already use, signed in the way they sign in today.
Every request passes one gateway that checks identity and permissions, logs what was asked, and applies limits — so any app can use AI without its own security work.
Open-weight models served on your GPUs, often a larger one for hard questions and a smaller, cheaper one for routine work.
Your content indexed for search, respecting who may see what, so answers are grounded in your data and cite where they came from.
Tests on your own data, plus usage and cost monitoring, so you know when quality drifts before your users do.
What you can build
Private AI is not a science project. These are the shapes it usually takes first.
A familiar chat interface over approved models, with sign-in, permissions and an audit trail.
Policies, contracts, manuals and tickets, searchable in plain language, with citations back to the source.
Invoices, claims, forms and reports turned into structured data your systems can use.
Replies drafted from your knowledge base and customer history, reviewed by a person before they go out.
Help across your codebase without sending proprietary source code to an outside service.
Questions about reports and data asked in plain language, answered inside your existing governance.

Clinical notes and records stay inside the hospital network.

Client data, statements and internal models stay under your control.

Privileged documents are searched and summarised without leaving the firm.

Manuals and plant knowledge on hand at the line, even offline.
How we deliver
Fixed-scope phases you can stop between, so every step is decided on evidence rather than momentum.
Assess
A one-to-two-week advisory sprint: the use cases worth doing, the data involved, and which deployment option fits. You get a written recommendation and a cost envelope.
Prove
A prototype measured against the accuracy the work actually needs, so the decision to invest is made on evidence.
Deploy
Model serving, retrieval, gateway and monitoring set up in your infrastructure. A first workflow in production typically takes six to ten weeks.
Hand over
Code, prompts, test sets and the evaluation harness belong to you, with runbooks for your team. Ongoing support only if you want it.
An honest fit check
If an enterprise plan from an AI provider is the better answer for you, we will say so.
Cost
There is no sticker price, because a handful of choices move it more than anything else. Scoping takes about a week and ends with a cost envelope before any build starts.
Bigger models need more GPU memory. Many business tasks run well on smaller models.
Concurrent users, not total users, set how much serving capacity you need.
Renting GPUs in your cloud account, or buying servers you run yourself.
Cleaning, structuring and indexing content is often the largest piece of work.
Monitoring, model updates and evaluation after launch, by your team or ours.
Common questions
Private AI means running AI models — usually open-weight large language models — inside infrastructure your organisation controls: your own servers, a dedicated account in your cloud, or an air-gapped network. Prompts, documents and answers stay inside that boundary instead of being sent to a third-party AI provider.
Yes. A private ChatGPT is a familiar chat interface over open-weight models that you host, with sign-in through your existing identity provider, answers grounded in your own documents, and an audit trail — all running inside your infrastructure.
It depends on your rules. OpenAI states that it does not train on business data by default and offers retention controls and data residency for eligible customers. Your prompts and files are still processed on OpenAI's infrastructure, though. If your policy, regulator or clients require that data never leaves systems you control, you need a private deployment; if not, an enterprise plan may be the simpler choice, and we will say so.
Open-weight model families such as Llama, Mistral, Qwen and Gemma, in sizes from small to very large. We pick by measuring candidates on your own tasks, and keep the setup model-agnostic so you can swap models as better ones are released.
For the hardest open-ended reasoning, the largest hosted models often still lead. For focused work grounded in your documents — search, extraction, drafting, classification — well-chosen open-weight models with good retrieval can meet the bar. The reliable answer is an evaluation on your own data, which is the first thing we build.
No. You can run on GPU instances in your own cloud account, buy servers for an on-premise deployment, or start in the cloud and move on-premise later once usage is predictable.
Scoping takes about a week and ends with a written recommendation and cost envelope. A first workflow in production typically takes six to ten weeks, with larger programmes split into phases you can stop between.
There is no fixed price, because cost depends on model size, how many people use it at once, where it runs, how much document preparation is needed, and who operates it afterwards. We give a cost envelope at the end of scoping, before any build starts.
It can make them easier to meet, because data stays where you decide and every request can be logged. Compliance still depends on your whole setup — access, retention, monitoring and process — so we design it with your security and compliance teams rather than treating deployment as the finish line.
Related
Assessment and feasibility before committing to a build.
The wider AI engineering practice private AI sits inside.
Assistants grounded in your data, private or hosted.
Why search quality decides answer quality.
Private AI
Tell us what you would like AI to help with and what cannot leave your network. We will come back with the options that fit — including if you don't need private AI at all.