Private LLM Setup
Run an LLM inside your own walls
Install a language model on your company's own servers, so the team can question internal documents, draft work and connect agents to it. Data never leaves the company, and every answer cites the file and page it came from.
Work in
Private LLM Setup+ human approval
To the customer
A person approves before it ever reaches the customer· drag to rotate
Timeline
Usually 6–10 weeks until the pilot department is using it for real, depending on how ready the hardware is and how many documents there are
Your time
One executive to decide the confidentiality tiers (about 2 hours) · one IT person for machines and user accounts · a pilot team of 10–40 people answering a before-and-after survey
Deliverables
6 deliverables
Starts with
A free 30-minute call
Example: Run an LLM inside your own walls
LIVE DEMO1 · Work arrives
2 · The AI agent does it
- Data map and where the model sits
- An LLM installed and ready to use
- RAG over internal documents, by user permission
- Every answer cites the document page
- Audit log and usage dashboard
3 · You just approve
Every group of data has a defined home in the table you approved, and the log proves that no question ever went somewhere it should not.
Result
-40 to -60%
Time spent finding documents
Who it is for
Is this service right for you?
A fit if you…
- Organisations with piles of internal documents (contracts, policies, manuals, regulations) that staff cannot find, so they ask each other in chat all day
- Businesses holding sensitive data — legal, financial, health, or customer data under PDPA — that are not comfortable sending it to public AI
- Companies whose staff already use public AI quietly, where management wants a safe option it can control
- Organisations planning several agents later, and wanting the central brain to sit in their own house
Not yet, if…
- Small teams whose data is not especially confidential and who only want AI to take over repetitive work — an API with a no-retention agreement plus Agentic Workflow Design is cheaper and faster
- Organisations that still do not know what they want AI to do — start with the AI Readiness Audit, then decide whether you need a model of your own
What it fixes
What this service unlocks
What you get
Real deliverables you can hold
Steps
How it runs step by step
Usually 6–10 weeks until the pilot department is using it for real, depending on how ready the hardware is and how many documents there are in total
Step 1 / 5 · A free 30-minute call + sorting your data by confidentiality
1 week
- 11 week
A free 30-minute call + sorting your data by confidentiality
We ask what you want the team to be able to ask, and how confidential each group of data is, then say whether you really need a model in-house, or only for part of it.
- 21–2 weeks
Choose where the model sits and prepare the machines
- 33–4 weeks
Install, then connect documents and permissions
- 44 weeks
Pilot one department, measure it
- 5Ongoing, monthly
Roll out across the organisation + hand over to IT
Sample of the output
Sample table: data type → where the model sits
An extract from the decision table we build for every organisation before installing. Management sees exactly which group of data goes where, and why (the numbers and data types are adjusted to your business).
Customer contracts, payroll, personal data
Where the model sitsOn-prem, company server
WhySensitive under PDPA; must not leave the network
Operating manuals, internal policies
Where the model sitsPrivate cloud in Thailand
WhyHigh volume, needs strong hardware, but not personal data
Email drafts, news summaries, public documents
Where the model sitsAPI with a no-retention agreement
WhyBetter language quality, pay for what you use
Financial data before results are announced
Where the model sitsOn-prem only + restricted users
WhyAffects the share price and shareholders; every question needs an audit log
| Data type | Where the model sits | Why |
|---|---|---|
| Customer contracts, payroll, personal data | On-prem, company server | Sensitive under PDPA; must not leave the network |
| Operating manuals, internal policies | Private cloud in Thailand | High volume, needs strong hardware, but not personal data |
| Email drafts, news summaries, public documents | API with a no-retention agreement | Better language quality, pay for what you use |
| Financial data before results are announced | On-prem only + restricted users | Affects the share price and shareholders; every question needs an audit log |
Results to expect
Measured in numbers — not feelings
-40 to -60%
Time spent finding documents
Measured on the pilot team, before and after, over 4 weeks (from cases we have run)
0 items
Data leaving the company
The model and the documents sit on servers you control, and the log proves it
-30 to -45%
Repeat questions to experts
Legal and HR answer only the hard ones, within 2 months of the pilot
Budget
How we charge — straight up
Priced as a one-off installation project plus a monthly support fee. Hardware or cloud costs are yours directly — we break them down before you decide, with nothing added on top.
SME on a no-retention API + RAG
The installation project costs about 2–3 months of one employee's salary; monthly support is low and there is no hardware to buy
Mid-sized company on a private cloud in Thailand
The project costs about 3–4 months of one senior employee; the monthly cloud bill is close to one employee's salary
Large organisation, full on-prem
A project of several months with an upfront hardware cost (about the price of a car), but the cost per question is near zero once many people use it
Included in the price
- A data classification and model placement table, approved by management
- Model installation, chat screen, SSO and RAG that follows user permissions
- A test set of your company's real questions, with an accuracy report before go-live
- Audit log, usage and cost dashboard
- A one-department pilot with before-and-after numbers
- Admin manual and training for your IT team
Guarantees
Safe — you can always back out
FAQ
Before you decide — ask anything
What is the difference between on-prem, a private cloud in Thailand and an API — which should we pick?
On-prem means the model runs on your company's own machines: the most control, but you buy the hardware. A private cloud in Thailand means renting machines hosted in the country and used only by you: cheaper and easier to grow. An API with a no-retention agreement has the best language quality and you pay for what you use, which suits data that is not confidential. Most organisations mix all three by confidentiality tier — you do not have to pick just one.
How powerful does the hardware need to be?
It depends on how many people use it at once and how big the model is. A small pilot team can generally start on a machine with one server-class GPU; you add more only once there are several hundred users. We size it for you before you buy, and if you are unsure about the volume we suggest starting on private cloud.
Are open-source models smart enough next to the famous AI?
For question-and-answer over documents, summaries and standard drafting, current open-source models are good enough, because the answer comes from your documents rather than from the model's own knowledge. For polished language or complex reasoning we route only non-confidential data out to a no-retention API.
How do we know the AI is not making things up?
Every answer links to the source page and you can open it straight away, and before go-live we test with a set of real questions from your team and measure how many it gets right. Questions it cannot answer come back as not found and go to the person who looks after them — it does not invent an answer.
Will staff see documents they should not?
No. People log in with their company account (SSO) and the system retrieves only the documents that person already has permission for in the existing system. Management or HR documents are never used to answer ordinary staff, and every question is logged for later review.
Who runs the system after handover, and will the monthly cost blow out?
We hand over with a manual and training so your IT team can add documents, add users and switch model versions themselves. The dashboard shows cost and usage day by day. If you want us to keep running it, there is a monthly package covering model updates and system health checks.
Related cases
See the real thing we already built
An in-house LLM that answers from your documents, with nothing leaving the company
-60% Time spent finding documents
See this case Multi-branch retail · Mid-sized · 200 peopleA leadership report at 07:00 every morning, with nobody writing it
0 hrs Time spent on the report
See this case Accounting practice · SME · 15 peopleRead client receipts, post them, reconcile, close the month sooner
2 days Monthly close
See this caseServices that usually come next
Want the team using AI without data leaving the company?
A free 30-minute call. Tell us how confidential your data is and we will say straight away whether you really need a model in-house, or which cheaper starting point makes more sense. No commitment.