Private LLM Setup

Run an LLM inside your own walls

Install a language model on your company's own servers, so the team can question internal documents, draft work and connect agents to it. Data never leaves the company, and every answer cites the file and page it came from.

Timeline

Usually 6–10 weeks until the pilot department is using it for real, depending on how ready the hardware is and how many documents there are

Your time

One executive to decide the confidentiality tiers (about 2 hours) · one IT person for machines and user accounts · a pilot team of 10–40 people answering a before-and-after survey

Deliverables

6 deliverables

Starts with

A free 30-minute call

Work arrivesAfraid of a leak, so nothing startsYour problem
The AI agent does itPrivate LLM SetupData map and where the model sits
A person decidesYou review and approveData stays where you decided
Result-40 to -60%Time spent finding documents

Example: Run an LLM inside your own walls

LIVE DEMO

1 · Work arrives

Contracts, payroll, customer data — send them to public AI and nobody knows where they end up. So management bans it, and staff use it anyway.

2 · The AI agent does it

  • Data map and where the model sits
  • An LLM installed and ready to use
  • RAG over internal documents, by user permission
  • Every answer cites the document page
  • Audit log and usage dashboard

3 · You just approve

Every group of data has a defined home in the table you approved, and the log proves that no question ever went somewhere it should not.

Approve in one click

Result

-40 to -60%

Time spent finding documents

Who it is for

Is this service right for you?

A fit if you…

  • Organisations with piles of internal documents (contracts, policies, manuals, regulations) that staff cannot find, so they ask each other in chat all day
  • Businesses holding sensitive data — legal, financial, health, or customer data under PDPA — that are not comfortable sending it to public AI
  • Companies whose staff already use public AI quietly, where management wants a safe option it can control
  • Organisations planning several agents later, and wanting the central brain to sit in their own house

Not yet, if…

  • Small teams whose data is not especially confidential and who only want AI to take over repetitive work — an API with a no-retention agreement plus Agentic Workflow Design is cheaper and faster
  • Organisations that still do not know what they want AI to do — start with the AI Readiness Audit, then decide whether you need a model of your own

What it fixes

What this service unlocks

Contracts, payroll, customer data — send them to public AI and nobody knows where they end up. So management bans it, and staff use it anyway.
The shared drive holds everything, but nobody knows which version is current. So people ask the one person who knows, who answers the same question every day.
The answer looks good but names no source. Staff use it, it turns out wrong, trust is gone, and the whole team stops using it.
Ordinary tools cannot split permissions. Either general staff can surface management documents, or it is locked down so hard nobody can use it.

What you get

Real deliverables you can hold

A table of which kind of your data belongs with which kind of model (on-prem / private cloud / API), with the reasons, ready for management to sign off.
An open-source model on your own servers or on a private cloud in Thailand, with a chat screen staff can use on day one, logging in with their company account (SSO).
Connected to your existing drive or document system, so each person can only ask about the files they are allowed to see. Update a document and the answers update with it.
Every answer says which file and which page it came from, and you can open the original. If it is not in the documents, the system says it found nothing instead of guessing.
A record of who asked what, when, and which documents the system pulled to answer. Reviewable afterwards for PDPA requirements and internal audit.
Your IT team knows how to add documents, add users, watch the cost and move to a newer model on their own, without rebuilding the system.

Steps

How it runs step by step

Usually 6–10 weeks until the pilot department is using it for real, depending on how ready the hardware is and how many documents there are in total

Step 1 / 5 · A free 30-minute call + sorting your data by confidentiality

1 week

  1. 11 week

    A free 30-minute call + sorting your data by confidentiality

    We ask what you want the team to be able to ask, and how confidential each group of data is, then say whether you really need a model in-house, or only for part of it.

  2. 21–2 weeks

    Choose where the model sits and prepare the machines

  3. 33–4 weeks

    Install, then connect documents and permissions

  4. 44 weeks

    Pilot one department, measure it

  5. 5Ongoing, monthly

    Roll out across the organisation + hand over to IT

Sample of the output

Sample table: data type → where the model sits

An extract from the decision table we build for every organisation before installing. Management sees exactly which group of data goes where, and why (the numbers and data types are adjusted to your business).

Customer contracts, payroll, personal data

Where the model sitsOn-prem, company server

WhySensitive under PDPA; must not leave the network

Operating manuals, internal policies

Where the model sitsPrivate cloud in Thailand

WhyHigh volume, needs strong hardware, but not personal data

Email drafts, news summaries, public documents

Where the model sitsAPI with a no-retention agreement

WhyBetter language quality, pay for what you use

Financial data before results are announced

Where the model sitsOn-prem only + restricted users

WhyAffects the share price and shareholders; every question needs an audit log

Results to expect

Measured in numbers — not feelings

-40 to -60%

Time spent finding documents

Measured on the pilot team, before and after, over 4 weeks (from cases we have run)

0 items

Data leaving the company

The model and the documents sit on servers you control, and the log proves it

-30 to -45%

Repeat questions to experts

Legal and HR answer only the hard ones, within 2 months of the pilot

Budget

How we charge — straight up

Priced as a one-off installation project plus a monthly support fee. Hardware or cloud costs are yours directly — we break them down before you decide, with nothing added on top.

SME on a no-retention API + RAG

The installation project costs about 2–3 months of one employee's salary; monthly support is low and there is no hardware to buy

Mid-sized company on a private cloud in Thailand

The project costs about 3–4 months of one senior employee; the monthly cloud bill is close to one employee's salary

Large organisation, full on-prem

A project of several months with an upfront hardware cost (about the price of a car), but the cost per question is near zero once many people use it

Included in the price

  • A data classification and model placement table, approved by management
  • Model installation, chat screen, SSO and RAG that follows user permissions
  • A test set of your company's real questions, with an accuracy report before go-live
  • Audit log, usage and cost dashboard
  • A one-department pilot with before-and-after numbers
  • Admin manual and training for your IT team
Get a quote / book a free call

Guarantees

Safe — you can always back out

Every group of data has a defined home in the table you approved, and the log proves that no question ever went somewhere it should not.
We set the model to answer only from the documents it retrieved, with page citations. With no evidence it says so plainly and passes the question to the person responsible.
Answers on legal or financial matters, or anything going outside the company, are always drafts. Someone accountable has to approve them before use.

FAQ

Before you decide — ask anything

What is the difference between on-prem, a private cloud in Thailand and an API — which should we pick?

On-prem means the model runs on your company's own machines: the most control, but you buy the hardware. A private cloud in Thailand means renting machines hosted in the country and used only by you: cheaper and easier to grow. An API with a no-retention agreement has the best language quality and you pay for what you use, which suits data that is not confidential. Most organisations mix all three by confidentiality tier — you do not have to pick just one.

How powerful does the hardware need to be?

It depends on how many people use it at once and how big the model is. A small pilot team can generally start on a machine with one server-class GPU; you add more only once there are several hundred users. We size it for you before you buy, and if you are unsure about the volume we suggest starting on private cloud.

Are open-source models smart enough next to the famous AI?

For question-and-answer over documents, summaries and standard drafting, current open-source models are good enough, because the answer comes from your documents rather than from the model's own knowledge. For polished language or complex reasoning we route only non-confidential data out to a no-retention API.

How do we know the AI is not making things up?

Every answer links to the source page and you can open it straight away, and before go-live we test with a set of real questions from your team and measure how many it gets right. Questions it cannot answer come back as not found and go to the person who looks after them — it does not invent an answer.

Will staff see documents they should not?

No. People log in with their company account (SSO) and the system retrieves only the documents that person already has permission for in the existing system. Management or HR documents are never used to answer ordinary staff, and every question is logged for later review.

Who runs the system after handover, and will the monthly cost blow out?

We hand over with a manual and training so your IT team can add documents, add users and switch model versions themselves. The dashboard shows cost and usage day by day. If you want us to keep running it, there is a monthly package covering model updates and system health checks.

Want the team using AI without data leaving the company?

A free 30-minute call. Tell us how confidential your data is and we will say straight away whether you really need a model in-house, or which cheaper starting point makes more sense. No commitment.

Free consultation

Book a free call 30 minutes