Guide

On-premise AI and self-hosted LLMs

On-premise AI means the model runs where your data already sits, so the data does not leave your infrastructure. The usual form is an open-weights model hosted on your own hardware and owned outright. There are alternatives: a ringfenced environment reached over SSH or VPN, commercial APIs under zero-retention terms, or anonymisation that preserves the statistical properties of the data while removing identifiers. The choice belongs to your security and governance function.

What does on-premise AI mean in practice?

It means the weights, the inference and the data all sit inside a boundary you control. No prompt, no document and no embedding crosses to a third party. In exchange you take on the hardware, the operations and the upgrade path.

Self-hosted LLM and on-premise LLM describe the same arrangement. The distinction that matters is not the phrase but where the boundary is drawn, and who is accountable for what crosses it.

What architectures are available?

  • an open-weights model hosted on your own infrastructure, which you own outright;
  • work conducted inside a ringfenced environment over SSH or VPN, with the data remaining where it is;
  • commercial APIs under zero-retention, zero-training terms; or
  • anonymisation that preserves the statistical properties of the data while removing identifiers, including a bridge approach in which no client data reaches a model provider at all.

These are not a ladder from worse to better. They are four positions on a trade-off, and different processes within the same organisation often sit in different places.

When is a self-hosted open-weights model the right choice?

When the data cannot leave, when the workload is steady enough that owned capacity costs less than metered inference, or when you need the model itself to be an asset rather than a subscription. Owning the weights also fixes behaviour: the model does not change underneath a process that has been validated against it.

It is the wrong choice when the workload is small and irregular, when the capability required is only available from a frontier commercial model, or when nobody in the organisation will own the operational burden.

What is the bridge approach to anonymisation?

Anonymisation removes identifiers while preserving the statistical properties that make the data useful. A bridge approach goes further. The mapping between real entities and their anonymised substitutes is held inside your boundary, and only the substituted data is ever processed, so no client data reaches a model provider at all.

It suits work where the analysis depends on structure rather than on identity. It does not suit work where the identity is the subject.

What are the trade-offs?

  • CostOwned capacity is a fixed cost and metered inference is a variable one. Which is cheaper depends on volume, and the crossing point is calculated rather than assumed.
  • CapabilityA frontier commercial model may do things an open-weights model cannot. Whether that matters depends on the quality threshold the use case has to meet.
  • Operational burdenSomeone has to run the hardware, patch it and own the upgrade path. That role is named before the architecture is chosen, not after.
  • Governance exposureEvery boundary the data crosses is an approval, a contract and a control. Fewer crossings is a simpler position to defend.
  • StabilityOwned weights do not change underneath a validated process. A hosted model can be deprecated on a vendor's schedule.

Most organisations end with a mixture. The architecture is chosen per use case rather than as a policy for the whole estate, and the reasoning for each is recorded so that it can be revisited when the volume or the capability changes.

Who decides which architecture is used?

You do. We set out the trade-offs in cost, capability, latency, operational burden and governance exposure, and build within whichever architecture your security and governance requirements permit.

The decision is a governance decision rather than a technical preference, and it should be recorded as one.

Questions

Common questions

  • What does on-premise mean?

    That the system runs on infrastructure the organisation controls, whether that is hardware in its own building or a private tenancy it administers. The word describes the boundary and who is accountable for it, rather than the physical location of a rack.

  • What is on-premise versus hosted?

    On-premise means you run the model and the data stays inside your boundary. Hosted means a provider runs it and the data crosses to them under whatever contractual terms apply. The technical difference is where inference happens. The governance difference is how many boundaries the data crosses, and each crossing is an approval, a contract and a control.

  • Can you self-host an LLM?

    Yes. Open-weights models can be downloaded, run on your own hardware and kept indefinitely. What is required is capacity sized to the concurrency, someone to own the operations and the upgrade path, and an evaluation written against the quality threshold the use case has to meet.

  • Is it worth self-hosting an LLM?

    It depends on volume, on the capability required and on whether the data can leave at all. Steady, sizeable, confidential workloads favour it. Small, irregular workloads that need frontier capability do not. The comparison is made per use case rather than as a policy for the whole estate.

  • What does self-hosting an LLM cost?

    It is a fixed cost set against a variable one, and we do not quote a figure in advance, because the answer turns on the model, the concurrency and the volume. What is calculated during Discovery is the crossing point: the volume above which owned capacity costs less than metered inference, together with the operational cost of running it.

  • Is an on-premise model less capable than a commercial API?

    Often, though the gap is task-dependent and narrower than headline benchmarks suggest for well-specified institutional work. The right comparison is against the quality threshold the use case has to meet, not against a general leaderboard.

  • Can we start on a commercial API and move on-premise later?

    Yes, provided the use case is specified independently of the provider and the evaluation is written down. Portability is a design decision taken early, not a migration undertaken late.

  • Does zero-retention mean the same as private?

    No. Zero-retention is a contractual commitment that inputs are not stored or used for training. The data still crosses the boundary. Whether that is acceptable is a governance question, not a technical one.

  • Who owns the model at the end?

    With an open-weights deployment, you do, outright, along with the specification and the reasoning behind it.

Begin

Begin with a Discovery

We begin with a paid Discovery: the application of the firm's method to your processes, a data-driven diagnostic across the functions in scope, typically three to six months, and closer to three where strong operational data already exists.

Tell us what you are trying to achieve commercially, and which functions are involved. We will return an initial assessment.

We reply to every enquiry.

Client work is confidential by default.