Zimmer
On-premise AI · Business infrastructure

On-premise AI for business

Run a shared AI service on hardware your company owns, keep prompts and documents inside your network, and give employees cited answers without building an inference, identity, retrieval, and remote-access stack from parts.

Published September 4, 2026 · Updated September 4, 2026 · By Omer Khan, Zimmer (Fihi Labs UG)

On-Premise AI for Business

On-premise AI is an operating model, not a privacy label

On-premise AI places model execution and the systems around it inside infrastructure the organisation controls. That boundary must cover more than the model file: prompts, document ingestion, embeddings, retrieval, generated answers, identity, logs, and administration all matter. Cohere's private deployment documentation draws the same distinction between on-premises hardware and a private cloud environment.

Zimmer AI provides a local-first workspace and an on-premise appliance around open-weight models. Zimmer Desktop is a complete personal workspace for one Mac or Windows PC. Zimmer Server is the separate multi-user appliance for private team networks, hosted on a company-owned Apple-silicon Mac, so organisational prompts, source material, and answers remain on the machine or network the company controls.

Start with the workload boundary, not the hardware catalogue

A defensible project begins with one recurring workload and its data path. Good first candidates include asking questions across controlled policy documents, comparing clauses across proposals, drafting from an approved knowledge collection, or finding evidence inside technical manuals. Define who may see each collection and what a useful citation must open before choosing a model.

Do not begin by promising that every cloud workflow will move on day one. A hosted frontier model such as Claude, GPT, or Gemini outperforms any model that fits on a small local server. The purpose of an on-premise deployment is to move work whose privacy, governance, availability, or cost pattern justifies that quality trade—not to pretend the trade does not exist.

The network boundary Zimmer Server actually provides

Zimmer Server runs inference, document indexing, vector retrieval, and agent tool calls on the organisation's server. There is no Zimmer content backend that holds customer prompts or files. Pilot registration sends the organisation name, administrator name, and work email; paid licensing carries billing identity. Those limited account events do not transfer document content.

Normal Team and pilot licences verify roughly every seven days and retain a 14-day grace period before a gentle read-only mode. A fully air-gapped Enterprise deployment uses a signed offline licence file. This is factual architecture, not a compliance certificate: Zimmer does not claim HIPAA, ISO 27001, or SOC 2 certification, and deployment on owned hardware does not satisfy a customer's regulatory duties by itself.

A cost model a finance lead can challenge

Compare annual cost per active employee and cost per completed workflow, not a cloud list price against a hardware receipt. Lenovo's 2026 TCO analysis shows why utilisation changes the result for large AI systems; its enterprise GPU figures are not a benchmark for Zimmer Server, but the method is useful. Owned capacity rewards steady use. Cloud capacity rewards short, variable demand.

Cost lineZimmer ServerHosted AI
CapacityBuy and operate an Apple-silicon Mac; a Mac Studio is recommended, while a Mac mini supports a small team.Capacity is rented through a seat, token, or infrastructure bill and can expand without buying hardware.
SoftwareTeam is $29 per user per month or $250 per user per year; the three-seat pilot is free.Use the provider's current seat and usage prices, including separate API or storage charges.
Usage growthInference is not metered, so more questions consume owned capacity without adding a token invoice.Additional seats, tokens, storage, or premium model use may increase the bill under the chosen plan.
OperationsInclude power, backups, updates, support time, model evaluation, and the hardware replacement cycle.Include vendor administration, data review, integrations, egress, and the cost of service dependency.
QualityLocal open-weight models trade peak capability for control, ownership, and predictable cost.Hosted frontier models generally provide stronger reasoning and are often cheaper for bursty or low-volume use.

Choose the server from concurrency and model fit

Zimmer Server currently hosts on Apple-silicon Mac hardware. A Mac mini is supported for a small team; a Mac Studio is recommended when several people share the service. Size for the model that answers the chosen workload and the number of simultaneous requests, then leave memory headroom for context and retrieval rather than buying only to a parameter count.

A second Mac can be enrolled as a node with a one-time code in about two minutes. Requests route to a node that is least busy and already has the model loaded; if a node drops, the request is rerouted. Each routing decision is shown with its reason. Linux or GPU nodes and native Windows server hosting remain roadmap items, not shipped options.

Shared documents need permissions inside retrieval

Zimmer Server keeps document collections private until they are shared with a group such as Legal, Finance, HR, or Everyone. The group filter is applied inside vector retrieval. A passage the employee cannot access is removed before model generation, so the model cannot disclose text that never entered its context. Grants and revocations take effect on the next question.

Document answers carry clickable citation chips that open the exact PDF page or the relevant Word or Markdown section. When the indexed documents do not contain an answer, Zimmer says so. That makes review faster, but it does not make a generated answer authoritative; the employee still checks the cited source and applies professional judgement.

Remote access without publishing the server

Zimmer Server opens zero inbound ports. An administrator creates a single-use enrolment payload valid for 60 minutes. The employee pastes it once, the device generates a keypair, and the private key is sealed in the Mac Keychain. Browser users receive a one-time sign-in code instead of a shared password.

The same enrolment joins the device to an end-to-end encrypted WireGuard-class mesh, with no public IP, port forwarding, or inbound firewall change. Revoking a device denies its next request. This is not the same as a physical air gap: an air-gapped deployment also needs an intentional process for moving approved model and software files across the boundary.

Buy an appliance or build the stack?

The build-or-buy decision is mostly about which operational layers the team wants to own. A roll-your-own platform can fit unusual hardware and integrate with an existing Kubernetes or GPU estate. An appliance is a better fit when a small IT team wants a bounded deployment and a single path for model service, identity, permissions, document retrieval, remote access, backup, and licensing.

Build from components

Choose this when the platform team needs custom runtimes, Linux or GPU nodes, and can own identity, network exposure, retrieval security, observability, and upgrades.

Use Zimmer Server

Choose this when Apple-silicon hardware fits and the team wants cited document Q&A, in-retrieval permissions, device enrolment, encrypted access, and backup in one product.

Keep hosted AI

Choose this when usage is light or bursty, the strongest frontier model is essential, or the organisation cannot operate supported local hardware responsibly.

Plan for failure, licensing, backup, and updates

A private server transfers operational responsibility to the organisation. Test what happens when a node disappears, a device is revoked, the public internet is unavailable, and a backup is restored. Zimmer Server backs up configuration, users, devices, permissions, the document index, sync tokens, and the licence; an integrity manifest is checked before restore writes data.

For an air-gapped system, document who approves and transfers new models and application updates. Zimmer supports a signed offline licence, but product facts do not promise an automatic offline update bundle. The customer's security process owns the transfer media, malware inspection, provenance checks, change record, and rollback decision.

A 30-day pilot should try to disprove the fit

Use the free three-seat pilot with one document collection and two permission groups. In week one, record a baseline for the current workflow. In week two, test cited answers and deliberately denied retrieval. In week three, enrol a remote device, revoke it, and disconnect the public internet. In week four, restore a backup and compare the result with the baseline.

Set pass criteria before the first question: the answer must cite a usable source, restricted text must never enter an unauthorised result, revocation must apply on the next request, and the team must accept the local model's quality on the selected workload. If any criterion fails, stop or narrow the deployment rather than broadening it.

Questions an on-premise AI buyer should ask

What does on-premise AI mean for a business?

On-premise AI runs the model and the application layer on infrastructure the organisation controls, rather than sending every prompt to a public model API. With Zimmer Server, inference, document indexing, retrieval, and answers run on a company-owned Apple-silicon Mac and stay on the organisation's private network.

When is on-premise AI cheaper than cloud AI?

On-premise AI becomes easier to justify when usage is steady, several employees share the same capacity, or per-token work is frequent. The fair comparison includes hardware, power, administration, software, and support—not only API charges. Bursty experiments and teams without suitable hardware often remain cheaper in the cloud.

Can employees use an on-premise AI server remotely?

Yes. Zimmer Server enrols devices with a single-use payload valid for 60 minutes and joins them to an end-to-end encrypted WireGuard-class mesh. The server opens zero inbound ports, needs no public IP or port forwarding, and rejects a revoked device on its next request.

What are the current limits of Zimmer Server?

Zimmer Server currently requires Apple-silicon Mac hardware; a Mac Studio is recommended and a Mac mini is supported for small teams. Native Windows server hosting, Linux or GPU nodes, and Enterprise SSO are not shipped. Hosted frontier models also remain stronger than models that fit on a small local server.

Evidence to bring to the security review

Bring a data-flow diagram, the list of outbound account and licence events, a demonstration of denied retrieval, citation behaviour, device enrolment and revocation, backup restoration, and the exact offline test result. The Zimmer security architecture explains the network and licence boundaries; the law-firm scenario shows the retrieval permission test with a concrete confidentiality boundary.

Treat unsupported capabilities as decision inputs. Zimmer Server does not ship Enterprise SSO, Windows server hosting, Linux or GPU nodes, or a compliance certificate. If any of those is mandatory now, document the gap and choose a different deployment rather than turning a roadmap item into an assumption.

Test one real workflow with three seats

Zimmer Server's self-service pilot runs on an Apple-silicon Mac and covers three users. Start with one representative collection, test citations and denied retrieval, and measure whether the privacy, governance, and cost boundary is worth the local-model trade-off.