Private AI document search
Search internal company files, contracts, and knowledge repositories on hardware your organisation owns. Zimmer Server enforces mathematical role-based access control inside vector retrieval, provides clickable page-level citations, and ensures confidential documents never leak to external cloud providers.
Published October 8, 2026 · Updated October 8, 2026 · By Omer Khan, Zimmer (Fihi Labs UG)

The multi-department scenario: why standard enterprise RAG leaks confidential files
Consider a 45-person company structured into Executive Leadership, Legal, Human Resources, Finance, and Product Engineering. The organisation maintains thousands of internal files: executive compensation models, pending litigation discovery bundles, proprietary software architecture blueprints, and general employee handbooks. To increase internal efficiency, the operations team introduces a centralized AI assistant capable of answering natural-language queries across the company repository.
In standard enterprise chat deployments or generic cloud-based Retrieval-Augmented Generation (RAG) frameworks, documents from every department are ingested into a shared vector index. When an engineer asks a technical question regarding system performance, a naive semantic similarity search sweeps across the entire vector space. If access restrictions rely merely on system instructions—such as prompting the language model to "only disclose compensation data if the user is in HR"—the security model has already failed.
Prompt-based security is fundamentally unprovable. Adversarial prompting, indirect prompt injection within third-party documents, and creative rephrasing can easily persuade a generative model to summarize or hint at sensitive passages sitting in its working context. Even without malicious intent, an LLM synthesizing broad operational context will routinely echo confidential metrics if those paragraphs were retrieved into its prompt. The breach does not happen during answer rendering; it happens the moment unauthorized text chunks enter the model's memory.
Under the hood, Zimmer AI is engineered as a local-first AI workspace and on-premise AI appliance running open-weight models on hardware your organisation owns—whether on a personal Mac or Windows PC or a dedicated company-owned Apple-silicon server—guaranteeing that prompts, confidential documents, vector indexes, and generated answers never leave your physical premises or private network. For single practitioners, Zimmer Desktop serves as your personal, local AI workspace for one machine. For collaborating organizations, Zimmer Server provides the dedicated multi-user appliance for private team networks.
In-retrieval RBAC: mathematical permission filtering inside vector search
True organizational confidentiality requires access control to operate as an immutable mathematical gate inside the retrieval engine, rather than a conversational filter after the fact. Security architectures in RAG fall into two distinct paradigms: post-retrieval filtering and in-retrieval pre-filtering. Understanding the difference is vital for any security review.
In post-retrieval filtering, the vector search engine fetches the top twenty nearest-neighbor text chunks based solely on semantic distance. An application middleware then inspects the metadata tags of each chunk and discards those belonging to restricted collections. This flawed approach introduces two catastrophic problems: first, if the top twelve most relevant chunks belong to a restricted HR document, the user receives an answer synthesized from remaining, low-quality fragments; second, timing and ranking discrepancies can leak the existence and keywords of restricted files.
Zimmer Server implements true in-retrieval Role-Based Access Control (RBAC). User accounts belong to defined administrative groups—such as Legal, Finance, Executive, HR, and a global Everyone tier. Collections remain strictly private until an administrator explicitly shares them with designated groups. When an employee queries the private search system, the pipeline executes a three-stage authorization workflow:
Client requests arrive cryptographically signed by the employee's hardware keypair sealed in their device Keychain, verifying active group memberships without session cookies or passwords.
Vector candidate generation executes only across embedding nodes tagged with the caller's allowed group IDs. Restricted collections are mathematically excluded before cosine distance is evaluated.
Because unauthorized passages are pruned inside index traversal, exactly zero restricted tokens enter the model context window. The model cannot leak information it physically never received.
Group grants and revocations take effect immediately on the subsequent question without index recomputation or daemon restarts. Furthermore, Zimmer Server records every retrieval query, verified group resolution, and denied access attempt directly into a tamper-resistant local audit log, providing compliance officers with a full evidentiary trail of internal retrieval activity.
Verifiable page-level citations and active hallucination refusal
In an enterprise environment, an AI summary without verifiable provenance is a liability. When employees make financial commitments, draft contracts, or verify technical architecture based on AI output, they must inspect the underlying source document with zero friction.
Zimmer Server binds every factual statement generated by the model to an interactive citation chip. When examining complex multi-page PDF files, the citation chip opens the exact document rendered to the specific page number containing the referenced paragraph. For Microsoft Word (.docx) and Markdown files, citations anchor directly to the relevant section heading. Employees can review the original context behind a generated synthesis in seconds, validating that quotes and numbers match executive records.
Equally crucial is the system's deliberate refusal behavior. Generic chatbots frequently exhibit confabulation—inventing plausible-sounding policies or financial figures when source data is ambiguous. Zimmer Server's retrieval prompt and system logic enforce strict evidence grounding: if the indexed documents do not contain direct evidence supporting the query, the model explicitly declares that the answer is not present in company files. Stating clearly what the documents do not say protects teams from costly mistaken assumptions.
Enterprise ingestion workflows and Apple Silicon hardware sizing
Document search is only as valuable as the recency of the files it indexes and the throughput of the underlying hardware. Ingesting company data must not require fragile manual exports, unencrypted staging servers, or per-page third-party OCR API charges. Zimmer Server supports native document workflows designed for corporate IT environments:
Point Zimmer Server directly at mounted network shares or local SSD directories. The indexing worker processes PDFs, scanned documents, Word documents (.docx), Markdown, and plain text files locally using on-device text extraction and Apple Silicon-accelerated embeddings.
- • Instant local parsing with $0 per-page processing fees
- • File-watch triggers automatically index modified or newly added files
- • Configurable chunking (512–1024 token segments with semantic overlap)
Connect corporate Microsoft 365 libraries via official Microsoft Graph delta polling. Zimmer Server syncs incremental changes in the background without re-indexing untouched corporate archives.
- • Scoped OAuth integration requiring minimal tenant read permissions
- • Outlook integration requests Mail.ReadWrite to draft replies, never Mail.Send
- • Annual and recurring proposals can regenerate section-by-section to .docx
Running team-wide private document search does not require a data center filled with high-voltage server racks. Apple Silicon's unified memory architecture provides high-bandwidth memory sharing across CPU and GPU cores at low thermal and electrical footprints, making compact desktop hardware like the Mac mini and Mac Studio ideal on-premise appliances.
The primary constraint when sizing hardware is unified memory capacity: memory must comfortably hold the active language model weights, the dense vector index and embedding model, and dynamic context windows for concurrent queries. The following deployment matrix outlines tested sizing baselines:
| Deployment Scale | Hardware Specification | Recommended Model Class | Target Concurrency |
|---|---|---|---|
| 5–15 seats | Mac mini (M4 / M4 Pro, 24 GB–32 GB Unified Memory) | 8B–14B models (e.g., Qwen 2.5 14B Q4_K_M) | 2–4 concurrent queries; ~28 tokens/sec generation |
| 20–50 seats | Mac Studio (M2 Max / M4 Max, 64 GB Unified Memory) | 14B–32B models or compact MoE architectures | 6–12 concurrent queries; sub-second vector filtering across 50k pages |
| 50–100+ seats | Mac Studio (M2 Ultra, 128 GB–192 GB) or dual Mac Studio cluster | 32B–70B models with dedicated fast embedding node | 15–30 concurrent queries with automatic failover and load balancing |
When an organisation grows beyond a single physical machine, Zimmer Server supports multi-node clustering. An administrator can enroll an additional Apple Silicon Mac in roughly two minutes using a one-time enrollment code. The appliance distributes query loads dynamically: fast embedding models and routine questions route to secondary nodes, while intensive multi-document synthesis queries route to nodes with larger memory headroom. If a node goes offline, pending requests automatically re-route without service interruption.
From a total cost of ownership perspective, a dedicated Mac Studio appliance running Zimmer Server operates at a flat hardware investment (amortised over three to four years) paired with Team licensing at $29 per user per month. In contrast, cloud document search platforms charge compounded per-seat license fees, variable per-token API costs, and ongoing per-gigabyte vector hosting fees that scale unpredictably as document volumes increase.
Frequently asked questions about private AI document search
IT leads, security reviewers, and operations managers evaluating on-premise document search require concrete operational answers. The following questions cover the technical boundaries, compliance scope, and practical administration of Zimmer Server:
How does in-retrieval RBAC prevent document leakage in AI search?
Zimmer Server evaluates group permissions as an active mathematical filter inside the vector database query before nearest-neighbor candidates are selected. Because restricted text chunks are pruned prior to candidate ranking, confidential passages never enter the model's context window, eliminating semantic leakage or prompt injection attacks across departments.
Can employees search internal documents without sending files to OpenAI or Microsoft?
Yes. Document parsing, embedding generation, vector indexing, retrieval, and open-weight model inference execute entirely on company-owned Apple Silicon hardware inside your private network. Prompts, extracted paragraphs, and answers never travel to external cloud APIs or third-party training pipelines.
What hardware does our company need to run private AI document search?
Zimmer Server requires an Apple-silicon Mac. A Mac mini with 24 GB to 32 GB unified memory supports 5 to 15 concurrent users for everyday document retrieval, while an M2 or M4 Max Mac Studio with 64 GB unified memory is recommended for 25 to 50 employees querying dense technical documents.
How does Zimmer Server handle remote employee access securely?
Remote devices join an end-to-end encrypted WireGuard-class mesh network via a single-use enrollment payload valid for 60 minutes. Device keypairs are sealed in the client's local Keychain, requiring zero passwords, no open firewall ports, and no public IP forwarding.
What are the operational trade-offs and limits of an on-premise AI search appliance?
An on-premise appliance requires owned hardware, local backup management, and relies on open-weight models that trail peak cloud frontier reasoning. Zimmer Server currently requires Apple Silicon, with native Windows server hosting and Linux GPU nodes remaining on the roadmap.
Deploying an on-premise AI search appliance involves real engineering trade-offs. The administrator owns hardware reliability, local physical access controls, and routine backup schedules. Zimmer Server provides one-click backup and restore with cryptographically verified integrity manifests, but storage redundancy remains an internal IT responsibility.
Furthermore, while open-weight models (such as Qwen 2.5 and Llama 3.3) excel at grounded document extraction, summarization, and citation verification, hosted frontier models (such as Claude 3.5 Sonnet or GPT-4o) retain higher general reasoning capabilities for complex abstract synthesis. Organizations adopt Zimmer Server when data confidentiality, regulatory compliance, and cost predictability outweigh peak frontier capabilities.
Finally, Zimmer AI does not assert third-party certifications such as HIPAA, ISO 27001, or SOC 2 on customer deployments. Deployment on internal hardware provides sovereign data residency by design, but organizations remain responsible for their own regulatory compliance and governance policies.
Test private document search on an Apple Silicon Mac inside your own network. Set up group permissions, index your internal PDF and SharePoint files, and verify in-retrieval RBAC firsthand with zero upfront licensing fees.