Local AI agent for Mac and Windows
A local AI agent is not merely a chat window attached to a model. It is a loop that reads context, chooses a bounded tool, observes the result, and decides what to do next. The useful question is not whether it feels autonomous. It is where each action happens and where you can stop it.

Start with one tool loop, not the word “agent”
Ask a local agent to explain a failing test. A useful loop reads the test, searches for the implementation, proposes a cause, requests permission before a write or command, observes the test result, and reports evidence. A model that only produces an answer has not completed that loop. A script that always follows fixed steps is a workflow, not an agent deciding its next action.
Anthropic's guide to effective agents makes the same practical distinction between predefined workflows and systems in which a model directs its own process and tool use. It also recommends simple, composable patterns and environmental feedback instead of complexity for its own sake.
Zimmer AI packages that loop as a personal local AI workspace for one Mac or Windows PC. With a local model, inference, repository context, tool calls, and proposed edits stay on that machine. Zimmer Server is a different product for a different situation: a multi-user appliance on company-owned Apple-silicon hardware for private team networks.
Trace the six boundaries of a local task
“Local” describes a data path, not a feeling. Follow a task across six checkpoints before deciding that an agent meets your privacy or control requirement.
| Checkpoint | Question to answer | Evidence to keep |
|---|---|---|
| Model | Does inference run locally or at a hosted endpoint? | Selected model and endpoint. |
| Context | Which files and conversation turns enter the request? | Named scope and exclusions. |
| Decision | Does the model choose the next action? | Plan and stated reason. |
| Tool | Can it read, write, run a command, or call a service? | Exact tool request. |
| Permission | What stops a risky action before it runs? | Allow, Ask, or Deny decision. |
| Result | Did the environment confirm the claimed outcome? | Test output and final diff. |
The model can be local while a connected tool is remote. An MCP connector to GitHub or Slack crosses a network boundary by design; a hosted OpenAI-compatible endpoint does too. Conversely, a local model can still make a bad edit. Privacy and correctness are separate checks, and both need evidence.
Match six roles to nine real tools
Zimmer ships six built-in agent types: Assistant, Coder, Reviewer, Tester, Refactorer, and Documenter. You can create custom roles with their own instructions, icons, and model assignments. Roles organize responsibility; they do not grant magical competence or replace the permission layer.
The harness exposes nine tools: read_file, grep, edit_file, write_file, run_command, todo_write, spawn_agent, use_skill, and mcp_tool. A read-only diagnosis may need only the first two. A verified fix may add one edit and one focused test command. A connected-service task adds a separate external data path.
The local model layer is based on bundled llama.cpp for GGUF inference. The official llama.cpp server documentation describes local CPU/GPU inference, OpenAI-compatible endpoints, schema-constrained output, and function calling. Those capabilities supply a model interface; the workspace still must decide which context, permissions, tools, and review steps surround it.
Put permission directly in the action path
Zimmer uses three permission states: Allow, Ask, and Deny. Read-only operations and safe shell commands can be pre-approved. Other actions show the exact request and offer allow once, allow always, or deny. Destructive commands such as disk formatting, privilege escalation, and piping a remote script into a shell are denied out of the box.
- Begin with read-only search over a named folder.
- Require the agent to explain the smallest intended change.
- Approve one write only after checking its target.
- Approve a focused verification command separately.
- Inspect the complete side-by-side diff before accepting it.
A permission prompt is useful only when it appears before the consequential action and names what will happen. “Agent mode enabled” is not a control. A Reviewer role is not a control either. The reliable boundary is the combination of narrow tool access, an explicit approval point, environmental output, and a diff you can reject.
Delegate without copying the whole conversation
Zimmer can hand work to another agent, run two agents in a split pane, or use spawn_agent. A spawned subagent receives isolated context, works for up to eight tool rounds, and returns only a summary. The parent agent can use up to 15 tool rounds in its own turn. These are loop limits, not promises that a task will succeed.
Give a subagent a contract rather than a vague role: one question, the relevant files, constraints, expected evidence, and a stopping condition. For example, ask Reviewer to find one unsupported assumption in a proposed patch and cite the file and test that expose it. Keep the parent responsible for deciding whether that evidence changes the plan.
Subagent contract - Question: why does the focused test fail? - Context: component.tsx and component.test.tsx only - Tools: read_file and grep - Evidence: quote the failing branch and test expectation - Stop: return a summary; do not edit or run commands
Context isolation keeps the delegate focused and reduces what it carries back. It is not a security sandbox by itself. Tool permissions still determine what the subagent can attempt, and the human still decides whether to accept an edit or command.
Choose a model that leaves room for the work
Zimmer Desktop and its multi-agent coding harness run on macOS and Windows. On Apple silicon, a 16 GB Mac with a 4B-class model at Q4 is a practical floor for useful local coding. The 32–36 GB band is more comfortable for 14B-class models or larger Mixture-of-Experts models. Those numbers are Mac guidance, not a claim about Windows GPU acceleration.
The app's Model Hub sorts candidates using RAM, chip, and free disk, and lets you choose quantization variants such as Q4_K_M or Q5_K_M. Default context is 16k, with presets from 8k through 128k. A bigger context consumes more memory, and the model must share the machine with the editor, browser, build, tests, and the agent workspace.
Do not choose a model from a leaderboard alone. Give it a real task that requires tool selection, a denied action, a small edit, and recovery from a failed check. Tool reliability is part of agent quality. Use the model and hardware-fit guide for the model layer and the agent controls page for the surrounding workflow.
Run a seven-minute action-boundary test
Use a small repository with a known failing test. Ask the agent to diagnose it without editing. Deny its first command request and check whether it recovers cleanly. Then approve one focused edit and one focused test. Finally, inspect every changed file before accepting the diff.
| Test question | Pass condition |
|---|---|
| Can it explain its next action? | The request names the tool, target, and reason. |
| Does Deny actually stop it? | Nothing runs, and the agent proposes a safer route. |
| Can it stay inside scope? | Only the named file changes. |
| Does it prove success? | The focused test output appears in the record. |
| Can you reject the change? | The diff is visible before it touches disk. |
Repeat the same task with the network disconnected after the model is downloaded. Local inference and repository tools should continue; model downloads, software updates, hosted endpoints, and connected services should fail at their own boundaries. For a longer rehearsal, use the offline coding acceptance test.
Know when a local agent is the wrong choice
A hosted frontier model from Claude, GPT, or Gemini will outperform a laptop-size local model on many difficult reasoning tasks. If your machine is not capable, cloud AI can also be cheaper to start. Local execution trades peak capability for privacy, ownership, offline continuity, and a predictable cost after the hardware is paid for.
Zimmer AI is not a full IDE or a replacement for version control, tests, or human review. Connecting a hosted endpoint changes the inference boundary. Calling an MCP service sends the data required for that tool to its provider. Granting a shell command can change more than the file named in the prompt. A strong local workflow states those limits instead of hiding them behind “private” or “autonomous.”
If your main need is coordinating specialist roles rather than understanding the local action loop, read the separate local AI agent manager guide. It focuses on role handoffs, review evidence, and when a second agent earns its complexity.
Questions about local AI agents
What is a local AI agent?
A local AI agent combines a model, selected context, and tools on your computer. It can choose actions such as reading a file, searching a repository, proposing an edit, or running an approved command. A useful local setup also exposes the endpoint, permissions, tool results, and final diff instead of hiding the action loop.
Can Zimmer AI agents run on both Mac and Windows?
Yes. Zimmer Desktop runs local GGUF models and its coding-agent harness on Apple-silicon macOS and Windows x64 or arm64. MLX, system-wide dictation, and JJ voice mode are Mac-only, but the agent roles, repository tools, permission prompts, and diff review are part of the Mac and Windows desktop workflow.
Which roles and tools does Zimmer AI include?
Zimmer includes Assistant, Coder, Reviewer, Tester, Refactorer, and Documenter roles, plus custom agents. Its nine tools cover file reading, search, editing, writing, commands, task tracking, subagent delegation, skills, and MCP connections. Choose the smallest role-and-tool combination that can produce the evidence your task needs.
How are local agent actions controlled?
Zimmer assigns tools an Allow, Ask, or Deny state. Read-only work and safe commands can be pre-approved; other actions show the exact request for one-time or persistent approval. Proposed edits appear in a side-by-side Monaco diff before they touch disk, while destructive command patterns are denied by default.
Does a spawned subagent receive my full conversation?
No. A Zimmer subagent receives isolated context, can work for up to eight tool rounds, and returns a summary to the parent. Give it the exact question, relevant files, constraints, expected evidence, and stopping point. Context isolation keeps delegation focused, but tool permissions and human diff review still remain necessary.
Test one bounded local agent loop
Zimmer Desktop is free forever for personal and commercial use. Download it, choose a model that fits, and verify the action boundary on a real task.