Run Claude Code 100% on-device with local AI on Apple Silicon.
No cloud, no API key, no proxy — an MLX-native server that speaks the Anthropic API.
🥊 Pick your fighter: Gemma 4 31B · Llama 3.3 70B · Qwen 3.5 122B · DeepSeek V4 Flash (1M context via ds4).
🛑 Usage Limit · 🤔 What Is This · 🚀 Quick Start · 🥊 Lineup · 🎮 Modes · 🔒 Privacy · 📊 Benchmarks · 🎤 Voice · 🌐 Browser · 📱 Phone · 🔌 MCP · 🛣️ Roadmap
If Claude Code just told you "you've reached your usage limit" and gave you a reset time hours away, that's what this is for. You keep working — same Claude Code, same terminal, same project — except the model answering is running on your own Mac.
curl -fsSL https://raw.githubusercontent.com/nicedreamzapp/claude-code-local/main/install.sh | bashNo API key. No second subscription. No waiting until 3pm. It works on a 16 GB MacBook and gets better the more RAM you have — see what runs on your Mac.
Real session, unedited. Claude Code reads and edits the file — the model answering is Gemma 4 31B on the laptop.
Your Mac has a powerful GPU built right into the chip. This project uses that GPU to run massive AI models — the same kind that power ChatGPT and Claude — entirely on your computer, and plugs them into Claude Code so the whole coding experience works offline.
🚫 No internet needed 💰 No monthly subscription 🔒 No one sees your code or data ✅ Full Claude Code experience — write code, edit files, manage projects, control your browser, or run a full hands-free voice session
📱 You (Mac or Phone)
│
🤖 Claude Code ← the AI coding tool you know
│ HTTP localhost:4000
⚡ MLX Native Server ← this repo (~1000 lines of Python)
│
🥊 Pick your fighter ← Gemma 4 31B · Llama 3.3 70B · Qwen 3.5 122B
│
🖥️ Apple Silicon GPU ← your M-series chip does all the work
The trick: Claude Code speaks the Anthropic API. Local model servers speak the OpenAI API. So everyone bolts a translation proxy in between — and the proxy is slow and fragile. This server speaks Anthropic natively. One process, zero translations:
| 🐌 What everyone else does | 🚀 What we did |
|---|---|
| Claude Code → Proxy → Ollama → Model | Claude Code → Our Server → Model |
| 3 processes, 2 API translations | 1 process, 0 translations |
| 133 seconds per task | 17.6 seconds per task |
🎯 That one change — eliminating the proxy — made it 7.5× faster.
A real NDA. Llama 3.3 70B. Wi-Fi physically OFF. lsof running live. Watch a 70-billion-parameter model audit a confidential legal document, on-device, with the receipts on screen.
AirGap is this whole build running as one private workstation — a capability, not a product. Everything you need is in this repo. If your firm needs one built, here's what it looks like.
More local-AI demos on the channel:
| Video | What happens |
|---|---|
| 🌌 The Rematch | 4 AI engines build northern lights, 3 fully local — the local challenger painted the best aurora |
| 🏁 Hexagon Shootout | Gemma 31B vs Llama 70B vs cloud Claude, same physics prompt, live counters — 2 of 3 with zero cloud calls |
| 🐳 DeepSeek Three-Way | DeepSeek V4 Flash local beats cloud Claude on wall-clock, same MacBook |
| 🎤 NarrateClaude | Speak to Claude Code, hear replies in a cloned voice — 100% on-device |
| 🏠 Mac mini as home AI | Chat with the Mac mini at home from any browser on any phone |
We started with one model. Now we ship a roster. Same MLX server, same Anthropic API — swap one env var and you swap the brain. Plus the ds4 engine for DeepSeek V4 Flash via its own native Metal runtime.
| 🟡 Hermes 4 14B | 🟢 Gemma 4 31B | 🟠 Llama 3.3 70B | 🔵 Qwen 3.5 122B | 🐳 DeepSeek V4 Flash ⭐ | |
|---|---|---|---|---|---|
| Nickname | The One That Runs On Your Laptop | The Quick One | The Wise One | The Beast | The 1M-Context Whale |
| Build | 4-bit abliterated | 4-bit IT abliterated | 8-bit abliterated | 4-bit MoE (A10B) | 2-bit asymmetric (ds4 GGUF) |
| Speed | not benchmarked yet | ~15 tok/s | ~7 tok/s | 65 tok/s 🚀 | ~32 tok/s |
| Params | 14 B dense (Qwen3 base) | 31 B dense | 71 B dense | 122 B / 10 B active | 284 B / 37 B active |
| Context | 40 K | 128 K | 128 K | 256 K | 1 M tokens |
| RAM | ~8 GB | ~18 GB | ~70 GB | ~75 GB | ~81 GB |
| Min RAM to run | 16 GB | 32 GB | 96 GB | 96 GB | 128 GB |
| Best at | Everyday edits on a stock MacBook | Daily coding | Hardest reasoning, full precision | Max throughput, active sparsity | Long context, agentic loops |
| Engine | MLX Native | MLX Native | MLX Native | MLX Native | antirez/ds4 |
| Launcher | Claude Local.command |
Gemma 4 Code.command |
Llama 70B.command |
Claude Local.command |
DeepSeek V4 Flash.app |
💻 Got a 16 GB MacBook Air? Start with Hermes.
setup.shpicks it for you automatically — you don't need 96 GB of RAM to use this.
💡 Fun fact: Qwen wins raw speed because it's an MoE — only 10B of 122B params activate per token. DeepSeek V4 Flash is even bigger (284B) but only ~37B active per token, and it ships with on-disk KV cache so a 25k-token Claude Code system prompt prefills exactly once, ever.
We tested it the day Antirez (the Redis guy) shipped ds4. Local DeepSeek beat cloud Claude on wall-clock time on the same MacBook, same prompt — watch the three-way.
| 🧠 Engine | antirez/ds4 — pure C + Metal kernels, ~few thousand lines |
| 🤗 Weights | antirez/deepseek-v4-gguf (q2: 81 GB, q4: 153 GB) |
| 📦 Server wrapper | ~/.local/bin/ds4-server-up (boots on demand) |
| 🚀 Claude Code wrapper | ~/.local/bin/claude-ds4 (drop-in replacement for claude) |
| 📏 Context | 1 M tokens; 200 K is sane for most agent runs |
| 💾 Disk KV cache | Persists across restarts — first prefill is the only one that ever happens |
The models in this lineup aren't from generic mirrors — we package and upload our own abliterated MLX builds to HuggingFace so anyone running this repo can pull them with one command. Browse the full set at huggingface.co/divinetribe.
# Llama 3.3 70B — full-precision feel
MLX_MODEL=divinetribe/Llama-3.3-70B-Instruct-abliterated-8bit-mlx \
bash scripts/start-mlx-server.sh
# Gemma 4 31B — fast daily driver
MLX_MODEL=divinetribe/gemma-4-31b-it-abliterated-4bit-mlx \
bash scripts/start-mlx-server.sh
# Hermes 4 14B — sweet spot for 16/32 GB Macs
MLX_MODEL=divinetribe/Hermes-4-14B-abliterated-4bit-mlx \
bash scripts/start-mlx-server.sh| Model | Quant | Disk | Params | Context | Best for |
|---|---|---|---|---|---|
Llama-3.3-70B-Instruct-abliterated-8bit-mlx |
8-bit, g64 | ~75 GB | 71 B dense | 128 K | Hardest reasoning on 96 GB+ Macs |
gemma-4-31b-it-abliterated-4bit-mlx |
4-bit, g64 | ~17 GB | 31 B dense | 128 K | Daily coding on a 32 GB+ Mac |
Hermes-4-14B-abliterated-4bit-mlx |
4-bit, g64 | ~8 GB | 14 B dense (Qwen3 base) | 40 K | 16 GB Macs, instruction-following, tool use |
Abliteration sources: huihui-ai (Llama, Gemma) and Babsie (Hermes). MLX conversion + quantization by us. See what abliteration means.
⚠️ Use it responsibly. "Abliterated" suppresses the model's built-in refusal direction so it doesn't refuse benign-but-edgy requests. It is not a general capability upgrade, and you remain bound by each upstream license (Llama 3.3, Gemma, Hermes/Qwen3).
Four ways to run the lineup. Each one is a double-clickable launcher in launchers/.
| Mode | What it does | Launcher |
|---|---|---|
| 🤖 Code | Run Claude Code with a local model — same UX, no API key | Claude Local.command, Gemma 4 Code.command, Llama 70B.command |
| 🌐 Browser | Local AI controls real Brave browser via Chrome DevTools | Browser Agent.command |
| 🎤 Hands-Free Voice | Speak in, hear replies in your cloned voice — full loop, 100% on-device | Narrative Gemma.command + NarrateClaude |
| 📱 Phone | iMessage in → text/image/video out, via claude-screen-to-phone | ~/.claude/imessage-*.sh |
| Your Mac | RAM | What setup.sh installs for you |
|---|---|---|
| MacBook Air / base M1-M4 | 16 GB | 🟡 Hermes 4 14B — yes, this works |
| M1/M2/M3/M4 Pro | 32-48 GB | 🟢 Gemma 4 12B |
| M2/M3/M4/M5 Max | 64-95 GB | 🟢 Gemma 4 31B |
| M3/M4/M5 Max · Ultra | 96 GB+ | 🔵 Qwen 3.5 122B, 🟠 Llama 70B, 🐳 DeepSeek |
Also need:
- 🐍 Python 3.12+ (for MLX)
- 🤖 Claude Code (
npm install -g @anthropic-ai/claude-code)
curl -fsSL https://raw.githubusercontent.com/nicedreamzapp/claude-code-local/main/install.sh | bashOr clone it yourself if you'd rather read the script first:
git clone https://github.com/nicedreamzapp/claude-code-local
cd claude-code-local
bash setup.shsetup.sh auto-detects your RAM, picks a model from the lineup, downloads it, installs the MLX server, and creates a Claude Local.command launcher on your Desktop.
Then double-click Claude Local.command. You're coding locally.
🐛 If the launcher asks you to sign in to a Claude account: your
claudeCLI is too old. The launchers pass--bareto force local-only API-key auth; older CLIs don't support it. Fix:npm install -g @anthropic-ai/claude-code
🛠️ Note for contributors:
setup.shinstalls the server as a symlink at~/.local/mlx-native-server/server.pypointing back at this repo'sproxy/server.py. Edit the file in the repo, restart the server, done — one source of truth, no silent drift.
# 1. Set up the MLX virtualenv
python3.12 -m venv ~/.local/mlx-server
~/.local/mlx-server/bin/pip install mlx-lm
# 2. Pick a fighter and download (one time, ~18-75 GB)
bash scripts/download-and-import.sh gemma # or 'llama' or 'qwen'
# 3. Start the server
MLX_MODEL=divinetribe/gemma-4-31b-it-abliterated-4bit-mlx \
bash scripts/start-mlx-server.sh
# 4. Launch Claude Code
ANTHROPIC_BASE_URL=http://localhost:4000 \
ANTHROPIC_API_KEY=sk-local \
claude --model claude-sonnet-4-6┌──────────────────────────────────────────────────┐
│ Your MacBook (M-series) │
│ │
│ 📝 You type ──> 🤖 Claude Code │
│ │ │
│ ▼ │
│ ⚡ MLX Server (port 4000) │
│ │ │
│ ▼ │
│ 🥊 Local model ──> 🖥️ GPU │
│ (Gemma·Llama·Qwen) │
│ │ │
│ ▼ │
│ 📝 Answer <─── ✨ Clean response │
│ │
│ 🔒 Nothing leaves this box. Ever. │
└──────────────────────────────────────────────────┘
The server (proxy/server.py) is one file, ~1000 lines. It does six things:
- 📦 Loads the model — Apple's MLX framework, native Metal GPU, unified memory. Handles Gemma's
RotatingKVCachequirk automatically. - 🔌 Speaks Anthropic API — Claude Code thinks it's talking to Anthropic's cloud. It's not.
- 🔧 Translates tool use — Three tool-call formats in and out: Gemma 4 native, Llama 3.3 raw JSON, and HuggingFace
<tool_call>JSON (Qwen and others). All converted ↔ Anthropictool_useblocks, with garbled-output recovery for small models. - 🧹 Cleans the output — A real-time
ThinkingFilterstrips<think>blocks token-by-token during generation, thenclean_responsehandles stop markers and reasoning preamble. - ⚡ Reuses prompt caches across requests — Claude Code's system prompt doesn't get re-prefilled every turn. Huge speedup for short questions.
- 🎯 Code mode — auto-detects Claude Code coding sessions, swaps the ~10K-token harness prompt for a slim ~150-token one, and strips verbose tool descriptions to name + parameter types. A 28× prompt reduction that cuts prefill from ~60 s to ~2 s on Gemma 4 31B.
We didn't start here. Three generations in one night:
| Gen | What We Tried | Speed | 💡 What We Learned |
|---|---|---|---|
| 1️⃣ | Ollama + custom proxy | 30 tok/s | Ollama works but Claude Code can't talk to it directly |
| 2️⃣ | llama.cpp TurboQuant + proxy | 41 tok/s | TurboQuant compresses KV cache 4.9x, but the proxy is the bottleneck |
| 3️⃣ | MLX native server | 65 tok/s | Kill the proxy. Speak Anthropic API directly. 7.5x faster. |
| 4️⃣ | The lineup | 65 / 15 / 7 tok/s | Three brains, one server — swap one env var to change the fighter |
This is the part we're proudest of. Your code never leaves your Mac. Not for a model call. Not for telemetry. Not for "anonymous analytics". Not ever.
┌─────────────────────────────────────────────────────────────┐
│ 🖥️ YOUR MACBOOK │
│ │
│ 📝 Your code ──> 🤖 Claude Code ──> ⚡ MLX Server │
│ (localhost:4000) │ │
│ ▼ │
│ 🧠 Local model ──> 🖥️ Apple GPU │
│ │
│ 🚫 ZERO outbound network calls │
│ 🚫 ZERO telemetry │
│ 🚫 ZERO phone-home │
└─────────────────────────────────────────────────────────────┘
│
✗ ← Nothing from our code crosses this line.
│
┌─────────────────────────────────────────────────────────────┐
│ ☁️ THE INTERNET │
│ (your code never goes here) │
└─────────────────────────────────────────────────────────────┘
| Component | Source | Outbound calls | Verdict |
|---|---|---|---|
| server.py (ours) | We wrote it line by line | 0 | ✅ Safe |
| browser agent | nicedreamzapp/browser-agent — we wrote it | 0 (localhost CDP only) | ✅ Safe |
| mlx-lm | Apple ML team | 0 | ✅ Safe |
| MLX framework | Apple | 0 | ✅ Safe |
| Model weights | HuggingFace verified repos | 0 at runtime | ✅ Safe |
| Claude Code CLI | Anthropic (closed-source binary) | 0 with our launchers — lsof-verified |
✅ Safe |
✅ Verified offline. Claude Code's own binary previously reached out to
api.anthropic.comon startup for telemetry, statsig feature flags, marketplace auto-install, and the autoupdater. The launchers plug all four channels via documented Anthropic env vars (thanks @tadrianonet, PR #32):CLAUDE_CODE_DISABLE_NONESSENTIAL_TRAFFIC=1 DISABLE_AUTOUPDATER=1 CLAUDE_CODE_DISABLE_OFFICIAL_MARKETPLACE_AUTOINSTALL=1 CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1Verify it yourself: run
lsof -p $(pgrep -f claude)during a session — you'll see onlylocalhost:4000. Runlsof -i -Pwhile the server is up — nothing leaves your Mac.
⚠️ We removed LiteLLM after supply-chain attack concerns. Every dependency was re-audited from scratch. If a package had unexplained network calls, it didn't ship.
| Situation | Use This? | Why |
|---|---|---|
| On a plane (no wifi) | ✅ | Full AI coding, no internet needed |
| NDA / sensitive client code | ✅ | Nothing leaves your machine — air-gapped, lsof-verified |
| Healthcare / legal / finance review | ✅ | 100% on-device, audit-friendly |
| Don't want API fees | ✅ | $0/month forever |
| Want fastest possible | ☁️ | Cloud Sonnet is still slightly faster |
| Need Claude-level reasoning | ☁️ | Local models are good, not Claude-level |
Three generations of optimization. Each one got faster.
| Generation | Approach | Speed | Real Claude Code task |
|---|---|---|---|
| 🐌 Gen 1 | Ollama + Proxy | 30 tok/s | 133 s |
| 🏃 Gen 2 | llama.cpp + Proxy | 41 tok/s | 133 s |
| 🚀 Gen 3 | MLX Native (ours) | 65 tok/s | 17.6 s |
| Model | tok/s | RAM | Best For |
|---|---|---|---|
| 🟢 Gemma 4 31B Abliterated | ~15 | ~18 GB | Daily coding on a 64 GB Mac |
| 🟠 Llama 3.3 70B Abliterated | ~7 | ~70 GB | Hardest reasoning, full precision |
| 🔵 Qwen 3.5 122B-A10B | 65 | ~75 GB | Maximum throughput, MoE sparsity |
| 🖥️ Our Local Setup | ☁️ Claude Sonnet | ☁️ Claude Opus | |
|---|---|---|---|
| Speed | 65 tok/s | ~80 tok/s | ~40 tok/s |
| Monthly cost | $0 🎉 | $20-100+ | $20-100+ |
| Privacy | 100% local 🔒 | Cloud | Cloud |
| Works offline | Yes |
No | No |
💡 Our local setup beats cloud Opus on raw speed (65 vs 40 tok/s) at $0/month. Qwen numbers measured on M5 Max 128 GB — full details in BENCHMARKS.md.
Local models don't format tool calls perfectly. They want to call a tool but mix XML and JSON syntax — Claude Code sees no valid tool call, re-prompts, and the model garbles it the same way again. The result: infinite loops where the AI says "let me do that" but never does anything.
We fixed this with 4 changes to server.py:
| Change | What | Why |
|---|---|---|
| KV Cache | 4-bit → 8-bit, quantization starts at token 1024 | Model retains conversation context |
| Temperature | 0.7 → 0.2 | Less randomness = more consistent tool formatting |
| Garbled Recovery | recover_garbled_tool_json() |
Catches XML-in-JSON hybrids, infers tool names from parameter keys |
| Retry Logic | Up to 2 retries when tool intent is detected but parsing fails | Re-prompts with explicit formatting instructions |
🧪 Results: 98/98 tests passed across 7 consecutive runs. Zero failures. The multi-step scenario that used to trigger infinite loops — create 12 month folders, delete all but September, verify — now passes every time. Run it yourself:
python3 scripts/test_mlx_server.py| Variable | Default | What It Does |
|---|---|---|
MLX_MODEL |
divinetribe/gemma-4-31b-it-abliterated-4bit-mlx |
Pick which fighter to load |
MLX_KV_BITS |
8 |
KV cache quantization bits (4 saves memory, 8 improves coherence) |
MLX_KV_QUANT_START |
1024 |
Token position where KV quantization begins |
MLX_TOOL_RETRIES |
2 |
Max retries when a garbled tool call is detected |
MLX_MAX_TOKENS |
8192 |
Max output tokens per response |
MLX_SUPPRESS_THINKING |
1 |
Skip the model's reasoning chain (~1 min/request saved). Set 0 to let it think. |
MLX_BROWSER_MODE |
0 |
Optimize for chrome-devtools MCP sessions — keeps only the 9 essential browser tools (~99% fewer tokens) |
Everything above gets you running. These live in docs/ so this page stays short:
| 🎤 Hands-Free Voice Mode | Talk to Claude Code, hear it answer in a cloned voice |
| 🌐 Browser Agent | Let the local model drive your real browser |
| 📱 Control From Your Phone | Run a session on your Mac from anywhere |
| 🔌 MCP Servers | Claude Code's whole plugin ecosystem, 100% local |
| 📁 What's In This Repo | File-by-file tour |
| 📊 Benchmarks · 🔧 Tool-Call Reliability | The numbers and how they were measured |
| 📱 Apps From The Same Workshop · 🙏 Credits | Everything else |
claude-code-local is the brain. It pairs with sibling repos — each stands alone, together they take Claude Code off the keyboard and off the screen:
| Repo | Role | What it does |
|---|---|---|
| 🤖 claude-code-local | Brain (you are here) | MLX Anthropic server · launcher lineup · tool-call translation |
| 🎤 NarrateClaude | Ears + Mouth | Talk to Claude, hear replies in your cloned voice — both directions on-device |
| 🌐 browser-agent | Hands | Drives real Brave via CDP — iframes, Shadow DOM, ProseMirror |
| 📱 claude-screen-to-phone | Remote | iPhone → Claude Code over iMessage; text/screenshots/videos back |
| 🛟 claude-failover | Backstop | Keep cloud Claude primary, flip one command to local when limits pinch or Anthropic is down |
We ship fast and in public. If any of these excite you, hit Watch to get the release ping.
- 🟡 Full Qwen 3.5 122B benchmark suite — reliability, tool-call pass rate, long-context behavior vs Gemma
- 🟡 Fully-local Whisper fallback — alternative to the Apple
SFSpeechRecognizerpath for older Macs and non-English voices - 🟡 One-click DMG installer — no terminal needed
- 🟡
MLX_MODEL=<hf-url>— point at any HuggingFace repo and auto-register a new fighter - 🟡 More fighters — open to PRs adding launchers for DeepSeek, Mistral, Phi, anything MLX-compatible
💡 Want something that's not on this list? Open an issue → Every serious request gets read and usually replied to within 24h.
Ideas, bug reports, a new launcher for a model I don't run, a better code-mode prompt — open an issue or a PR, I read them all. Especially interested in: folks on older Apple Silicon (M1/M2, 16–36 GB) who know which models actually fit; anyone stress-testing the voice loop on different hardware or accents; TTS recipes beyond Pocket TTS (Piper, MLX-TTS, Kyutai Moshi); and edge cases I'll never hit on an M5 Max with 128 GB.
Every one of these landed on hardware I don't own, on a bug I hadn't hit. Thank you.
| Who | What they fixed |
|---|---|
| @0xshugo | Client disconnects handled, retries skipped when there are no tools (#4) |
| @asdmoment | Gemma inference crash — auto-disable KV quantization (#7) |
| @kulveersingh | ArraysCache has no attribute offset (#10) |
| @tripathiprateek | uninstall.sh — reverses setup.sh cleanly (#23) |
| @tadrianonet | Mac base/Pro 16 GB support: Qwen 2.5 14B, ChatML stop markers, <tools> parser, offline leak fix (#32) |
| @kevbarns | Gemma 4 thinking suppression + slimmer tool descriptions — ~4× latency cut (#33) |
| @KaoCSC | Stop on the tokenizer's real EOS, and tolerate empty env ints (#41) · bare JSON tool calls, which took Qwen 2.5 Coder from 0/12 to 14/14 (#43) |
Tested on Apple M5 Max with 128 GB unified memory.
Built by Matt Macosko in Arcata, CA — part of Nice Dreamz LLC. More open-source at nicedreamzwholesale.com/software · demos at youtube.com/@nicedreamzapps.
📜 MIT License — Use it however you want.
💬 Builders hang out on Discord — share what you're building, swap MLX tips.
⭐ Star this repo if it helped you! ⭐
Upstream projects this is built on are listed in docs/CREDITS.md.
