Weighing hardware + infra investment against Anthropic API rates, and whether a hybrid split beats either extreme. Prepared by keeper:hearth, generated 2026-07-08.
The short version: self-hosting doesn't beat the API at forge's current scale and usage pattern — the constraint is utilization, not price. A small, cheap pilot is worth running to learn whether that changes; a big hardware bet is not worth making yet.
The open-weight field has closed most of the gap on routine coding and reasoning. It has not closed the gap on the hardest, most agentic single-benchmark work. Numbers below are cross-cited across several 2026 model-comparison sites (BenchLM.ai, aimadetools, MindStudio-style aggregators) rather than primary release papers in every case — treat exact percentages as EST directionally-right, not precise. Artificial Analysis's blended index is the more credible cross-check where cited.
| Model | Size | License | Notable benchmarks |
|---|---|---|---|
| DeepSeek V4-Pro Apr 2026 | 1.6T MoE / 49B active | MIT | SWE-bench Verified 80.6%, LiveCodeBench 93.5%, Codeforces 3206. Cheapest frontier-adjacent coding model via API ($0.44/$0.87 per M tok) — the benchmark it sets for what self-hosting is competing against. |
| GLM-5.2 (Zhipu) Jun 2026 | ~744B MoE / 40B active | MIT | 1M context, DeepSeek Sparse Attention. Leads SWE-Bench Pro / Terminal-Bench 2.1 / FrontierSWE among open models; predecessor GLM-5.1 hit SWE-Bench Pro 58.4 vs. Opus 4.6's 57.3 — a rare case of an open model edging a recent Opus point release on one benchmark. |
| Kimi K2.6 (Moonshot) | 1T MoE / 32B active | Open | 256K context, MLA attention + vision encoder. Best-regarded open model for long-horizon/autonomous coding; co-leads the open field on Artificial Analysis's blended Intelligence Index (~54). |
| Qwen3.5 / 3.6 (Alibaba) | 0.8B–397B-A17B flagship | Apache 2.0 | Most permissive license of the group. Qwen3.5-Coder-72B: LiveCodeBench 61.8% vs. GPT-5's 64.2% — genuinely close on this one. Flagship: GPQA-D 88.4, AIME 91.3. Notably, the 27B dense Qwen3.6 variant now beats Alibaba's own 397B MoE flagship on coding — size isn't a clean proxy for capability anymore. |
| Llama 4 (Meta) Apr 2025, still current — no Llama 5 | Scout 109B/17B active; Maverick 128 experts | Custom, EU-restricted | Highest open MMLU (85.5%), but a signal worth naming: Meta's own proprietary Muse Spark model scores 52 on the AA Index vs. Maverick's 18 — Meta itself appears to be de-prioritizing its open line as the frontier path. |
| Mistral Large 3 | 675B MoE / 41B active | Open-weight | Strong HumanEval (~92%) but weak GPQA-D (~44%) — a non-reasoning model, clearly behind specifically on hard reasoning. |
| Precision | VRAM/RAM needed | What that looks like |
|---|---|---|
| Full FP16 | ~2TB | A 1T-class model at full precision needs a multi-node GPU cluster — out of reach for any single-box setup, ours included. |
| INT4 / Q4 the realistic self-host target | ~500–600GB combined | E.g. Kimi K2.6 at Q4 — needs 8×H200, 4×A100-80GB, or an 8×RTX-4090 aggregate. Q2 gets to ~350GB with real quality loss. |
| CPU-only, no GPU matches our forge box: 12-core / 31GB RAM | 1–5 tok/s | Memory-bandwidth-bound (~100GB/s DDR → a theoretical floor around 2.8s/token for an unquantized 70B dense model). A 31GB box can't even load any model in the table above at 4-bit. Ceiling is a small (<14B), heavily-quantized model at low single-digit tok/s. Not viable for real work. |
| Mac Studio, 512GB unified memory | 16–20 tok/s | The realistic prosumer ceiling. A DeepSeek-V3-class 671B MoE model at 4-bit runs at this rate via mlx-lm/llama.cpp — usable, but far slower than API latency. AVAILABILITY Apple reportedly pulled the 512GB config in 2026 amid a DRAM price spike; no M4/M5 Ultra exists yet, so this ceiling may not be purchasable new right now. |
The honest gap, in two numbers. On Artificial Analysis's blended Intelligence Index (Jul 2026), Opus 4.8 leads at 55.7% with the best open models (Kimi K2.6, Xiaomi MiMo V2.5) around 54 and DeepSeek V4 Pro around 52 — genuinely close, a 3–6 point gap on a broad blended measure. But on SWE-bench Verified specifically — the single benchmark closest to what forge actually does — Opus 4.8's 88.6% vs. the best open score around 77.8% is an ~11-point gap, and it's concentrated exactly where it matters: long-horizon autonomous multi-file work requiring self-correction with minimal supervision, plus multimodal reasoning and long tool-use chains. Open models have gotten good at routine coding and reasoning. They have not closed the gap on the hardest agentic work forge actually leans on Opus for.
All figures EST unless noted — sourced from 2026 GPU/cloud pricing pages, cross-checked across several vendors. The headline finding isn't a price table, it's a structural one: self-hosting's economics are dominated by utilization pattern, not token volume.
| GPU | Buy (new) | Buy (used) | Rent (typical) |
|---|---|---|---|
| H100 80GB | $25–30K | $12–18K | $2–3/hr (neo-cloud) · $7–12/hr (hyperscaler) |
| H200 | $30–40K | — | $3.72–10.60/hr |
| B200 / B300 | $30–50K DGX B300 8-GPU system: $300–350K | — | $5–18/hr launch premium |
| RTX 5090 | $3,000–4,000 | — | not typically rented |
| RTX 4090 | $1,600–2,200 | — | — |
| RTX 3090 24GB | — | $600–850 | — |
Hopper (H100/H200) is in a price-decline cycle now that Blackwell is shipping; neo-clouds/GPU marketplaces run 50–75% cheaper than hyperscalers for identical hardware.
Mac Studio's 2026 lineup is worse than it was, not better. The top current config is M3 Ultra at $5,299 with 96GB unified memory — Apple removed the 256GB and 512GB options in March/May 2026 as RAM prices spiked. An M5 Ultra is rumored but unreleased. A sourced benchmark on a (no-longer-purchasable) 512GB unit ran DeepSeek R1 671B at ~17–18 tok/s; at 800GB/s memory bandwidth, a 70B-class model at Q4 derives to a ceiling around 21 tok/s (~13–16 tok/s realistic single-stream) EST, not directly benchmarked. That's slower than a single RTX 4090 (36 tok/s sourced) on the same model class — modern GDDR-based consumer GPUs now out-bandwidth the M3 Ultra. The Mac Studio's edge is unified-memory capacity (it can load a huge model at all), not speed — a poor fit for the low-latency, interactive back-and-forth of a coding agent.
An 8×H100 node draws ~10–11kW under load. US commercial electricity averages 13.5¢/kWh (7–35¢ by state). Colocation for a standard 3–5kW rack runs $900–2,500/month; a high-density 10–30kW+ rack runs $3,000–6,000+/month — power availability, not floor space, is now the binding constraint at most colo providers.
Self-hosted cost-per-token depends almost entirely on how continuously the hardware is kept busy — and the difference between the two realistic cases is enormous:
Forge's own usage sits in Case A, not Case B. Forge's non-cache monthly volume (198M output + 44M fresh input = 242M tokens) smoothed evenly over a month is only ~93 tokens/second on average — roughly one GPU's single-stream capacity. But real usage isn't smooth: it's concentrated in active sessions with idle gaps between them. A dedicated GPU sized for forge's peak pace sits mostly idle the rest of the time, and idle GPU-hours are a sunk cost that a metered API never charges for. Reaching Case B's cheap-per-token economics needs sustained, batched, multi-tenant demand — pooling many concurrent agents/hearths onto shared always-warm infrastructure — which forge doesn't currently have and would need to build toward, not something a hardware purchase alone provides.
Working the numbers for a single dedicated GPU running a modest (not near-frontier) open model, sized for forge's "bulk" stages (implement/test/polish — the Sonnet/Haiku-eligible slice) rather than the hardest design/crypto work:
Honest answer to "at what volume does hardware beat API": for forge's actual usage pattern (bursty, effectively single-tenant), there isn't a clean volume threshold where owning hardware wins — the constraint is utilization, not scale. A dedicated GPU's fixed cost (~$1,000–1,800/month) is cheap enough that it plausibly pencils out against the Sonnet/Haiku tier specifically if quality holds up and the ops burden is accepted — but it does not touch the Opus tier, where the near-frontier open models that could substitute need 4–8 GPUs (~$200–400K capex), a scale that only makes sense with sustained batched demand forge doesn't have.
This is the part that keeps the analysis honest: none of the three boxes forge has today were bought for this, and it shows.
No GPU. Memory-bandwidth-bound at ~1–5 tok/s even on models that fit — and nothing in the near-frontier open-model table (Section 1) fits in 31GB even at 4-bit quantization. Ceiling is a small (<14B), heavily-quantized model at low single-digit tok/s.
The only one of the three that can load a serious open model at all. Best case (a max-config unit) runs a 70B-class model around 13–16 tok/s — usable for a single interactive session, but slower than a single consumer GPU on the same model class, and far short of API latency. 2026 note: current retail config tops out at 96GB, well below what's needed for the near-frontier 700B–1T-class models in Section 1.
Same constraint as the forge box, likely with less RAM and shared/lower-bandwidth CPU. Useful for orchestration, routing, or a tiny embedding/classifier model — not for serving a coding-capable LLM.
The honest bottom line: forge cannot self-host anything near-frontier today. Getting to the "single dedicated GPU for bulk work" scenario in Section 2 means buying a GPU box we don't have — a real capex decision, not a config change. Getting to a near-frontier open model (DeepSeek V4-Pro / Kimi K2.6 / GLM-5.2 class) means a multi-GPU cluster two orders of magnitude more expensive than that, and still trailing Opus by the ~11-point SWE-bench gap from Section 1.
MeshLLM turns out to be real and actively maintained — not vaporware, not a rumor. Worth naming: it comes from the same engineer at Block who built goose, the agent framework forge already spiked and shelved pending a graduation call.
Mesh-LLM/mesh-llm on GitHub, Apache-2.0, copyright Block Inc. It pools GPU/memory capacity across machines and exposes the combined result as a single OpenAI-compatible API endpoint. Nodes find each other via Nostr (public meshes) or invite tokens (private meshes); transport rides on iroh relay, which works through NAT/cellular/WiFi. Built on llama.cpp. One outside analysis frames it as "closer to BitTorrent than Bittensor" — protocol-first, tokenless, permissionless — a decentralized-inference thesis, not a training system, and not an agent-messaging metaphor despite the "mesh" name inviting that guess.
Actively released: 123 releases, latest v0.72.2 (Jul 2026). Core routing and an experimental "Mixture of Agents" mode both work today. Backends: CPU, CUDA, ROCm, Vulkan, and Metal (macOS — relevant to our Mac Studio). Multimodal support and mobile/desktop apps are roadmap items, not shipped.
Honest read for our specific hardware. The Mac Studio is the only strong node in this picture — real unified-memory bandwidth, and MeshLLM has a native Metal backend. The forge box (CPU-only) and the small VPS (no GPU) can technically join as CPU-backend peers for stage-splitting, but pipeline throughput is capped by its slowest stage, and CPU token generation is far slower per-layer than Metal/GPU compute. Putting any real fraction of a model's layers on the forge box or the VPS will likely make the whole pipeline crawl, even though the Mac Studio's own stage stays fast. This is a genuine capacity-for-speed trade: it could let forge load a bigger model than fits in the Mac Studio's own memory alone — at the cost of throughput dropping to low single-digit tokens/sec or worse EST — no MeshLLM-specific benchmark found for this pass. A remote VPS adds real network latency on top of that.
Verdict: don't expect MeshLLM to make the CPU-only forge box or the no-GPU VPS pull meaningful inference weight — they're slow enough to risk being a net drag on a shared pipeline rather than a net add. It is a legitimate, free, actively-maintained way to test whether pooling helps once (if) forge adds a real GPU box alongside the Mac Studio — worth a cheap experiment at that point, not a substitute for buying hardware now. Comparable projects worth knowing about if MeshLLM doesn't pan out: exo (Mac-cluster-focused), Petals, and llama.cpp's own native --rpc multi-host mode, which MeshLLM's stage-splitting likely builds on.
"Mixture of models" here means routing or ensembling across separately-trained models at the API/orchestration layer — distinct from mixture-of-experts (MoE), which is an architecture choice a model's trainer makes inside a single model's forward pass (why e.g. DeepSeek V4-Pro can have 1.6T total parameters but only 49B active per token). Routing across models is something forge could build; MoE is baked into the open models forge would be choosing between.
OpenRouter now moves an estimated ~20–25 trillion tokens/week (up from ~5T six months earlier), with its own "Auto Router" built on Not Diamond's meta-model. Chinese open-weight models (DeepSeek, MiniMax, Kimi, Qwen, Xiaomi) now account for 45–61%+ of OpenRouter's traffic — US frontier models' combined share has fallen to roughly 30%. Commercial routing products (Not Diamond, Martian) report 40–85% cost reductions in production enterprise deployments, and an estimated 37% of enterprises now run 5+ models in production.
The honest verdict: routing is a real cost lever, not a capability multiplier. RouterArena — the 2026 standard benchmark for comparing routers — explicitly finds that systems "lean heavily on expensive models for accuracy," and that cheaper routers "achieve competitive performance at much lower budgets, though they plateau earlier." That plateau is the frontier gap, made concrete: routing well within a pool of non-frontier models gets you close to that pool's own ceiling, not past it — hard queries still get sent to the frontier-tier model whenever one is in the pool. Separately, Mixture-of-Agents' headline claim of beating a frontier model using only open models is measured on AlpacaEval 2.0, MT-Bench, and FLASK — LLM-judged helpfulness/style benchmarks, not verifiable hard-reasoning or coding benchmarks. No evidence turned up of an open-model ensemble beating a frontier single model on SWE-bench- or GPQA-style tasks — the kind of hard design/crypto/coding work this analysis actually cares about.
For forge specifically: routing/cascading is already the right tool for the cost problem the prior pricing model solved (routing Opus/Sonnet/Haiku by lifecycle stage, saving 27–42%) — and the same logic extends cleanly to routing bulk-tier work to a self-hosted open model as one more (very cheap, or free-marginal-cost) rung on that ladder. It does not mean an ensemble of cheap open models can stand in for Opus on hard design, crypto, or novel-system work. That gate stays where it already is.
Putting Sections 1–5 together against forge's own numbers (198M output · 44M fresh-input · heavy cache-read, per the prior pricing model).
| Strategy | $/mo | New capex | Verdict |
|---|---|---|---|
| Pure API — all-Opus | $28,820 | $0 | Baseline. Frontier quality everywhere, zero ops burden, zero capex risk. |
| Pure API — aggressive routing prior model's recommendation | $16,716 | $0 | Current best option with zero new investment. Already bakes in Opus/Sonnet/Haiku routing by lifecycle stage. |
| Pure self-host attempt full replacement, all tiers | not viable | $200–400K+ | NOT RECOMMENDED Near-frontier open models need a multi-GPU cluster we don't own, still trail Opus by ~11 points on SWE-bench, and forge's bursty/low-tenant usage (Case A) means that cluster would mostly sit idle — worse $/token than API, not better, absent a fundamental change in how forge pools demand. |
| Hybrid — small pilot self-host the Haiku-eligible slice only | ~$15,500–16,700 ongoing | ~$15–30K one-time | Modest near-term savings (~$1,000/mo — the Haiku slice is only ~$1,150/mo of the current bill). The real return isn't the dollar figure; it's learning our real utilization pattern and validating quality on the lowest-risk work before betting bigger. |
| Hybrid — graduated aspirational, if the pilot proves out | ~$5,000–9,000 ongoing | ~$15–30K one-time | Opus slice stays untouched (~$4,300/mo for hard design/crypto). Self-hosted GPU takes over most of the Sonnet-eligible bulk tier (implement/test/polish). Depends entirely on the pilot validating quality, ops burden, and real utilization — not a starting assumption. |
Recommendation: stay on aggressive API routing now; run a small, cheap self-host pilot alongside it — don't bet on a big GPU build. Adopt (or confirm) the prior pricing model's aggressive-routing strategy first — it's the highest-confidence, zero-capex move already on the table. In parallel, if forge wants to learn whether self-hosting is real for this fleet, the right-sized experiment is one GPU (~$15–30K, buy don't rent past ~15 months of use), running a strong small open coding model (Qwen3.5-Coder-class), scoped to the lowest-stakes slice already routed to Haiku today (polish/docs, mechanical debug). That's cheap enough to lose without regret, concrete enough to teach us our real utilization pattern (is forge's usage really Case A, or would pooling multiple lanes/hearths push it toward Case B?), and low-stakes enough that a quality miss doesn't cost anything that matters. Graduating further — taking on Sonnet-eligible bulk work, or MeshLLM-pooling the GPU with the Mac Studio — is worth doing after that pilot reports back, not before. A multi-GPU near-frontier build is not worth it at forge's current scale and usage pattern, full stop.
The money/crypto gate does not move. Nothing here touches it. FROST, Lightning/ecash custody, settlement paths, and their adversarial verification stay on Opus/Fable regardless of which routing or self-host strategy forge adopts — that's a floor, not a dial, in every scenario on this page, same as the prior pricing model.