◂ FORGE
Hardware & infra · extends the pricing model · not deployed

Self-Hosting Models vs. API Keys

Weighing hardware + infra investment against Anthropic API rates, and whether a hybrid split beats either extreme. Prepared by keeper:hearth, generated 2026-07-08.

Executive summary

The short version: self-hosting doesn't beat the API at forge's current scale and usage pattern — the constraint is utilization, not price. A small, cheap pilot is worth running to learn whether that changes; a big hardware bet is not worth making yet.

Cheapest option today — API, aggressive routing
$16,716/mo
Zero new capex, zero new ops burden — the prior pricing model's recommendation, unchanged by this analysis
Pure self-host, forge's actual usage pattern
$14–83/M tok
Worse than every API tier — bursty, low-tenant usage is the expensive utilization case, not the cheap one
Recommended move
~$15–30K one-time
One GPU, one narrow low-stakes slice (Haiku-eligible work) — a learning pilot, not a cost-cut project

1. Open models near frontier (mid-2026)

The open-weight field has closed most of the gap on routine coding and reasoning. It has not closed the gap on the hardest, most agentic single-benchmark work. Numbers below are cross-cited across several 2026 model-comparison sites (BenchLM.ai, aimadetools, MindStudio-style aggregators) rather than primary release papers in every case — treat exact percentages as EST directionally-right, not precise. Artificial Analysis's blended index is the more credible cross-check where cited.

ModelSizeLicenseNotable benchmarks
DeepSeek V4-Pro Apr 2026 1.6T MoE / 49B active MIT SWE-bench Verified 80.6%, LiveCodeBench 93.5%, Codeforces 3206. Cheapest frontier-adjacent coding model via API ($0.44/$0.87 per M tok) — the benchmark it sets for what self-hosting is competing against.
GLM-5.2 (Zhipu) Jun 2026 ~744B MoE / 40B active MIT 1M context, DeepSeek Sparse Attention. Leads SWE-Bench Pro / Terminal-Bench 2.1 / FrontierSWE among open models; predecessor GLM-5.1 hit SWE-Bench Pro 58.4 vs. Opus 4.6's 57.3 — a rare case of an open model edging a recent Opus point release on one benchmark.
Kimi K2.6 (Moonshot) 1T MoE / 32B active Open 256K context, MLA attention + vision encoder. Best-regarded open model for long-horizon/autonomous coding; co-leads the open field on Artificial Analysis's blended Intelligence Index (~54).
Qwen3.5 / 3.6 (Alibaba) 0.8B–397B-A17B flagship Apache 2.0 Most permissive license of the group. Qwen3.5-Coder-72B: LiveCodeBench 61.8% vs. GPT-5's 64.2% — genuinely close on this one. Flagship: GPQA-D 88.4, AIME 91.3. Notably, the 27B dense Qwen3.6 variant now beats Alibaba's own 397B MoE flagship on coding — size isn't a clean proxy for capability anymore.
Llama 4 (Meta) Apr 2025, still current — no Llama 5 Scout 109B/17B active; Maverick 128 experts Custom, EU-restricted Highest open MMLU (85.5%), but a signal worth naming: Meta's own proprietary Muse Spark model scores 52 on the AA Index vs. Maverick's 18 — Meta itself appears to be de-prioritizing its open line as the frontier path.
Mistral Large 3 675B MoE / 41B active Open-weight Strong HumanEval (~92%) but weak GPQA-D (~44%) — a non-reasoning model, clearly behind specifically on hard reasoning.

Hardware required to self-host EST

PrecisionVRAM/RAM neededWhat that looks like
Full FP16~2TBA 1T-class model at full precision needs a multi-node GPU cluster — out of reach for any single-box setup, ours included.
INT4 / Q4 the realistic self-host target~500–600GB combinedE.g. Kimi K2.6 at Q4 — needs 8×H200, 4×A100-80GB, or an 8×RTX-4090 aggregate. Q2 gets to ~350GB with real quality loss.
CPU-only, no GPU matches our forge box: 12-core / 31GB RAM1–5 tok/sMemory-bandwidth-bound (~100GB/s DDR → a theoretical floor around 2.8s/token for an unquantized 70B dense model). A 31GB box can't even load any model in the table above at 4-bit. Ceiling is a small (<14B), heavily-quantized model at low single-digit tok/s. Not viable for real work.
Mac Studio, 512GB unified memory16–20 tok/sThe realistic prosumer ceiling. A DeepSeek-V3-class 671B MoE model at 4-bit runs at this rate via mlx-lm/llama.cpp — usable, but far slower than API latency. AVAILABILITY Apple reportedly pulled the 512GB config in 2026 amid a DRAM price spike; no M4/M5 Ultra exists yet, so this ceiling may not be purchasable new right now.

The honest gap, in two numbers. On Artificial Analysis's blended Intelligence Index (Jul 2026), Opus 4.8 leads at 55.7% with the best open models (Kimi K2.6, Xiaomi MiMo V2.5) around 54 and DeepSeek V4 Pro around 52 — genuinely close, a 3–6 point gap on a broad blended measure. But on SWE-bench Verified specifically — the single benchmark closest to what forge actually does — Opus 4.8's 88.6% vs. the best open score around 77.8% is an ~11-point gap, and it's concentrated exactly where it matters: long-horizon autonomous multi-file work requiring self-correction with minimal supervision, plus multimodal reasoning and long tool-use chains. Open models have gotten good at routine coding and reasoning. They have not closed the gap on the hardest agentic work forge actually leans on Opus for.

2. Hardware economics: capex + opex vs. API

All figures EST unless noted — sourced from 2026 GPU/cloud pricing pages, cross-checked across several vendors. The headline finding isn't a price table, it's a structural one: self-hosting's economics are dominated by utilization pattern, not token volume.

GPU costs, 2026

GPUBuy (new)Buy (used)Rent (typical)
H100 80GB$25–30K$12–18K$2–3/hr (neo-cloud) · $7–12/hr (hyperscaler)
H200$30–40K$3.72–10.60/hr
B200 / B300$30–50K DGX B300 8-GPU system: $300–350K$5–18/hr launch premium
RTX 5090$3,000–4,000not typically rented
RTX 4090$1,600–2,200
RTX 3090 24GB$600–850

Hopper (H100/H200) is in a price-decline cycle now that Blackwell is shipping; neo-clouds/GPU marketplaces run 50–75% cheaper than hyperscalers for identical hardware.

Mac Studio's 2026 lineup is worse than it was, not better. The top current config is M3 Ultra at $5,299 with 96GB unified memory — Apple removed the 256GB and 512GB options in March/May 2026 as RAM prices spiked. An M5 Ultra is rumored but unreleased. A sourced benchmark on a (no-longer-purchasable) 512GB unit ran DeepSeek R1 671B at ~17–18 tok/s; at 800GB/s memory bandwidth, a 70B-class model at Q4 derives to a ceiling around 21 tok/s (~13–16 tok/s realistic single-stream) EST, not directly benchmarked. That's slower than a single RTX 4090 (36 tok/s sourced) on the same model class — modern GDDR-based consumer GPUs now out-bandwidth the M3 Ultra. The Mac Studio's edge is unified-memory capacity (it can load a huge model at all), not speed — a poor fit for the low-latency, interactive back-and-forth of a coding agent.

Power and hosting

An 8×H100 node draws ~10–11kW under load. US commercial electricity averages 13.5¢/kWh (7–35¢ by state). Colocation for a standard 3–5kW rack runs $900–2,500/month; a high-density 10–30kW+ rack runs $3,000–6,000+/month — power availability, not floor space, is now the binding constraint at most colo providers.

The real question: which utilization case are we in?

Self-hosted cost-per-token depends almost entirely on how continuously the hardware is kept busy — and the difference between the two realistic cases is enormous:

Case A — bursty, single-stream MATCHES FORGE
$14–83/M tok
1 rented H100 @ ~$2.50/hr, ~50 tok/s single-stream. $14/M if you could cleanly power down between sessions (you can't — cold-start/model-load costs make that impractical for an active agent); realistically billed ~24/7 for the same volume → $83/M. EST
Case B — saturated, batched serving
~$0.35/M tok
Same H100, continuous-batching many concurrent requests (vLLM-style) — ~2,000 tok/s aggregate, 24/7. Far cheaper than API — if you have enough concurrent demand to fill that throughput. EST

Forge's own usage sits in Case A, not Case B. Forge's non-cache monthly volume (198M output + 44M fresh input = 242M tokens) smoothed evenly over a month is only ~93 tokens/second on average — roughly one GPU's single-stream capacity. But real usage isn't smooth: it's concentrated in active sessions with idle gaps between them. A dedicated GPU sized for forge's peak pace sits mostly idle the rest of the time, and idle GPU-hours are a sunk cost that a metered API never charges for. Reaching Case B's cheap-per-token economics needs sustained, batched, multi-tenant demand — pooling many concurrent agents/hearths onto shared always-warm infrastructure — which forge doesn't currently have and would need to build toward, not something a hardware purchase alone provides.

Is there a volume-based break-even?

Working the numbers for a single dedicated GPU running a modest (not near-frontier) open model, sized for forge's "bulk" stages (implement/test/polish — the Sonnet/Haiku-eligible slice) rather than the hardest design/crypto work:

  • Fixed monthly cost, one GPU: buying an H100 (~$27K new, amortized 30 months) + power/overhead (~$100–150/mo) ≈ $1,000–1,050/month, roughly flat regardless of how much it's actually used. Renting the equivalent (~$2.50/hr × 24 × 30) runs ~$1,800/month — buying wins past ~15–18 months of continuous use. EST
  • What that buys, in dollars saved: in the aggressive routing strategy from the prior pricing model, Sonnet+Haiku together already cost an estimated ~$12,400/month (65% Sonnet + 20% Haiku of the $16,716/mo total, backed out from the 0.6×/0.2× rate ratios). If a single GPU's open model could fully substitute for that tier at acceptable quality, the theoretical ceiling is ~$11,000–11,400/month saved — a big number, but a best-case ceiling, not a plan.
  • Why it isn't that clean: (a) quality — a single-GPU-class open model (70B-ish dense, or a small MoE) is not a drop-in Sonnet replacement across all of implement/test/polish, especially on anything requiring multi-file reasoning; (b) ops burden — uptime, monitoring, request queuing, and model upgrades are new work forge doesn't do today; (c) none of forge's current hardware can run this — it requires new capex forge doesn't have (see Section 3).

Honest answer to "at what volume does hardware beat API": for forge's actual usage pattern (bursty, effectively single-tenant), there isn't a clean volume threshold where owning hardware wins — the constraint is utilization, not scale. A dedicated GPU's fixed cost (~$1,000–1,800/month) is cheap enough that it plausibly pencils out against the Sonnet/Haiku tier specifically if quality holds up and the ops burden is accepted — but it does not touch the Opus tier, where the near-frontier open models that could substitute need 4–8 GPUs (~$200–400K capex), a scale that only makes sense with sustained batched demand forge doesn't have.

3. What our actual hardware can do

This is the part that keeps the analysis honest: none of the three boxes forge has today were bought for this, and it shows.

Forge box 31GB RAM · 12-core · CPU-only

No GPU. Memory-bandwidth-bound at ~1–5 tok/s even on models that fit — and nothing in the near-frontier open-model table (Section 1) fits in 31GB even at 4-bit quantization. Ceiling is a small (<14B), heavily-quantized model at low single-digit tok/s.

Not viable for real work

Mac Studio unified memory

The only one of the three that can load a serious open model at all. Best case (a max-config unit) runs a 70B-class model around 13–16 tok/s — usable for a single interactive session, but slower than a single consumer GPU on the same model class, and far short of API latency. 2026 note: current retail config tops out at 96GB, well below what's needed for the near-frontier 700B–1T-class models in Section 1.

Only real candidate — still slow, still capacity-limited

Small Ubuntu VPS no GPU

Same constraint as the forge box, likely with less RAM and shared/lower-bandwidth CPU. Useful for orchestration, routing, or a tiny embedding/classifier model — not for serving a coding-capable LLM.

Not viable for model serving

The honest bottom line: forge cannot self-host anything near-frontier today. Getting to the "single dedicated GPU for bulk work" scenario in Section 2 means buying a GPU box we don't have — a real capex decision, not a config change. Getting to a near-frontier open model (DeepSeek V4-Pro / Kimi K2.6 / GLM-5.2 class) means a multi-GPU cluster two orders of magnitude more expensive than that, and still trailing Opus by the ~11-point SWE-bench gap from Section 1.

4. MeshLLM, idle compute, and combining our boxes

MeshLLM turns out to be real and actively maintained — not vaporware, not a rumor. Worth naming: it comes from the same engineer at Block who built goose, the agent framework forge already spiked and shelved pending a graduation call.

What it is CONFIRMED

Mesh-LLM/mesh-llm on GitHub, Apache-2.0, copyright Block Inc. It pools GPU/memory capacity across machines and exposes the combined result as a single OpenAI-compatible API endpoint. Nodes find each other via Nostr (public meshes) or invite tokens (private meshes); transport rides on iroh relay, which works through NAT/cellular/WiFi. Built on llama.cpp. One outside analysis frames it as "closer to BitTorrent than Bittensor" — protocol-first, tokenless, permissionless — a decentralized-inference thesis, not a training system, and not an agent-messaging metaphor despite the "mesh" name inviting that guess.

Maturity CONFIRMED

Actively released: 123 releases, latest v0.72.2 (Jul 2026). Core routing and an experimental "Mixture of Agents" mode both work today. Backends: CPU, CUDA, ROCm, Vulkan, and Metal (macOS — relevant to our Mac Studio). Multimodal support and mobile/desktop apps are roadmap items, not shipped.

Two distinct modes

  • Federation/routing — each node independently serves whatever model already fits on it; a router sends each request to a capable peer by model/complexity. This is closer to load-balancing than to pooling capacity.
  • Cross-node model splitting — for a model too big for any single node, MeshLLM loads it as contiguous "layer stages" spread across machines. This is genuine pipeline parallelism (splitting a model's layers across machines, not splitting individual matrix operations the way tensor-parallelism does) — inherently more tolerant of slow/high-latency links, since activations only cross machines at stage boundaries, not on every operation.

Honest read for our specific hardware. The Mac Studio is the only strong node in this picture — real unified-memory bandwidth, and MeshLLM has a native Metal backend. The forge box (CPU-only) and the small VPS (no GPU) can technically join as CPU-backend peers for stage-splitting, but pipeline throughput is capped by its slowest stage, and CPU token generation is far slower per-layer than Metal/GPU compute. Putting any real fraction of a model's layers on the forge box or the VPS will likely make the whole pipeline crawl, even though the Mac Studio's own stage stays fast. This is a genuine capacity-for-speed trade: it could let forge load a bigger model than fits in the Mac Studio's own memory alone — at the cost of throughput dropping to low single-digit tokens/sec or worse EST — no MeshLLM-specific benchmark found for this pass. A remote VPS adds real network latency on top of that.

Verdict: don't expect MeshLLM to make the CPU-only forge box or the no-GPU VPS pull meaningful inference weight — they're slow enough to risk being a net drag on a shared pipeline rather than a net add. It is a legitimate, free, actively-maintained way to test whether pooling helps once (if) forge adds a real GPU box alongside the Mac Studio — worth a cheap experiment at that point, not a substitute for buying hardware now. Comparable projects worth knowing about if MeshLLM doesn't pan out: exo (Mac-cluster-focused), Petals, and llama.cpp's own native --rpc multi-host mode, which MeshLLM's stage-splitting likely builds on.

5. Mixture-of-models: does routing close the gap?

"Mixture of models" here means routing or ensembling across separately-trained models at the API/orchestration layer — distinct from mixture-of-experts (MoE), which is an architecture choice a model's trainer makes inside a single model's forward pass (why e.g. DeepSeek V4-Pro can have 1.6T total parameters but only 49B active per token). Routing across models is something forge could build; MoE is baked into the open models forge would be choosing between.

The technique

  • Routing — classify a query, send it to the cheapest model predicted to be good enough. RouteLLM (the academic baseline) hits ~95% of a strong model's quality while routing only 14–26% of queries to it — a 75–85% cost cut.
  • Cascades — try the cheap model first, escalate to the expensive one only on a failed confidence/verification check. FrugalGPT reports up to 98% cost reduction at equal/better accuracy on its own benchmark suite.
  • Ensembling (Mixture-of-Agents) — several models answer in parallel or in layered rounds; one model (often the strongest) synthesizes the final answer.

Current landscape (2026)

OpenRouter now moves an estimated ~20–25 trillion tokens/week (up from ~5T six months earlier), with its own "Auto Router" built on Not Diamond's meta-model. Chinese open-weight models (DeepSeek, MiniMax, Kimi, Qwen, Xiaomi) now account for 45–61%+ of OpenRouter's traffic — US frontier models' combined share has fallen to roughly 30%. Commercial routing products (Not Diamond, Martian) report 40–85% cost reductions in production enterprise deployments, and an estimated 37% of enterprises now run 5+ models in production.

The honest verdict: routing is a real cost lever, not a capability multiplier. RouterArena — the 2026 standard benchmark for comparing routers — explicitly finds that systems "lean heavily on expensive models for accuracy," and that cheaper routers "achieve competitive performance at much lower budgets, though they plateau earlier." That plateau is the frontier gap, made concrete: routing well within a pool of non-frontier models gets you close to that pool's own ceiling, not past it — hard queries still get sent to the frontier-tier model whenever one is in the pool. Separately, Mixture-of-Agents' headline claim of beating a frontier model using only open models is measured on AlpacaEval 2.0, MT-Bench, and FLASK — LLM-judged helpfulness/style benchmarks, not verifiable hard-reasoning or coding benchmarks. No evidence turned up of an open-model ensemble beating a frontier single model on SWE-bench- or GPQA-style tasks — the kind of hard design/crypto/coding work this analysis actually cares about.

For forge specifically: routing/cascading is already the right tool for the cost problem the prior pricing model solved (routing Opus/Sonnet/Haiku by lifecycle stage, saving 27–42%) — and the same logic extends cleanly to routing bulk-tier work to a self-hosted open model as one more (very cheap, or free-marginal-cost) rung on that ladder. It does not mean an ensemble of cheap open models can stand in for Opus on hard design, crypto, or novel-system work. That gate stays where it already is.

6. The balance: pure-API, pure-self-host, or hybrid

Putting Sections 1–5 together against forge's own numbers (198M output · 44M fresh-input · heavy cache-read, per the prior pricing model).

Strategy$/moNew capexVerdict
Pure API — all-Opus$28,820$0 Baseline. Frontier quality everywhere, zero ops burden, zero capex risk.
Pure API — aggressive routing prior model's recommendation$16,716$0 Current best option with zero new investment. Already bakes in Opus/Sonnet/Haiku routing by lifecycle stage.
Pure self-host attempt full replacement, all tiersnot viable$200–400K+ NOT RECOMMENDED Near-frontier open models need a multi-GPU cluster we don't own, still trail Opus by ~11 points on SWE-bench, and forge's bursty/low-tenant usage (Case A) means that cluster would mostly sit idle — worse $/token than API, not better, absent a fundamental change in how forge pools demand.
Hybrid — small pilot self-host the Haiku-eligible slice only~$15,500–16,700 ongoing~$15–30K one-time Modest near-term savings (~$1,000/mo — the Haiku slice is only ~$1,150/mo of the current bill). The real return isn't the dollar figure; it's learning our real utilization pattern and validating quality on the lowest-risk work before betting bigger.
Hybrid — graduated aspirational, if the pilot proves out~$5,000–9,000 ongoing~$15–30K one-time Opus slice stays untouched (~$4,300/mo for hard design/crypto). Self-hosted GPU takes over most of the Sonnet-eligible bulk tier (implement/test/polish). Depends entirely on the pilot validating quality, ops burden, and real utilization — not a starting assumption.

Recommendation: stay on aggressive API routing now; run a small, cheap self-host pilot alongside it — don't bet on a big GPU build. Adopt (or confirm) the prior pricing model's aggressive-routing strategy first — it's the highest-confidence, zero-capex move already on the table. In parallel, if forge wants to learn whether self-hosting is real for this fleet, the right-sized experiment is one GPU (~$15–30K, buy don't rent past ~15 months of use), running a strong small open coding model (Qwen3.5-Coder-class), scoped to the lowest-stakes slice already routed to Haiku today (polish/docs, mechanical debug). That's cheap enough to lose without regret, concrete enough to teach us our real utilization pattern (is forge's usage really Case A, or would pooling multiple lanes/hearths push it toward Case B?), and low-stakes enough that a quality miss doesn't cost anything that matters. Graduating further — taking on Sonnet-eligible bulk work, or MeshLLM-pooling the GPU with the Mac Studio — is worth doing after that pilot reports back, not before. A multi-GPU near-frontier build is not worth it at forge's current scale and usage pattern, full stop.

The money/crypto gate does not move. Nothing here touches it. FROST, Lightning/ecash custody, settlement paths, and their adversarial verification stay on Opus/Fable regardless of which routing or self-host strategy forge adopts — that's a floor, not a dial, in every scenario on this page, same as the prior pricing model.

7. Assumptions & caveats — read before trusting the numbers

  • This is a research synthesis, not measured production data. Every dollar figure in Sections 2–6 is a model built on 2026 web-sourced pricing and benchmark data, not forge's own measured spend on hardware forge doesn't yet own. Treat all of it as directional.
  • Open-model benchmark precision is uncertain EST. Most sources for Section 1's model comparisons are 2026 aggregator/comparison sites, not primary release papers — model names and version numbers (V4 vs V4-Pro, 5.1 vs 5.2) are moving fast and cross-cited among these sites. Artificial Analysis's blended index was the more credible cross-check where cited, and is what the "~11-point SWE-bench gap" finding leans on most.
  • Mac Studio and self-hosted throughput figures are largely derived, not directly benchmarked EST. The 671B DeepSeek-R1 benchmark (~17–18 tok/s) is sourced; the 70B-class Mac Studio estimate and the MeshLLM cross-node throughput are both extrapolated from hardware specs and general pipeline-parallel behavior, not measured on forge's actual hardware.
  • GPU rental/purchase prices move fast and vary by provider/region. The ranges in Section 2 are a 2026 snapshot; Hopper-class GPU prices specifically are in a decline cycle as Blackwell ships, so re-check before committing capex.
  • Quality comparisons are all on public benchmarks, not on forge's actual workload. Nothing in this analysis tests an open model against a real forge task (a real PR, a real debug session, a real crypto review). The pilot recommended in Section 6 exists partly to generate that missing data point.
  • Engineer/ops time isn't priced in. Running self-hosted inference — uptime, monitoring, model updates, request queuing across concurrent lanes — is real ongoing labor forge doesn't currently do. On the small pilot specifically, that labor cost could easily exceed the ~$1,000/month in raw savings; the pilot's value is learning, not near-term dollars.
  • Cache economics don't automatically transfer. Anthropic's cache-read discount (~0.1×) assumes a specific caching implementation; a self-hosted stack would need an equivalent (e.g. vLLM automatic prefix caching, SGLang RadixAttention) to get comparable reuse economics on repeated context. This isn't modeled in detail here — the throughput figures in Section 2 focus on fresh generation (output + fresh-input tokens), not the full 31.4B-token cache-inclusive volume.
  • The Case A / Case B utilization framing is the single most load-bearing finding in this analysis. Every self-host dollar figure depends on which case applies, and forge's own usage pattern (bursty, low-tenant) sits in the expensive one. If that pattern changes — more concurrent lanes, pooled demand across hearths — the self-host economics improve substantially and this analysis should be revisited.