Mac Mini AI Subfactory

Sovereign Compute Console · qwen3:32b · deepseek-r1:32b/70b · nomic-embed-text

Concept console · all data simulated 4 / 4 nodes modelled --:--:--
The Hardware
Apple Mac mini — one node of the AI Subfactory
Apple Mac mini (M4), front view — official Apple product image
Apple M4 Pro 64 GB unified memory 20-core GPU 16-core Neural Engine Thunderbolt 5
Product image © Apple Inc.
Inside the Mini — silicon at work
Animated cutaway (cover off) · driven by the simulated telemetry below · component layout representative, not exact
Simulation
SoC PACKAGE M4 PRO P-CORES E-CORES GPU · 20-CORE NEURAL ENGINE SLC CACHE UNIFIED MEMORY LPDDR5X 32 GB LPDDR5X 32 GB NVMe SSD TB5 TB5 TB5 RDMA → CLUSTER FAN 1200 rpm
Memory read Memory write Thunderbolt · RDMA NVMe I/O
Modelled savings vs. cloud API pricing
$0
▲ modelled accumulation · not a measured figure
Loading today's win…
Data Sovereignty
100%
0 bytes left the building today
Tokens Processed Today
0
Avg Time-to-First-Token
— ms
vs ~850ms typical cloud RTT
Cluster Uptime
99.97%
24 / 7 dedicated agent box
Rate Limits Hit
0
unlimited local throughput
Hardware Payback
—%
of $8,000 cluster recouped

Apple Silicon Telemetry — simulated stream, cluster-wide

Throughput per node (tok/s)
4 nodes · rolling 60s window
Timeqwen3:32bdeepseek-r1:32bdeepseek-r1:70bnomic-embed
Compute utilization
Cluster average · GPU vs Neural Engine
TimeGPU %ANE %
Unified memory bandwidth
DRAM read / write · GB/s, cluster sum
TimeRead GB/sWrite GB/s
Power & thermal envelope
Separate scales by design — never dual-axis
Power draw (W)
Peak node temperature (°C)

Cluster Topology

Why Local — the argument, made visible

Cost per 1M tokens
Local marginal cost vs. blended cloud API rate
Response latency
On-device inference vs. typical cloud round trip
Sovereignty rate
100%
all inference on-devicetarget 100%
Hardware payback progress
—%
$8,000 cluster costbreakeven ≈ 18–24 mo
Concept console. Every number on this page is simulated; none of it is live telemetry from the subfactory. The panels are shaped for the real feed (per-node tokens/sec, GPU and Neural Engine utilisation, memory bandwidth, power and temperature per node, streamed from the machines themselves). When that collector ships it will appear in the build log first, and this caption will change. Cost, latency and payback figures are planning assumptions, not measurements.