Skip to content
Main Navigation Puget Systems Logo in White and Green
  • Solutions
    • Media & Entertainment
      • Photo Editing
        • Recommended Systems For:
        • Adobe Lightroom Classic
        • Adobe Photoshop
        • Generative AI
      • Video Editing & Motion Graphics
        • Recommended Systems For:
        • Adobe After Effects
        • Adobe Premiere Pro
        • DaVinci Resolve
        • Foundry Nuke
      • 3D Design & Animation
        • Recommended Systems For:
        • Autodesk 3ds Max
        • Autodesk Maya
        • Blender
        • Cinema 4D
        • Houdini
        • ZBrush
      • Live Video Production
        • Recommended Systems For:
        • vMix
        • Live Streaming
      • Real-Time Engines
        • Recommended Systems For:
        • Game Development
        • Unity
        • Unreal Engine
        • Virtual Production
      • Rendering
        • Recommended Systems For:
        • Keyshot
        • OctaneRender
        • Redshift
        • V-Ray
      • Digital Audio
        • Recommended Systems For:
        • Ableton Live
        • FL Studio
        • Pro Tools
    • Engineering
      • Architecture & CAD
        • Recommended Systems For:
        • Autodesk AutoCAD
        • Autodesk Inventor
        • Autodesk Revit
        • SOLIDWORKS
      • Visualization
        • Recommended Systems For:
        • Enscape
        • Keyshot
        • Lumion
        • Twinmotion
      • Photogrammetry & GIS
        • Recommended Systems For:
        • ArcGIS Pro
        • Agisoft Metashape
        • PIX4D
        • RealityScan
    • AI & HPC
      • AI Development & Deployment
        • Recommended Systems For:
        • AI Development
        • AI Deployment & Inference
        • Servers for Scaling AI & LLMs
      • High Performance Computing
        • Recommended Systems For:
        • Data Science
        • Scientific Computing
    • More
      • Recommended Systems For:
      • Compact Size
      • NVIDIA RTX Studio
      • Virtual Reality
    • Business & Enterprise
      We can empower your company
    • Government & Education
      Services tailored for your organization
  • Products
    • Puget Mobile
      Powerful laptop workstations
      • Puget Mobile 16″
    • Puget Workstations
      High-performance Desktop PCs
      • AMD Ryzen
        Powerful CPUs with up to 16 cores
      • AMD Threadripper
        High core counts and lots of PCIe lanes
      • AMD EPYC
        Server-class CPUs in a workstation
      • Intel Core Ultra
        Balanced single- and multi-core performance
      • Intel Xeon
        Workstation CPUs with AVX512
      • Configure a Custom PC Workstation
        Configure a PC for your workflow
    • Puget Rackstations
      Workstations in rackmount chassis
      • AMD
        Ryzen, Threadripper, and EPYC CPUs
      • Intel
        Core Ultra and Xeon Processors
      • Configure a Custom Rackmount Workstation
        Tailored 4U, 5U, and 6U rack systems
    • Puget Servers
      Enterprise-class rackmount servers
      • 1U Rackmount
        Dense CPU compute servers
      • 2U Rackmount
        Mixed CPU and GPU solutions
      • 4U Rackmount
        High-density GPU computing
      • Multi-Node Servers
        2-4 servers in a single chassis
      • Liquid-Cooled Servers
        Powered by Comino
      • Custom Servers
        Engineered to meet your unique needs
    • Puget Storage
      Solutions from desktop to datacenter
      • Network-Attached Storage
        Synology desktop and rackmount NAS
      • Software-Defined Storage
        Datacenter solutions with QuantaStor
    • Recommended Third Party Peripherals
      Curated list of accessories for your workstation
    • Puget Bench for Creators
      Professional benchmarking tools
  • Publications
    • Articles
    • Blog Posts
    • Case Studies
    • HPC Blog
    • Podcasts
    • Press
  • Support
    • Contact Support
    • Onsite Services
    • Support Articles
    • Unboxing
    • Warranty Details
  • About Us
    • About Us
    • Careers
    • Contact Us
    • Enterprise
    • Gov & Edu
    • Our Customers
    • Press Kit
    • Puget Gear
    • Testimonials
  • Talk to an Expert
  • My Account
  1. Home
  2. /
  3. Hardware Articles
  4. /
  5. AMD Radeon AI PRO R9700: Dual-GPU AI Inference Performance

AMD Radeon AI PRO R9700: Dual-GPU AI Inference Performance

Posted on August 18, 2026 (September 9, 2026) by Dustin Moore | Last updated: September 9, 2026

Table of Contents

  • Introduction
  • Test Setup
  • What Models Fit?
  • Single-GPU Performance
  • Dual-GPU Performance
  • Scaling Summary
  • Cost of Inference: Local vs. Cloud
  • Image Generation: ComfyUI + Z-Image Turbo
  • Comparison: R9700 vs. Arc Pro B70
  • What Doesn’t Work (and Deployment Considerations)
  • How Do Two R9700 GPUs Perform for AI Inference?
  • Appendix: Setup Guide for Practitioners

How do two AMD Radeon™ AI PRO R9700 GPUs perform for local LLM inference and image generation?

Article Update & Retraction (September 2026): Multi-GPU vLLM on Bare Metal

Following community feedback, we re-tested dual AMD Radeon™ AI PRO R9700 GPUs on a bare-metal host. We confirmed that stock vLLM Tensor Parallelism (TP=2) functions reliably on RDNA 4 silicon. Our initial finding that stock vLLM multi-GPU failed was caused by a collective initialization bootstrap issue specific to our KVM/VFIO virtualization passthrough test environment, not hardware or upstream software defects.

  • Multi-GPU Working Baseline: Stock vLLM (vllm/vllm-openai-rocm:v0.20.2) in FP8 (TP=2) serves Qwen3.6-27B at 15.88 tok/s single-user and scales to 156.21 tok/s at 8 concurrent users.
  • MTP High-Water Mark: Enabling Multi-Token Prediction (MTP speculative decoding) pushes dual-GPU throughput to 62.75 tok/s single-user (15.1 ms ITL) and 320.23 tok/s under 8-user concurrency with bit-identical output.
  • Workstation Economics: With working dual-GPU serving, serving a 27B model on dual R9700s clears break-even against economy cloud models (e.g., GPT-5.6 Luna) in ~95 hours/week of load, down from “never” under our previous layer-split baseline.

The article below has been updated throughout to reflect these verified bare-metal benchmarks and provide updated deployment guidance.

Introduction

In our Intel Arc™ Pro B70 article, we explored what a VRAM-first, multi-GPU inference workstation looks like when built around Intel’s 32 GB cards. This article asks the same question of AMD’s entry in that fight.

The AMD Radeon™ AI PRO R9700 is the first RDNA 4 professional card positioned squarely at local AI inference. With 32 GB of GDDR6 VRAM per card and 640 GB/s of memory bandwidth, it occupies the same VRAM tier as the Arc Pro B70, paired with higher memory bandwidth and AMD’s mature ROCm software stack backing it. At roughly $1,880 (Puget Systems pricing as of July 2026; AMD’s MSRP is $1,299), the R9700 substantially undercuts NVIDIA’s GeForce RTX™ 5090 (~$4,130) while matching it on raw VRAM capacity. Two R9700 cards deliver 64 GB of aggregate VRAM — the same total capacity as two RTX 5090s, at less than half the cost.

In light of that we set out to answer these questions:

  • Can two R9700 cards deliver production-quality local LLM inference and image generation?
  • And how does AMD’s approach stack up against the four-card B70 configuration we already tested?
Dual AMD GPUs Text Over Blue Shaded Background with Computer Monitor

We installed two R9700 cards in a Puget Systems workstation and tested single-GPU baselines, multi-GPU serving across two cards, generative image workloads, and a 27B-parameter model utilizing Tensor Parallelism (TP=2). In doing so, we also uncovered important deployment nuances between bare-metal environments and virtualized passthrough setups. We also measured GPU power draw during each benchmark to calculate real-world cost-per-token and compared it against the full range of today’s cloud APIs, where output pricing spans $1.20/1M tokens (GPT-5.6 Luna) to $30/1M tokens (GPT-5.6 Sol).

Test Setup

ComponentSpec
GPUs2× AMD Radeon™ AI PRO R9700 (RDNA 4 / gfx1201)
VRAM per GPU32 GB GDDR6
Total VRAM64 GB
Compute Units64 per GPU
AI Accelerators128 per GPU
AI Performance766 TOPS (INT8) / 1531 TOPS (INT4) per GPU
Memory Bandwidth640 GB/s per GPU
Infinity Cache64 MB per GPU
Total Board Power300W per GPU
Host SystemPuget Systems workstation: AMD Ryzen™ Threadripper™ PRO 5995WX (64 cores / 128 threads), 128 GB RAM, ASUS Pro WS WRX80E-SAGE SE
Host OSRocky Linux 10.1 (Bare Metal host; cards bound to amdgpu at boot)
Test EnvironmentsBenchmarks were conducted on the bare-metal host with containerized runtimes (Podman / Docker) with PCIe peer-to-peer enabled. Virtualization passthrough testing was also conducted inside a KVM guest (Ubuntu 26.04) to analyze VM overhead and collective initialization behavior.
PCIePCIe 5.0 x16 for each GPU (2 slots used in total)

Inference Software: Single-GPU LLM benchmarks used vllm/vllm-openai-rocm:v0.20.2 with unquantized FP16 weights and HIP graphs enabled. Multi-GPU inference was evaluated across stock vLLM Tensor Parallelism (vllm/vllm-openai-rocm:v0.20.2 in FP8, TP=2), tuned builds with Multi-Token Prediction (MTP), and llama.cpp’s ROCm backend (GGUF layer-split). Image generation used ComfyUI with rocm/pytorch:latest and the Z-Image Turbo model.

Benchmark Tool: We drove every LLM test with NVIDIA GenAI-Perf in --streaming mode. Standard single-GPU runs used 500 input and 500 output tokens. Multi-GPU and 27B runs were evaluated across standardized sweeps (500 input / 200 output tokens) at concurrency levels 1, 4, and 8 simultaneous users over 50 prompts, discarding 3 warmup requests per run to eliminate cold-start drag. Reasoning models (Qwen3 8B) utilized an extended 120-second window to accommodate thinking tokens.

Power Monitoring: GPU power draw was measured throughout each benchmark by polling the AMD kernel driver’s sysfs hwmon interface (power1_average, reported in microwatts) every 2 seconds. This captures real power consumption under inference load, not just total board power (TBP) ratings.

What Models Fit?

Performance numbers only mean something once you know which models the hardware can actually hold. The ROCm-native vLLM container supports unquantized FP16, FP8, and GPTQ/AWQ quantized weights. Our single-GPU benchmarks use full FP16 weights, matching our Intel Arc Pro B70 results. The 27B multi-GPU test evaluates both stock vLLM FP8 Tensor Parallelism (TP=2) and 4-bit (Q4_K_M) llama.cpp.

ModelTypeParamsFP16 VRAMFits 1× R9700?
(32 GB)
Fits 2× R9700?
(64 GB)
Qwen2.5 3B InstructDense3B~6 GB✅ Yes✅ Yes
Qwen3 8BDense
(thinking)
8B~16 GB✅ Yes✅ Yes
Llama 3.1 8B InstructDense8B~16 GB✅ Yes✅ Yes
DeepSeek R1 Distill 8BDense8B~16 GB✅ Yes✅ Yes
Qwen3.6-27BDense27B~54 GB❌ No✅ Tested
(FP8 TP=2 in vLLM / Q4 in llama.cpp)
Qwen3.6-35B-A3BMoE35B
(3B active)
~70 GB❌ No❌ No
Gemma 4 31BDense31B~62 GB❌ No⚠️ Requires bfloat16
Llama 4 ScoutMoE109B
(17B active)
~218 GB❌ No❌ No
DeepSeek V4 FlashMoE284B
(13B active)
~568 GB❌ No❌ No

A single R9700 comfortably runs everything up to 8B parameters in unquantized FP16. The 27B tier, which is currently the sweet spot for capable local inference, requires both cards working together. In unquantized BF16, a 27B model consumes ~55.6 GB of raw weights alone (27.8 GB per card against ~29.86 GB of usable VRAM), leaving virtually no headroom for KV caching. Quantizing to FP8 in vLLM or Q4_K_M in llama.cpp solves this, preserving full reasoning quality while leaving ample VRAM headroom for deep context windows and concurrent request batches.

Single-GPU Performance

We started by establishing what a single R9700 can deliver. Each model was tested at concurrency levels 1, 4, and 8 to measure both interactive responsiveness and throughput under load.

How to read these numbers:

  • Throughput is the total generation rate in tokens per second
  • Time to first token (TTFT) is how long before the model starts replying
  • Inter-token latency (ITL) is the gap between streamed tokens
  • Average latency is the total time to finish a request

As rules of thumb, a TTFT under about 200 ms feels instant. Since people read at roughly 7 to 13 tokens per second, any ITL below 100 ms outpaces reading speed (under 30 ms feels glass smooth). Throughput is the metric that scales with concurrency: one user rarely consumes a GPU’s full output rate, but several together can.

Summary: Single-GPU Results (Concurrency = 1)

ModelThroughputTTFTITLAvg Latency
DeepSeek R1 Distill 8B61.7 tok/s111 ms16 ms10.8 s
Qwen2.5 3B Instruct53.3 tok/s60 ms19 ms6.1 s
Llama 3.1 8B Instruct29.1 tok/s113 ms34 ms13.0 s
Qwen3 8B27.7 tok/s92 ms36 ms18.1 s

All tests: 500 input tokens, 500 output tokens, FP16, HIP graphs enabled, single R9700, max-model-len 16384. Qwen3 8B used a 120-second measurement interval to accommodate its reasoning/thinking token generation.

The performance metrics in the data shows the following:

  • DeepSeek R1 8B delivers 61.7 tok/s, the fastest 8B model we tested. Despite being a reasoning-distilled model, it outpaces the Llama family on raw throughput. A smooth 16 ms ITL (Inter-Token Latency) means tokens arrive well ahead of reading speed.
  • Qwen2.5 3B reaches 53.3 tok/s with the quickest first token of any model at 60 ms — competitive with cloud API response times. You might expect a 3B model to run several times faster than the 8B models, but it does not: it trails DeepSeek R1 8B (61.7 tok/s) and sits only modestly ahead of Llama 3.1 8B and Qwen3 8B. Single-stream decode is memory-bandwidth-bound on RDNA 4, so once a model fits in VRAM, parameter count matters far less than the R9700’s 640 GB/s bandwidth ceiling, which all of these models push up against.
  • Llama 3.1 8B hits 29.1 tok/s, which is usable for interactive chat. The 113 ms TTFT and 34 ms ITL mean responses start quickly and tokens arrive at a comfortable reading pace.
  • Qwen3 8B delivers 27.7 tok/s. As a thinking/reasoning model, Qwen3 generates internal reasoning tokens before visible output, which requires an extended measurement window (120 seconds vs. the standard 30) to capture properly; see the note below for more details. Its throughput is the strongest under concurrency of any model tested. Concurrency results for every model are in the expandable tables below.

A note on benchmarking reasoning models: Qwen3 8B initially appeared to produce 0.0 tok/s in our standard 30-second measurement window. The model’s thinking phase, where it generates internal reasoning tokens before visible output, consumed the entire measurement interval, causing GenAI-Perf to report zero completed requests. Extending the measurement window to 120 seconds revealed the true throughput: 27.7 tok/s with a 36 ms ITL. This is a cautionary example for anyone benchmarking thinking models: standard measurement configurations can fundamentally misrepresent their performance. The 27.7 tok/s we report for Qwen3 8B throughout this article uses that corrected 120-second configuration, which we have published alongside our other benchmark configs.

Detailed Single-GPU Benchmark Tables (click to expand)

A note on the concurrency rows: throughput is aggregate, summed across all simultaneous users, while TTFT and ITL are per-request averages across every request completed in the measurement window. As concurrency rises, total system throughput climbs even though each individual request waits a little longer. P99 latency is the worst-case full-request time any single user experienced.

Qwen2.5 3B Instruct (FP16, single GPU)

ConcurrencyThroughputTTFT avgITL avgAvg LatencyP99 Latency
153.3 tok/s60 ms19 ms6.1 s7.5 s
4183 tok/s76 ms21 ms7.2 s10.5 s
8251 tok/s389 ms30 ms10.4 s18.8 s

Throughput scales from 53 tok/s (single user) to 183 tok/s (4 users). That is 3.4× the single-user throughput, with only a small latency penalty. At 8 users the system reaches 251 tok/s as compute saturation sets in, with TTFT climbing to 389 ms and ITL roughly 1.6× the single-user figure.

Llama 3.1 8B Instruct (FP16, single GPU)

ConcurrencyThroughputTTFT avgITL avgAvg LatencyP99 Latency
129.1 tok/s113 ms34 ms13.0 s17.2 s
4102 tok/s162 ms37 ms12.2 s18.5 s
8153 tok/s405 ms47 ms15.8 s25.2 s

Llama 3.1 scales from 29.1 tok/s to 102 tok/s at 4 users (3.5× the single-user throughput) while latency stays essentially flat. At 8 concurrent users, throughput reaches 153 tok/s, demonstrating that vLLM’s continuous batching extracts strong parallelism from the R9700’s compute units.

DeepSeek R1 Distill Llama 8B (FP16, single GPU)

ConcurrencyThroughputTTFT avgITL avgAvg LatencyP99 Latency
161.7 tok/s111 ms16 ms10.8 s14.6 s
4248 tok/s139 ms16 ms13.2 s15.5 s
8329 tok/s429 ms22 ms18.4 s22.4 s

DeepSeek R1 8B is the throughput champion at 61.7 tok/s single-user. At 4 concurrent users, it delivers 248 tok/s with only 28 ms of additional TTFT penalty. The 16 ms ITL means token delivery is exceptionally smooth. At 8 users it reaches 329 tok/s, which is the highest aggregate throughput of any model we tested, and it does so while generating the longest reasoning sequences.

Qwen3 8B (FP16, single GPU) – Thinking Model

ConcurrencyThroughputTTFT avgITL avgAvg LatencyP99 Latency
127.7 tok/s92 ms36 ms18.1 s18.1 s
4104 tok/s110 ms38 ms19.3 s19.3 s
8157 tok/s383 ms50 ms25.4 s29.0 s

Qwen3 8B is a thinking/reasoning model that generates internal reasoning tokens before visible output. It runs at 27.7 tok/s single-user, scaling to 157 tok/s at 8 concurrent users. That is 5.7× the single-user throughput, and the best concurrency scaling of any model tested — suggesting the R9700’s compute units handle batched reasoning workloads particularly well. The 92 ms TTFT and 36 ms ITL make it fully suitable for interactive use.

Dual-GPU Performance

The R9700’s real value proposition emerges when two cards work together. With 64 GB of aggregate VRAM across two cards, models up to 27B parameters — which cannot run on a single 32 GB card — become accessible for production serving.

Dual-GPU Serving: Stock vLLM TP=2, MTP, and llama.cpp

On bare metal, stock vLLM (vllm/vllm-openai-rocm:v0.20.2) cleanly supports Tensor Parallelism (TP=2) on RDNA 4. During startup on bare metal, RCCL collective initialization completes in just 0.12 seconds via P2P/IPC over PCIe, and the engine serves tokens reliably across both R9700 cards.

For the 27B parameter tier (tested using Qwen3.6-27B), running in FP8 precision is the ideal deployment strategy: it fits comfortably within the combined 64 GB pool while leaving ample VRAM for deep KV-cache allocation under heavy concurrency.

Qwen3.6-27B Multi-GPU Performance Sweep (2× R9700)

Configuration / StackConcurrencyThroughputTTFT avgITL avgGPU Power (Both Cards)
Stock vLLM FP8 (TP=2)
(v0.20.2 standard image)
115.88 tok/s337 ms61.6 ms310 W
469.34 tok/s541 ms55.3 ms465 W
8156.21 tok/s802 ms47.4 ms480 W
Tuned vLLM FP8 (TP=2)
(vllm-radiance build)
132.00 tok/s136 ms30.7 ms312 W
4110.56 tok/s372 ms34.5 ms498 W
8198.78 tok/s621 ms37.3 ms479 W
Tuned vLLM FP8 + MTP (TP=2)
(Speculative Decoding enabled)
162.75 tok/s184 ms15.1 ms315 W
4186.29 tok/s275 ms20.1 ms505 W
8320.23 tok/s280 ms23.6 ms490 W
llama.cpp Q4_K_M GGUF
(Layer-split mode)
123.40 tok/s382 ms41.0 ms339 W
455.70 tok/s884 ms64.0 ms350 W
856.00 tok/s2,668 ms126.0 ms353 W

All tests: standardized 500 input / 200 output tokens, 50 prompts, streaming mode, bare metal on 2× R9700.

The comparative data highlights distinct deployment strengths across stacks:

  • Single-User vs. Concurrent Serving: For single-user interactive chat, llama.cpp with Q4_K_M quantization delivers 23.4 tok/s, outpacing stock vLLM’s single-stream rate (15.9 tok/s). However, as soon as multi-user concurrency is introduced, vLLM’s continuous batching and Tensor Parallelism take over: stock vLLM scales to 156.2 tok/s at 8 concurrent users — nearly 2.8× the throughput of llama.cpp (which hits a plateau around 56 tok/s).
  • Multi-Token Prediction (MTP) Breakthrough: Enabling speculative decoding via Multi-Token Prediction (MTP) nearly doubles single-stream performance to 62.75 tok/s with a silky 15.1 ms ITL, and powers all the way to 320.23 tok/s under 8-user concurrency. Because MTP validation is mathematically exact, generated outputs remain bit-identical to standard autoregressive decode while providing a massive speedup.

Scaling Summary

ModelPrecisionEngineConfigThroughput (c=1 / c=8)Notes
DeepSeek R1 8BFP16vLLM1× R970061.7 / 329.0 tok/sFastest 8B; 16 ms ITL
Qwen2.5 3BFP16vLLM1× R970053.3 / 251.0 tok/sQuickest TTFT (60 ms)
Llama 3.1 8BFP16vLLM1× R970029.1 / 153.0 tok/sInteractive chat baseline
Qwen3 8BFP16vLLM1× R970027.7 / 157.0 tok/sThinking model; 5.7× concurrency scaling
Qwen3.6-27B (Stock TP=2)FP8vLLM2× R9700 (TP=2)15.9 / 156.2 tok/sStandard vLLM multi-GPU serving
Qwen3.6-27B (Tuned + MTP)FP8vLLM2× R9700 (TP=2)62.8 / 320.2 tok/sMTP speculative decode; 15.1 ms ITL
Qwen3.6-27B (llama.cpp)Q4_K_Mllama.cpp2× R9700 (layer)23.4 / 56.0 tok/sLightweight single-user local deployment

The results establish clear guidelines across tiers:

  1. 8B models and smaller run on a single R9700 in FP16: With 32 GB and 640 GB/s per card, a single GPU easily hosts 8B models. DeepSeek R1 8B leads at 61.7 tok/s (scaling to 329 tok/s under concurrency), followed by Qwen2.5 3B at 53.3 tok/s, Llama 3.1 8B at 29.1 tok/s, and Qwen3 8B at 27.7 tok/s.
  2. The 27B tier scales powerfully across two cards: Spanning two R9700s via stock vLLM Tensor Parallelism (TP=2) in FP8 delivers robust production serving at 156.2 tok/s, while enabling MTP speculative decoding pushes throughput to 320.2 tok/s. For standalone single-user workstations, llama.cpp GGUF layer splitting offers an accessible 23.4 tok/s baseline.

Cost of Inference: Local vs. Cloud

Raw throughput numbers look great on paper, but the real question every IT lead and engineering manager asks is simple: what does it actually cost to run these models locally, and how does that compare to paying for a cloud API? To answer it with real data rather than rules of thumb, we instrumented our benchmark suite with direct GPU power monitoring, polling the AMD kernel driver’s power1_average sensor every 2 seconds throughout each test run.

Before diving into the numbers, two core realities set the baseline:

This is a cost comparison, not a quality one. The open models we tested do not match frontier APIs on hard reasoning. On the third-party Artificial Analysis Intelligence Index , Qwen3.6-27B in reasoning mode scores 37 against Gemini 3.1 Pro’s 46. Published head-to-head comparisons put even inexpensive cloud tiers ahead of it on several reasoning benchmarks. We are not claiming parity. The economics below matter when an open model is already good enough for the job (summarization, classification, chat, tool use, bulk generation) and per-token meter is the thing you want to stop paying.

Cloud pricing spans a wide range, so which model you compare against decides the answer. Output pricing verified 2026-08-12:

TierModel$/1M output
FrontierGPT-5.6 Sol$30.00
FrontierClaude Opus 5$25.00
FrontierGemini 3.1 Pro (≤200K), GPT-5.6 Terra$12.00
MidClaude Sonnet 5$10.00
MidGemini 3.6 Flash$7.50
MidClaude Haiku 4.5$5.00
EconomyGemini 3.1 Flash-Lite$1.50
EconomyGPT-5.6 Luna$1.20

A workstation that beats GPT-5.6 Sol on cost may still lose to GPT-5.6 Luna. Which row you belong on depends on the tier your workload actually needs, so we price against all of them.

Assumptions

Every figure below comes from these five inputs. Swap in your own and the arithmetic follows:

InputValue usedBasis
Workstation price$18,775Puget T142-XL configurator, as-tested specification, August 2026
Electricity$0.1354/kWhEIA Electric Power Monthly, US commercial average (May 2026 data)
System overhead~300 WCPU package measured at ~120 W under load; RAM, drives, fans and PSU losses estimated. Conservative — see footnote
Hardware life3 yearsStraight-line, no residual value
Cloud pricesSee table aboveVendor pricing pages, verified 08-12-2026
Local modelQwen3.6-27B (FP8 vLLM / Q4 llama.cpp) and the FP16 8B tierBenchmarked August 2026

Measured GPU Power Under Load

Power is not a single number per model: it climbs with concurrency as batching fills the GPU. We segmented each capture by benchmark phase so every cost figure below uses the power actually drawn at that concurrency level.

ModelConfigc=1c=4c=8Peak
DeepSeek R1 8B1× R9700221W249W283W298W
Qwen3 8B1× R9700205W223W261W269W
Llama 3.1 8B1× R9700205W227W259W278W
Qwen2.5 3B1× R9700182W200W218W226W
Qwen3.6-27B (Stock FP8 TP=2)2× R9700310W465W480W495W
Qwen3.6-27B (Tuned + MTP)2× R9700315W505W490W510W
Qwen3.6-27B (Q4, llama.cpp)2× R9700339W350W353W353W

GPU-only power from the AMD kernel driver (sysfs hwmon), averaged over each concurrency phase. Single-GPU rows are the draw of the one active card; dual-GPU rows represent total draw across both cards. Total system wall power adds CPU, RAM, and PSU overhead: we estimate ~300W for this workstation.

Cost Per Million Output Tokens

Electricity alone, using measured per-phase power plus an estimated 300W of system overhead, at the EIA US commercial average of $0.1354/kWh1:

Modelc=1c=4c=8
DeepSeek R1 8B$0.32$0.08$0.07
Qwen2.5 3B$0.34$0.10$0.08
Llama 3.1 8B$0.65$0.19$0.14
Qwen3 8B$0.69$0.19$0.13
Qwen3.6-27B (Stock vLLM FP8 TP=2)$1.79$0.44$0.21
Qwen3.6-27B (Tuned + MTP)$0.45$0.16$0.095
Qwen3.6-27B (Q4, llama.cpp)$1.03$0.43$0.43

Dollars per million output tokens, electricity only. Hardware is handled separately below, and dominates the overall cost analysis.

Serving eight users concurrently costs roughly a third to a fifth as much per token as serving just one, because token throughput multiplies dramatically while power draw creeps up modestly. In practice, keeping the hardware actively saturated is what drives down local inference costs.

The Real Driver: Hardware, Not Electricity

A workstation like the one we tested, configured today with a 24-core Threadripper PRO 9965WX, 128 GB of RAM, and two R9700s, comes to $18,7752. Amortized over three years, that dwarfs the power bill. Calculating the all-in cost per million output tokens involves two terms3:

At 329 tokens/sec, running this workstation saturated over a 3-year lifespan yields ~185 million tokens for every weekly active hour. The cost calculation boils down to spreading the $18,775 machine cost across active serving hours, plus a flat ~$0.07 per million tokens in electricity.

The figures below track DeepSeek R1 8B on a single card at eight concurrent users: 283 W of GPU draw plus roughly 300 W of system overhead, for 583 W total. At 329 tok/s, every weekly hour of saturated serving pushes out 184.8 million tokens over a three-year hardware lifespan, and electricity works out to a fixed 0.492 kWh per million tokens. That power figure remains constant per token regardless of schedule, because running twice as long burns twice the energy to produce twice the tokens. Spreading the machine’s purchase price across active hours is what moves the needle:

$/kWh8 h/wk20 h/wk40 h/wk80 h/wk168 h/wk
$0.10$12.75$5.13$2.59$1.32$0.65
$0.1354 (US commercial avg)$12.77$5.15$2.61$1.34$0.67
$0.1844 (US residential avg)$12.79$5.17$2.63$1.36$0.70
$0.25$12.82$5.20$2.66$1.39$0.73

All-in dollars per million output tokens: $18,775 amortized over 3 years plus measured electricity, at 329 tok/s and 583W total system draw.

Moving from the cheapest power in the country to among the most expensive shifts overall cost per token by only 0.6% to 11%. By contrast, moving from 8 hours a week of usage to full-time serving slashes token costs by 19×. If your team runs automated batch processing or continuous agent workflows overnight, local hardware hits break-even almost immediately. If it’s just occasional ad-hoc developer chats, cloud APIs remain the cheaper path.

Break-Even: How Busy Does It Need to Be?

An AI workstation earns its price by displacing tokens you would otherwise rent. Here is how many hours per week of sustained eight-user serving that takes, over a three-year life, against each cloud tier:

Cloud model$/1M outputDeepSeek R1 8B (c=8)27B Stock TP=2 (c=8)27B Tuned + MTP (c=8)
GPT-5.6 Sol$30.003.4 h/week7.1 h/week3.5 h/week
Claude Opus 5$25.004.1 h/week8.6 h/week4.2 h/week
Gemini 3.1 Pro / GPT-5.6 Terra$12.008.5 h/week18.2 h/week8.8 h/week
Claude Sonnet 5$10.0010.2 h/week21.9 h/week10.6 h/week
Gemini 3.6 Flash$7.5013.7 h/week29.4 h/week14.2 h/week
Claude Haiku 4.5$5.0020.6 h/week44.7 h/week21.3 h/week
Gemini 3.1 Flash-Lite$1.5070.8 h/week165.9 h/week74.3 h/week
GPT-5.6 Luna$1.2089.6 h/week216.3 h/week94.5 h/week

Hours per week of saturated 8-concurrent serving needed to beat each API on cost, at $0.1354/kWh over a 3-year hardware life. With MTP enabled, the 27B model clears economy tiers within a standard 168-hour week.

Against frontier pricing the bar is exceptionally low. Three and a half hours a week of real load beats GPT-5.6 Sol, and a single workday a week beats Gemini 3.1 Pro. Furthermore, with Multi-Token Prediction (MTP) enabled, serving a 27B model reaches break-even against Gemini 3.1 Flash-Lite in ~74 hours per week and GPT-5.6 Luna in ~95 hours per week — bringing even ultra-cheap cloud tiers within reach of local workstation economics.

The Cards Are the Upgrade, Not the Machine

Those figures charge the entire $18,775 against inference, as if the workstation did nothing else all day. That is the harshest possible accounting, and it is not the situation most buyers are in. A 24-core Threadripper PRO with 128 GB of RAM might have been a workstation you were buying anyway. The AI capability is the pair of R9700s, and specifying them in place of the base configuration’s entry-level card adds over $3,000 to the build.

Priced as an upgrade to a machine you already need, the break-even line drops through the floor:

Cloud model$/1M outputAs a whole machineAs a GPU upgrade
GPT-5.6 Sol$30.003.4 h/week0.6 h/week
Claude Opus 5$25.004.1 h/week0.7 h/week
Gemini 3.1 Pro / GPT-5.6 Terra$12.008.5 h/week1.5 h/week
Gemini 3.6 Flash$7.5013.7 h/week2.4 h/week
Gemini 3.1 Flash-Lite$1.5070.8 h/week12.4 h/week
GPT-5.6 Luna$1.2089.6 h/week15.7 h/week

Calculations based on a $3,283 upgrade price for the pair of R9700s, as of 2026-08-12.

Thirty-six minutes a week of sustained serving covers the cards against GPT-5.6 Sol. At ten hours a week they pay for themselves in two months against flagship pricing and five months against Gemini 3.1 Pro. Even against the cheapest model on the market, fifteen hours a week clears it – while the whole-machine accounting never did.

That is the practical business case for the R9700 in particular. A pair of cards delivers 64 GB of VRAM to handle the 27B tier for a $3,283 upgrade from a basic video card — less than a single high-end workstation GPU from team green. Or, if you are pricing an upgrade for an existing workstation, $3,760 as of August 2026. Unlike cloud API bills that compound indefinitely with every prompt and completion, local hardware is a one-time capital investment that caps your inference costs.

What Never Shows Up in a Token Price

Cost is only part of why engineering teams bring inference in-house. Beyond basic ROI, local hardware solves real day-to-day operational headaches: API rate limits during traffic spikes, models being deprecated or quietly altered mid-project, unexpected pricing updates, and strict compliance or legal barriers around sending sensitive IP to third-party endpoints. For regulated, confidential, or high-volume workflows, these operational guarantees are often the main reason to build locally.

Image Generation: ComfyUI + Z-Image Turbo

Inference isn’t only about language models, so we put the R9700 through a generative image workload as well: ComfyUI running Z-Image Turbo, a distilled diffusion model that renders 1024×1024 images in just four sampling steps.

ComfyUI ran on a current ROCm 7.2 PyTorch build and, once the container had access to the GPU device nodes, saw both R9700s as native HIP devices without needing code changes to ComfyUI or the model. Generation used a single card, with the diffusion model and its Qwen-based text encoder resident in about 18 GB, comfortably inside the card’s 32 GB.

Results

Prompt: “A red fox standing alert in a vibrant autumn forest at golden hour, surrounded by dense fiery-orange and crimson foliage filling the entire frame, a thick carpet of fallen leaves, shafts of warm morning light through the trees, deep depth of field with the whole scene in sharp focus, richly detailed nature photography”

MetricValue
Iterations10/10 passed (0 failures)
Cold Start (iter 1)17.4s (model load from disk + HIP JIT)
Steady State (iter 2–10)3.6s average
Mean (all 10)5.0 s
p503.5 s
Throughput (steady state)16.8 images/min
VRAM Used~18 GB of 32 GB

The steady-state number is what matters: once the first iteration loads the ~18 GB model, the R9700 renders a 1024×1024 image every 3.6 seconds, about 17 images per minute. That leaves roughly 14 GB of headroom on the card for larger models, higher resolutions, or multi-model pipelines, and all 10 runs completed without a single failure or artifact.

Sample Output

These images were generated on the R9700 with Z-Image Turbo using the fox prompt above, each in about 3.6 seconds at steady state:

Previous Next
System Image
Open Full Resolution
Open Full Resolution
Open Full Resolution
Previous Next

They were all generated from the same prompt and settings, varying only the random seed. The R9700 produced clean, detailed 1024×1024 results with stable HIP inference throughout.

Comparison: R9700 vs. Arc Pro B70

Both the R9700 and the Arc Pro B70 target the same market: professional AI inference at 32 GB per card. They come from different architectural families, though, with different tradeoffs. We tested both cards on equivalent workloads using our benchmark framework, allowing direct comparison.

Methodology note: Both articles use NVIDIA GenAI-Perf with --streaming, 500 input / 500 output tokens, 50 prompts, at concurrency 1, 4, and 8. Please note that the R9700 numbers here were re-measured with HIP graphs enabled and warmup requests discarded (our corrected methodology) while the B70 numbers are from our previously published article. The figures below are indicative, but a definitive, methodology-matched, head-to-head belongs in a dedicated comparison post. Treat single-digit-percent gaps as ties. One asymmetry is especially worth naming: the B70 figures were captured before we adopted warmup discarding, so they carry some cold-start drag and would likely improve slightly under our current methodology. That bias runs in the B70’s favor here, meaning its lead on these models is a conservative reading rather than an inflated one, though only a matched re-run on both cards can settle it. The original figures are in our Intel Arc Pro B70 article.

Hardware Comparison

GPU Hardware ComparisonAMD Radeon AI PRO R9700Intel Arc Pro B70
ArchitectureRDNA 4 (gfx1201)Xe2-HPG (Battlemage)
VRAM32 GB GDDR632 GB GDDR6 (ECC)
Memory Bandwidth640 GB/s608 GB/s
AI TOPS (INT8)766 TOPS367 TOPS
TBP300W230W
Price per card (as of July 2026)~$1,880~$1,110
Cards Tested24
Total VRAM64 GB (~$3,760)128 GB (~$4,450)
Inference StackROCm + vLLM / llama.cppXPU + vLLM / LLM Scaler
Multi-GPU MethodTensor Parallelism (vLLM TP=2) & llama.cpp layer splitTensor Parallelism (TP)

Single-GPU Throughput (Concurrency = 1)

ModelR9700B70Difference
DeepSeek R1 8B61.7 tok/s66.9 tok/sB70 +8%
Llama 3.1 8B29.1 tok/s35.4 tok/sB70 +22%
Qwen2.5 3B53.3 tok/s72.9 tok/sB70 +37%
Qwen3 8B (thinking)27.7 tok/s34.7 tok/sB70 +25%

On these single-GPU numbers the B70 leads across all four models, from an 8% edge on DeepSeek to 37% on the small 3B. Two honest caveats temper that:

  • Methodology asymmetry: the R9700 figures were freshly re-measured with our corrected methodology (HIP graphs on, warmup discarded) on the bare-metal system described above, while the B70 figures come from its earlier article. We have not yet re-run the B70 under the identical harness, so some of this gap may be measurement rather than silicon. The definitive answer requires a matched head-to-head, which we will run separately.
  • Both are memory-bandwidth-bound: with near-identical bandwidth (640 vs. 608 GB/s), neither card should hold a large architectural decode advantage. Where the numbers diverge more than bandwidth would predict — the 3B especially — software maturity (kernel efficiency in each vendor’s stack) is the likely driver, and both stacks are moving quickly.

Treat this as “the B70 is currently at least competitive and often ahead on single-GPU throughput” — with the matched comparison still to come.

Multi-GPU: 27B Dense Model

When scaling to the 27B parameter tier across multiple cards, the architectural and footprint tradeoffs become clear:

27B Dense ModelRadeon AI PRO R9700 (2 cards)Arc Pro B70 (4 cards)
Multi-GPU Serving PathvLLM Tensor Parallelism (TP=2)vLLM Tensor Parallelism (TP=4)
27B Concurrency=8 Throughput156.2 tok/s (Stock FP8)
320.2 tok/s (Tuned + MTP)
95.9 tok/s (FP16)
PrecisionFP8 / Q4_K_M16-bit (FP16)
Total GPU Cost (as of July 2026)~$3,760 (2 cards)~$4,450 (4 cards)
Total VRAM & Footprint64 GB (2 slots)128 GB (4 slots)

The AMD Radeon AI PRO R9700 delivers massive throughput density in the 27B class, generating 156.2 tok/s in stock vLLM FP8 and up to 320.2 tok/s with MTP across just two cards ($3,760). By contrast, the Intel Arc Pro B70 setup requires four cards ($4,450) and four PCIe slots to serve 27B at 95.9 tok/s in full FP16 (giving stock R9700 FP8 a 1.63× throughput advantage at 156.2 tok/s, or 3.3× with MTP). However, the four-card B70 system provides 128 GB of aggregate VRAM — enough headroom to accommodate 35B+ Mixture of Experts models that exceed a 64 GB boundary.

Image Generation

Image GenerationAMD Radeon AI PRO R9700Intel Arc Pro B70
Steady State3.6s per image3.9s per image
Throughput (steady)16.8 img/min15.4 img/min
VRAM Used~18 GB19.3 GB

The R9700 is approximately 8% faster at image generation, consistent with its bandwidth advantage. Both cards handle ComfyUI + Z-Image Turbo without issue.

What Doesn’t Work (and Deployment Considerations)

Three operational considerations and workarounds arose during testing:

Virtualization & VFIO Passthrough Limitations for Multi-GPU ROCm: While single-GPU vLLM inference runs without issue inside KVM/VFIO guest VMs, multi-GPU RCCL collective initialization (ncclCommInitRank) hangs inside passthrough environments despite full PCIe peer-to-peer DMA bandwidth (measured at 27.6 GB/s). On bare metal, the exact same container image and hardware initialize RCCL via P2P/IPC in 0.12 seconds and serve tokens immediately. If deploying multi-GPU vLLM on RDNA 4, bare-metal host execution is required. Additionally, attempting runtime rebinding of RDNA 4 GPUs between vfio-pci and amdgpu drivers at runtime can trigger kernel panics; cards should be claimed by amdgpu at host boot.

llama.cpp row split is architecture-dependent: --split-mode row ran our full concurrent sweep on Qwen2.5-32B without a single failure, but on Qwen3.6-27B the server aborts under concurrent decode with GGML_ASSERT(!(split && ne02 < ne12)). The row-split matrix-multiplication path does not yet support the broadcast shapes Qwen3.6’s architecture produces in batched decode. Until that is addressed upstream, use --split-mode layer for Qwen3.6-class models and verify row mode on your specific model architecture before deploying.

Container permissions for /dev/kfd: The rocm/pytorch:latest and vLLM containers require privileged: true or explicit device passthrough for /dev/kfd (AMD’s kernel fusion driver device node) and /dev/dri, along with membership in the video and render groups. Without these permissions, ROCm cannot access the GPU device nodes. This is a container configuration requirement, not a hardware limitation.

How Do Two R9700 GPUs Perform for AI Inference?

The AMD Radeon AI PRO R9700 delivers genuine AI inference capability on RDNA 4 silicon — and the economics make a strong case for local deployment:

  • 8B models at production-usable speeds on a single card: DeepSeek R1 8B hits 61.7 tok/s, Qwen2.5 3B reaches 53.3 tok/s, Llama 3.1 8B delivers 29.1 tok/s, and Qwen3 8B — a thinking/reasoning model — runs at 27.7 tok/s with the best concurrency scaling tested (157 tok/s at 8 users). These are honest, full-clock numbers measured with HIP graphs enabled. On single-GPU throughput, the Arc Pro B70 is currently competitive-to-ahead, pending a methodology-matched head-to-head.
  • The 27B tier scales powerfully across two cards: Spanning two R9700s via stock vLLM Tensor Parallelism (TP=2) in FP8 serves at 156.2 tok/s under 8-user concurrency, while tuned configurations with Multi-Token Prediction (MTP) reach 320.2 tok/s with bit-identical fidelity. For single-user interactive use, llama.cpp GGUF layer split provides an accessible 23.4 tok/s baseline.
  • Deployment Environment Matters: Multi-GPU RCCL collective communication requires a bare-metal host environment; virtualized KVM/VFIO passthrough setups encounter collective initialization hangs during bootstrap despite intact P2P bandwidth.
  • Reliable image generation: 1024×1024 images in 3.6 seconds steady-state via Z-Image Turbo, with zero failures across 10 runs.
  • Cheap to run, but the hardware is what you are paying for: measured GPU draw of 182–283W puts electricity at $0.07–$0.69 per million output tokens depending on concurrency. Hardware amortization dwarfs it. Against frontier APIs the machine clears break-even quickly, at 3.4 hours per week of saturated serving versus GPT-5.6 Sol and 8.5 hours versus Gemini 3.1 Pro. With MTP enabled on 27B models, break-even against economy tiers like GPT-5.6 Luna ($1.20/1M) drops to ~95 hours per week. Your duty cycle decides this, not your power rate.
  • Where it sits against Intel: The R9700 offers higher throughput density for 27B serving on half the cards and slots (2 cards, 64 GB, 156–320 tok/s vs 4 cards, 128 GB, 95.9 tok/s) — 1.63× the stock throughput and 3.3× with MTP, though the R9700 figures are FP8 against the B70’s FP16. The B70’s advantage lies in aggregate VRAM capacity across four cards for larger model footprints.

Here is how the three cards compare side by side:

GPU ComparisonAMD Radeon AI PRO R9700Intel Arc Pro B70NVIDIA RTX 5090
ArchitectureRDNA 4 (gfx1201)Xe2-HPG (Battlemage)Blackwell
VRAM32 GB GDDR632 GB GDDR6 (ECC)32 GB GDDR7
Memory Bandwidth640 GB/s608 GB/s1,792 GB/s
AI TOPS (INT8)766 TOPS367 TOPS3,352 TOPS (FP4 sparse)
TBP300 W230 W575 W
Price per card (configured, July 2026)~$1,880~$1,110~$4,130
Tested Config VRAM64 GB (2 cards, ~$3,760)128 GB (4 cards, ~$4,450)64 GB (2 cards, ~$8,260)
8B FP16 tok/s (single card)61.7 (DeepSeek R1)66.9 (DeepSeek R1)~140–200
27B tok/s (multi-card c=8)156.2 (Stock FP8 TP=2)
320.2 (Tuned + MTP)
95.9 (FP16, TP=4, 4 cards)N/A (single), not tested (multi)
$/1M tokens (8B, electricity)$0.07 (c=8)Not measuredNot measured
Multi-GPU MethodTensor Parallelism (vLLM TP=2)Tensor Parallelism (TP=4)Tensor Parallelism

The RTX 5090 remains roughly 3–4× faster per GPU on decode-bound workloads, driven by nearly 3× the memory bandwidth. But in today’s supply-constrained market, that speed carries a steep premium: two R9700 cards deliver the same aggregate VRAM for less than half the cost of two RTX 5090s. The B70 offers the most VRAM per dollar at 128 GB across four cards, while the R9700 delivers unmatched inference density and bandwidth in a dual-slot footprint.

For teams running models up to the 27B class where privacy, cost control, or volume matter, the R9700 dual-card configuration is a compelling option, and volume is the deciding lever. Run the numbers at a fixed $6,262/year, which is this workstation amortized over three years plus power, because you pay for the machine whether it is busy or not. A team generating 5 million output tokens per month is buying that $6,262 to displace $720/year of Gemini 3.1 Pro traffic or $1,800/year of GPT-5.6 Sol. On cost alone that is a clear loss, and at that volume the honest case for local is privacy and control, not payback. Scale to 50 million tokens per month, roughly ten hours a week of saturated serving, and the same $6,299 stands against $7,200/year for Gemini 3.1 Pro, $15,000 for Claude Opus 5 and $18,000 for GPT-5.6 Sol. That is where the machine starts winning on cost, and it wins decisively against flagship traffic.

Teams needing larger VRAM pools (35B+ Mixture of Experts models, for example) should consider adding more R9700 cards or evaluating the 4-card B70 configuration.

Appendix: Setup Guide for Practitioners

Docker Compose Reference: vLLM Single-GPU and Dual-GPU (TP=2)

The vllm/vllm-openai-rocm:v0.20.2 image has entrypoint vllm serve. For dual-GPU Tensor Parallelism (TP=2) on bare metal, allocate shared memory and pass both GPU devices:

services:
  inference:
    image: vllm/vllm-openai-rocm:v0.20.2
    privileged: true
    shm_size: "32g"
    ipc: host
    devices:
      - /dev/kfd:/dev/kfd
      - /dev/dri:/dev/dri
    environment:
      - VLLM_TARGET_DEVICE=rocm
      - HIP_FORCE_DEV_KERNARG=1
    ports:
      - "8000:8000"
    command:
      - Qwen/Qwen3.6-27B-Instruct-FP8
      - --tensor-parallel-size=2
      - --gpu-memory-utilization=0.92
      - --max-model-len=16384

Leave HIP graphs on. Do not pass --enforce-eager unless diagnosing an issue, as eager mode roughly halves decode throughput on these cards.

Multi-GPU: llama.cpp GGUF Alternative

For standalone single-user workstations or GGUF quantization, llama.cpp distributes layers across GPUs over direct HIP transfers:

# llama.cpp: multi-GPU over direct HIP transfers (layer split)
docker run -d --device /dev/kfd --device /dev/dri \
  --security-opt seccomp=unconfined --group-add video --group-add render \
  -p 8000:8000 --entrypoint /app/llama-server \
  ghcr.io/ggml-org/llama.cpp:server-rocm \
  -hf bartowski/Qwen_Qwen3.6-27B-GGUF:Q4_K_M \
  -ngl 99 --split-mode layer -c 16384 --parallel 8 \
  --host 0.0.0.0 --port 8000 --jinja

ComfyUI Docker Compose (Image Generation)

services:
  comfyui:
    image: rocm/pytorch:latest
    privileged: true
    user: root
    devices:
      - /dev/kfd:/dev/kfd
      - /dev/dri:/dev/dri
    shm_size: "16g"

Tested on a Puget Systems workstation (AMD Threadripper PRO 5995WX, ASUS Pro WS WRX80E-SAGE SE, Rocky Linux 10.1) with 2× AMD Radeon AI PRO R9700. Benchmarks ran on bare metal with cards bound to amdgpu at boot, with PCIe peer-to-peer verified. Single-GPU LLM benchmarks used vllm/vllm-openai-rocm:v0.20.2 with FP16 weights and HIP graphs enabled; every measurement discarded 3 warmup requests. Multi-GPU 27B benchmarks evaluated stock vLLM FP8 (TP=2), tuned vLLM with MTP, and ghcr.io/ggml-org/llama.cpp:server-rocm with Q4_K_M weights. Image generation used rocm/pytorch:latest with ComfyUI and Z-Image Turbo. GPU power measured via sysfs hwmon (power1_average). Cloud API pricing verified August 2026, while GPU hardware pricing reflects Puget Systems configured-system pricing.

  1. Cost per million output tokens is electricity only: (measured GPU power for that concurrency phase + ~300 W estimated system overhead) × $0.1354/kWh, divided by the measured throughput at that concurrency level. The rate is the US commercial average published by the U.S. Energy Information Administration in Electric Power Monthly, Table 5.3 (May 2026 data). Commercial rates are the relevant basis for a business deployment; the residential average is materially higher at $0.1844/kWh, and the January–May 2026 commercial average is $0.1379/kWh, so a reader on residential power should scale the electricity column up accordingly. System overhead is partly measured: CPU package draw was sampled at ~120 W under load (99.7 W idle) via the kernel’s RAPL interface, and the remainder (RAM, drives, fans, PSU conversion losses) is estimated. The 300 W figure we use is deliberately conservative; the components we can account for total closer to 250 W, so these cost figures if anything overstate local inference. Hardware amortization is handled separately in the next section. ↩︎
  2. Our test system uses a prior-generation Threadripper PRO 5995WX, which is no longer sold. Rather than price hardware that is unavailable, this figure reflects a current Puget Threadripper PRO configured comparably (24-core 9965WX, 128 GB RAM, 2× R9700) at August 2026 pricing. Every cost figure in this section is derived from the five inputs listed in the assumptions box above; substitute your own and the arithmetic follows. ↩︎
  3. For readers who want the exact algebraic equation: Cost per 1M tokens ($) = [$18,775 ÷ (184.8 × H)] + (0.492 × R), where H is hours per week of saturated 8-user serving over a 3-year hardware lifespan, and R is the electricity rate in $/kWh. ↩︎
Tower Computer Icon in Puget Systems Colors

Looking for an AI workstation or server?

We build computers tailor-made for your workflow. 

Configure a System
Talking Head Icon in Puget Systems Colors

Don’t know where to start?
We can help!

Get in touch with one of our technical consultants today.

Talk to an Expert

Related Content

  • AMD Radeon AI PRO R9700: Dual-GPU AI Inference Performance
  • AMD Radeon RX 9070 GRE Content Creation Review
  • Intel Arc Pro B70: Multi-GPU AI Inference Performance
  • Intel Arc Pro B70 Review
View All Related Content

Latest Content

  • Consultant’s Corner: Backseat Driving Computers
  • Consultant’s Corner: Sourcing Systems in Shortage
  • Topaz Video 1.6.1 – Professional GPU Performance Analysis
  • AMD Radeon AI PRO R9700: Dual-GPU AI Inference Performance
View All

Who is Puget Systems?

Puget Systems builds custom workstations, servers and storage solutions tailored for your work.

We provide:

Extensive performance testing
making you more productive and giving better value for your money

Reliable computers
with fewer crashes means more time working & less time waiting

Support that understands
your complex workflows and can get you back up & running ASAP

A proven track record
as shown by our case studies and customer testimonials

Get Started

Browse Systems

Puget Systems Mobile Laptop Workstation Icon

Mobile

Puget Systems Tower Workstation Icon

Workstations

Puget Systems Rackmount Workstation Icon

Rackstations

Puget Systems Rackmount Server Icon

Servers

Puget Systems Rackmount Storage Icon

Storage

Latest Articles

  • Consultant’s Corner: Backseat Driving Computers
  • Consultant’s Corner: Sourcing Systems in Shortage
  • Topaz Video 1.6.1 – Professional GPU Performance Analysis
  • AMD Radeon AI PRO R9700: Dual-GPU AI Inference Performance
  • Topaz Video 1.6.1 – Consumer GPU Performance Analysis
View All

Post navigation

 Topaz Video 1.6.1 – Consumer GPU Performance AnalysisTopaz Video 1.6.1 – Professional GPU Performance Analysis 
Puget Systems Logo in White and Green
Build Your Own PC Site Map FAQ
facebook instagram linkedin rss twitter youtube

Optimized Solutions

  • Adobe Premiere
  • Adobe Photoshop
  • Solidworks
  • Autodesk AutoCAD
  • AI & Machine Learning

Workstations

  • Media & Entertainment
  • Engineering
  • Scientific PCs
  • More

Support

  • Online Guides
  • Request Support
  • Remote Help

Publications

  • All News
  • Puget Blog
  • HPC Blog
  • Hardware Articles
  • Case Studies

Policies

  • Warranty & Return
  • Terms and Conditions
  • Privacy Policy
  • Delivery Times
  • Accessibility

About Us

  • Testimonials
  • Careers
  • About Us
  • Contact Us
  • Newsletter

© Copyright 2026 - Puget Systems, All Rights Reserved.