Estimate — sign in for exact rates.

 Trusted GPU & AI infrastructure since 2015 — Huntsville, AL
32 GB GDDR6 · RDNA 4 · gfx1201

AMD Radeon AI PRO R9700 32GB

The Radeon AI PRO R9700 has 32 GB of GDDR6 memory, but a running model sees only the portion currently free. The proposed single-card Qwen3.8-27B Q6 profile asks whether that memory tier can support a larger quantization and 32K catalog context without sacrificing useful response time. A separate Q8 candidate uses a shorter 16K context; the difference is a catalog choice, not a tested card limit.

If you have two R9700s, decide whether you need two independent requests or a larger model split across cards. Those need different launchers and measurements. Check both cards’ actual PCIe links, power connections and airflow before considering a pair. Flash-Next with host-memory offload remains separate research, not a recipe shown here.

AMD specifications

Recipes

Windows

Linux

Community reports — not measured or reproduced by Bitworks

Reported by others, linked to the original. Bitworks has not verified these numbers; they can differ on your system.

Related reports from the same author or project — not independent replication

Reported on Blog by ElioVP / Eliovp-BV — posted 2026-09-16 Original report

Not reproduced by Bitworks

Paiton Qwen3.8 MXFP4/DFlash2 R9700 release concurrency

Method (as reported): One BetterBench 0.6.0 validation run, 290 successes including warmups, 48 requests per concurrency level.

HEAD seen at retrieval: 8f56157c05eb (not necessarily the commit the author measured)

  • Reported GPU setup Radeon AI PRO R9700 32GB
  • GPU count 1
  • Operating system Linux (Linux x86-64)
  • Environment Container
  • Engine vLLM 0.29.0 with Paiton plugin 65K ROCm10 release image
  • Backend ROCm 10
  • Model Qwen3.8-27B Unsloth NVFP4 checkpoint via MXFP4 path (Unsloth)
  • Configured context 65,536 tokens
  • Actual input input length not reported
  • Full window tested no
  • Reported context Concurrent prompt occupancy unreported; separate prefill sweep reached median 47,016.5.
  • Server slots 8
  • Topology single
  • Prefix cache off
  • Speculative decoding DFlash2
  • 115 tok/s — generation speed (aggregate across clients) · API server · 1 client
    C1; completed output/common full-workload wall; 48 requests per level; prompt occupancy unreported.
    This row: configured context: 65,536 tokens · engine: vLLM 0.29.0 + Paiton · speculative decoding: DFlash2
  • 203.2 tok/s — generation speed (aggregate across clients) · API server · 2 clients
    C2; completed output/common full-workload wall; 48 requests per level; prompt occupancy unreported.
    This row: configured context: 65,536 tokens · engine: vLLM 0.29.0 + Paiton · speculative decoding: DFlash2
  • 296.5 tok/s — generation speed (aggregate across clients) · API server · 4 clients
    C4; completed output/common full-workload wall; 48 requests per level; prompt occupancy unreported.
    This row: configured context: 65,536 tokens · engine: vLLM 0.29.0 + Paiton · speculative decoding: DFlash2
  • 400.7 tok/s — generation speed (aggregate across clients) · API server · 8 clients
    C8; completed output/common full-workload wall; 48 requests per level; prompt occupancy unreported.
    This row: configured context: 65,536 tokens · engine: vLLM 0.29.0 + Paiton · speculative decoding: DFlash2

Not reported: model file checksum, CPU, system memory, concurrent prompt occupancy and per-client rates, independent replication on Bitworks hardware

  • Same Paiton project/author as the original register's long-KV4 entry; do not count as independent corroboration.
  • Different checkpoint, runtime, sampling and workload from Paiton's earlier 8K matched Radiance comparison; no cross-release speedup percentage follows.
  • 65K is configured capacity, not eight simultaneous full-65K conversations; Linux result is not native Windows.

Setup guide (external)

Related link (external)

Reported on Reddit by evp-cloud / Eliovp-BV Original report

Not reproduced by Bitworks

Qwen3.8-27B on one R9700 with long prefix cache

Method (as reported): Project author reports BetterBench 20-pass runs and a repeated-prefix, 20-turn agent session.

HEAD seen at retrieval: 918be25e10f1 (not necessarily the commit the author measured)

  • Reported GPU setup Radeon AI PRO R9700
  • GPU count 1
  • Operating system Linux (x86-64 Docker ROCm)
  • Environment Container
  • Engine stock vLLM plus Paiton plugin
  • Backend ROCm
  • Model Qwen3.8-27B project 3-bit mode
  • Configured context 262,144 tokens
  • Actual input input length not reported
  • Full window tested no
  • Reported context Agent session grew from about 50K to 253K; exact input per metric unknown.
  • Workers 1
  • Topology single
  • 173.4 tok/s — generation speed (per request) · 1 client
    Compared with 174.3 tok/s (baseline prior mode; reported long-kv4); reported difference: other
    Single-stream decode comparison; cache conditions not fully specified.
    This row: configured context: 262,144 tokens
  • 478.5 tok/s — generation speed (aggregate across clients) · 8 clients
    Compared with 462.4 tok/s (prior mode vs long-kv4); reported difference: other
    Aggregate across eight requests; not single-client decode speed.
    This row: configured context: 262,144 tokens
  • 204 s — task time (scope not stated) · agent session · 1 client
    Compared with 1237 s (baseline prior mode; reported warm repeated-prefix agent task); reported difference: prefix cache
    Total for a 20-turn warm repeated-prefix session, not one request; benefit does not isolate decode.
    This row: configured context: 262,144 tokens · prefix cache: on
  • 569,878 tokens — reusable prefix cache (server total) (server-wide)
    Total reusable prefix tokens, not a per-request window.
    This row: configured context: 262,144 tokens · prefix cache: on

Not reported: engine build, CPU, system memory, container digest, model digest, independent replication, Bitworks workload eviction behavior

  • Warm prefix reuse explains task-time change; it is not decode acceleration.
  • Cold first read remains expensive; short or cold prompts can be slower.
  • Linux Docker path is not native Windows.

Setup guide (external)

Related link (external)

Reported on GitHub by JohnTDI-cpu Original report

Not reproduced by Bitworks

Direct P2P AllReduce on dual R9700

Method (as reported): llama-bench -sm tensor -fa on -n 128 -r 5; author compared perplexity; prefill remained on stock butterfly path.

HEAD seen at retrieval: f6f66106f2a9 (not necessarily the commit the author measured)

  • Reported GPU setup Radeon AI PRO R9700
  • GPU count 2
  • Operating system Linux (ROCm 7.2.4)
  • Engine llama.cpp HIP tensor plus direct-P2P patch latest master at author test; commit unknown
  • Backend ROCm
  • Model Qwen3-27B Q8_0
  • Actual input input length not reported
  • Simultaneous clients 1
  • Workers 1
  • Topology dual-split (tensor split; peer access available, link layout unknown)
  • PCIe peer access reported; layout unknown
  • 30.43 tok/s — generation speed (per request) · Synthetic benchmark (not API serving) · 1 client
    Compared with 28.3 tok/s (baseline stock tensor; reported direct-P2P patch); reported difference: engine build
    Synthetic llama-bench tg128; five repetitions, not API serving.
    This row: topology: dual-split

Not reported: configured context, actual input length, CPU, system memory, tested engine commit, artifact digest, PCIe topology, API concurrency

  • Model is Qwen3-27B, not Qwen3.8-27B.
  • One-machine synthetic decode does not establish API serving throughput.
  • Patch depends on peer access and was not tested on Windows.

Setup guide (external)

Reported on GitHub by bkvargyas Original report

Not reproduced by Bitworks

Dual R9700 vLLM Proxmox recipe and benchmarks

Method (as reported): BetterBench 0.4.0 on two passthrough cards with shared emulated PCIe switch and repaired RCCL/P2P path.

HEAD seen at retrieval: e2a38bf4ab5d (not necessarily the commit the author measured)

  • Reported GPU setup Radeon AI PRO R9700
  • GPU count 2
  • Operating system Linux (Proxmox 9.2 VM; Debian 13 guest)
  • Environment VM guest
  • Engine vLLM plus DFlash2 FP8 and Radiance vLLM 0.28; BetterBench 0.4.0
  • Backend ROCm
  • Model Qwen3.8-27B MXFP4-W4A8
  • Actual input input length not reported
  • Reported context Source spans multiple depths; per-metric context differs or is unreported.
  • Workers 1
  • Topology dual-split (PCIe 5.0 x16 each; VM P2P path)
  • CPU EPYC 7543
  • PCIe PCIe 5.0 x16 each with VM P2P plumbing
  • 159.7 tok/s — other (per request) · 1 client
    Compared with 106.2 tok/s (baseline FP8+DFlash2; reported MXFP4-W4A8+DFlash2-FP8); changed between the two runs: quantisation, engine
    BetterBench weighted single-stream generation; weight and kernel paths differ.
    This row: quantisation: MXFP4-W4A8 · engine: vLLM plus DFlash2 FP8 and Radiance
  • 495 tok/s — generation speed (aggregate across clients) · 8 clients
    Aggregate across eight simultaneous streams; not per-request generation.
    This row: quantisation: MXFP4-W4A8
  • ≈5153 tok/s — prompt speed (per request) · 1 client
    Approximate input depth around 16K; not exact occupancy.
    This row: quantisation: MXFP4-W4A8

Not reported: configured context, system memory, model artifact digest, bare-metal behavior, Windows behavior, per-request context in summary

  • FP8 and MXFP4 comparison changes weights, kernels, and runtime path.
  • Reported total KV cache capacity is not per-request tested context.
  • VM-specific P2P workaround is not a general installer instruction.

Setup guide (external)

Reported on GitHub by charlie12345 · measured 2026-08-26 Original report

Not reproduced by Bitworks

Native Windows dual R9700 custom-vLLM reference concurrency

Method (as reported): OpenAI-compatible serving; warmup wave per level, forced 128 output, no prefix cache; TP1 one GPU vs TP2 RCCL two GPUs.

HEAD seen at retrieval: 9c9087be48a5 (not necessarily the commit the author measured)

  • Reported GPU setup 2 × Radeon AI PRO R9700 32GB
  • GPU count 2
  • Operating system Windows (Windows 11)
  • Engine custom native-Windows vLLM fork prior reference build; current RC2 TP2 requalification pending
  • Backend ROCm Direct-RCCL
  • Model Qwen3.8-27B abihsoro AWQ INT4 W4A16 group-128 symmetric (abihsoro)
  • Configured context 512 tokens
  • Actual input 32 tokens
  • Output 128 tokens
  • Full window tested no
  • Reported context Short 32-input/128-output workload; not long-context agents.
  • Server slots 32
  • Topology dual-split (Reported arm TP2 RCCL on two GPUs; baseline TP1 on one GPU.)
  • Prefix cache off
  • 13.65 tok/s — generation speed (aggregate across clients) · API server · 1 client · input 32 tokens
    Compared with 21.31 tok/s (TP1 one R9700 vs TP2 RCCL two R9700); changed between the two runs: GPU count, split mode
    C1; 8 requests/arm; completed output/common full-workload wall; TP2 RCCL two GPUs.
    This row: configured context: 512 tokens · topology: dual-split · cards: 2
  • 69.52 tok/s — generation speed (aggregate across clients) · API server · 8 clients · input 32 tokens
    Compared with 66.65 tok/s (TP1 one R9700 vs TP2 RCCL two R9700); changed between the two runs: GPU count, split mode
    C8; 16 requests/arm; completed output/common full-workload wall; TP2 RCCL two GPUs.
    This row: configured context: 512 tokens · topology: dual-split · cards: 2
  • 179.73 tok/s — generation speed (aggregate across clients) · API server · 32 clients · input 32 tokens
    Compared with 170.48 tok/s (TP1 one R9700 vs TP2 RCCL two R9700); changed between the two runs: GPU count, split mode
    C32; 64 requests/arm; completed output/common full-workload wall; TP2 RCCL two GPUs.
    This row: configured context: 512 tokens · topology: dual-split · cards: 2
  • 1067.01 ms — time to first token (per request) · API server · 32 clients · input 32 tokens
    Compared with 780.45 ms (TP1 one R9700 vs TP2 RCCL two R9700); changed between the two runs: GPU count, split mode
    C32 mean TTFT across 64 requests/arm; not p95; TP2 RCCL two GPUs.
    This row: configured context: 512 tokens · topology: dual-split · cards: 2

Not reported: CPU, system memory, exact reference-run binary digest and full machine topology in this summary, performance at useful agent contexts or with mixed long prompts, performance on RX 7900 XT, RX 7900 XTX or current RC2 runtime

  • Fresh TP2 qualification of final RC2 vLLM/ROCm 10 commits is pending after a power interruption. These are prior reference measurements, not a current RC2 benchmark.
  • Only dual-R9700 gfx1201 was historically end-to-end exercised; gfx1100 is build-selectable but unvalidated. No dual-XT recipe or performance claim follows.
  • 512 configured context and 32-token inputs do not represent long-context subagent workloads; TP2 gains only 5.4% aggregate at C32 and has higher mean TTFT in this row.
  • Custom native-Windows fork/transport is not FastLLM's pinned llama.cpp Vulkan lane.

Setup guide (external)

Related link (external)

Reported on Blog by Umberto Breglia / Lucebox Original report

Not reproduced by Bitworks

R9700 Lucebox continuous-batching comparison

Method (as reported): Ten synchronized HumanEval waves/concurrency, median of three fresh-server repetitions after warmup; output/common full request wall.

  • Reported GPU setup Radeon AI PRO R9700 32GB
  • GPU count 1
  • Operating system OS not reported — GPU-only evidence
  • Engine Lucebox HIP and separate llama.cpp
  • Backend ROCm 7.2.4
  • Model Qwen3.8-27B IQ4_XS target; Q8_0 DFlash2 draft
  • Configured context 16,384 tokens
  • Actual input input length not reported
  • Output 256 tokens
  • Full window tested no
  • Reported context HumanEval short prompts; separate retrieval depth 1,046–15,695.
  • Server slots 6
  • Topology single
  • Prefix cache off
  • Speculative decoding DFlash2
  • 106.6 tok/s — generation speed (aggregate across clients) · API server · 1 client
    Compared with 75.2 tok/s (Separate llama.cpp run vs Lucebox HIP); reported difference: engine
    C1 HumanEval; total completed output / common full-request wall, not isolated decode.
    This row: configured context: 16,384 tokens · engine: Lucebox HIP · speculative decoding: DFlash2
  • 300.9 tok/s — generation speed (aggregate across clients) · API server · 5 clients
    Compared with 193 tok/s (Separate llama.cpp run vs Lucebox HIP); reported difference: engine
    C5 HumanEval; total completed output / common full-request wall, not per-client decode.
    This row: configured context: 16,384 tokens · engine: Lucebox HIP · speculative decoding: DFlash2
  • 6.2 tok/s — generation speed (aggregate across clients) · API server · 5 clients · input 15,695 tokens
    C5 separate 15,695-token retrieval prompt; 128 output/request; median TTFT 100.0 s.
    This row: configured context: 16,384 tokens · engine: Lucebox HIP · speculative decoding: DFlash2

Not reported: operating system, engine build, CPU, system memory, exact OS distribution and model artifact digest, raw requests/prompt fixtures (author says retained on measurement host), performance under mixed/random arrivals and cancellation

  • ROCm R9700 result, not native Windows or RX 7900 XT guidance.
  • Lucebox and llama.cpp were separate runs; no single-resident paired process.
  • Short coding-prompt scaling does not predict long-context throughput; the long-prompt result is a separate workload.

Setup guide (external)

Related link (external)

Reported on GitHub by Niko1221 / Strata maintainers Original report

Not reproduced by Bitworks Project documentation — capability, not a measurement

Strata AMD HIP guide

Method (as reported): Maintainer documents Windows bundled HIP runtime, per-card checks and Linux layer split; no discrete AMD Windows model-load validation.

HEAD seen at retrieval: 6f32ec070f23 (not necessarily the commit the author measured)

  • Reported GPU setup Windows: one GPU per model; Linux: selectable layer split
  • GPU count not reported
  • Operating system Windows and Linux (Windows HIP and Linux HIP)
  • Engine Strata HIP
  • Backend HIP
  • Model Qwen3.8-Flash-Next pack
  • Actual input input length not reported
  • Topology not reported — Windows one GPU per model; Linux selectable layer split

Not reported: model quantisation, configured context, actual input length, simultaneous clients, topology, CPU, system memory, release archive digest and revision, Windows discrete-card load, Windows dual-card execution, quality and memory curves

  • Maintainer says ready-made Windows archive has not run a model on a discrete AMD card in documented validation.
  • Capability documentation, not a measured compatibility or performance result.
  • RX 9060 XT memory variant is unspecified in the guide, so neither variant is mapped.

Setup guide (external)

Discussion

Share setups, questions and your own results. Posts are reviewed before they appear. Replies from Bitworks staff are marked.

Log in Create an account to join the discussion.
Select the fields to be shown. Others will be hidden. Drag and drop to rearrange the order.
  • Image
  • SKU
  • Rating
  • Price
  • Stock
  • Availability
  • Add to cart
  • Description
  • Content
  • Weight
  • Dimensions
  • Additional information
Click outside to hide the comparison bar
Compare
Compare ×
Let's Compare! Continue shopping