Estimate — sign in for exact rates.

 Trusted GPU & AI infrastructure since 2015 — Huntsville, AL
16 GB GDDR6 · RDNA 3 · gfx1101

AMD Radeon RX 7800 XT 16GB

The RX 7800 XT offers 16 GB of GDDR6 memory. Its candidate Qwen3.5-9B Q8 profile uses the same pinned artifact as several other 16 GB cards, making a future matched comparison possible. The card still needs its own driver, free-memory observation and completed-output evidence; equal memory labels do not establish equal speed.

For a desktop used for both display and inference, record what else occupies GPU memory before testing the 32K catalog context. A different quantization may trade file size against response quality, so compare representative tasks rather than only download size. The larger Qwen3.8-27B profile is a separate experiment, not an established alternative on this card.

AMD specifications

Recipes

Windows

Linux

Community reports — not measured or reproduced by Bitworks

Reported by others, linked to the original. Bitworks has not verified these numbers; they can differ on your system.

Reported on GitHub by HarlonOna — posted 2026-08-09 Original report

Not reproduced by Bitworks

RX 7800 XT Vulkan Qwen3.5-9B field report

Method (as reported): llama-bench -pg 512,128 -r 20; separate server chat with MTP2 and one slot.

  • Reported GPU setup RX 7800 XT 16GB
  • GPU count 1
  • Operating system Linux (Fedora 44, Mesa RADV)
  • Engine llama.cpp
  • Backend Vulkan
  • Model Qwen3.5-9B Q4_K_M
  • Configured context 102,400 tokens
  • Actual input input length not reported
  • Reported context Reported sessions occupied about 20000 to 30000 tokens, not the configured full window.
  • Simultaneous clients 1
  • Server slots 1
  • Workers 1
  • Topology single
  • Speculative decoding MTP2 on server; off in llama-bench
  • CPU Ryzen 5 7500F
  • 71.95 tok/s ± 0.03 (as reported, n = 20) — generation speed (per request) · Synthetic benchmark (not API serving) · 1 client
    llama-bench tg128; MTP off; author reports ±0.03
    This row: configured context: not reported for this row · speculative decoding: off
  • 1700.48 tok/s ± 4.65 (as reported, n = 20) — prompt speed (per request) · Synthetic benchmark (not API serving) · 1 client · input 512 tokens
    llama-bench pp512; MTP off; author reports ±4.65
    This row: configured context: not reported for this row · speculative decoding: off
  • ≈86 tok/s — generation speed (per request) · agent session · 1 client
    Approximate average across ten chat questions with MTP2
    This row: speculative decoding: MTP2

Not reported: system memory, Exact llama.cpp revision, GGUF digest, Driver revision, Physical GPU residency, Raw server timings

  • Configured context was not filled in the reported sessions.
  • Synthetic no-MTP and chat with MTP are different workloads.
  • Author describes a concurrent slowdown but not a controlled aggregate throughput rate.
  • Linux RADV does not qualify native Windows Vulkan.

Setup guide (external)

Reported on GitHub by Niko1221 / Strata maintainers Original report

Not reproduced by Bitworks Project documentation — capability, not a measurement

Strata AMD HIP guide

Method (as reported): Maintainer documents Windows bundled HIP runtime, per-card checks and Linux layer split; no discrete AMD Windows model-load validation.

HEAD seen at retrieval: 6f32ec070f23 (not necessarily the commit the author measured)

  • Reported GPU setup Windows: one GPU per model; Linux: selectable layer split
  • GPU count not reported
  • Operating system Windows and Linux (Windows HIP and Linux HIP)
  • Engine Strata HIP
  • Backend HIP
  • Model Qwen3.8-Flash-Next pack
  • Actual input input length not reported
  • Topology not reported — Windows one GPU per model; Linux selectable layer split

Not reported: model quantisation, configured context, actual input length, simultaneous clients, topology, CPU, system memory, release archive digest and revision, Windows discrete-card load, Windows dual-card execution, quality and memory curves

  • Maintainer says ready-made Windows archive has not run a model on a discrete AMD card in documented validation.
  • Capability documentation, not a measured compatibility or performance result.
  • RX 9060 XT memory variant is unspecified in the guide, so neither variant is mapped.

Setup guide (external)

Reported on Reddit by Specific-Pomelo-5455 Original report

Not reproduced by Bitworks

Corrected RX 7800 XT Qwen3.8 occupancy benchmarks

Method (as reported): Structured server harness; parallel one, FA, q4_0 K/V; four to six repeats per point.

  • Reported GPU setup RX 7800 XT 16GB
  • GPU count 1
  • Operating system Windows (Windows, version not reported)
  • Engine llama.cpp
  • Backend HIP/ROCm
  • Model Qwen3.8-27B Unsloth-Dynamic-derived IQ4_XS mixed tensors
  • Configured context 65,536 tokens
  • Actual input input length not reported
  • Reported context One row occupied about 7000 tokens; the other occupied about 48000.
  • Simultaneous clients 1
  • Server slots 1
  • Workers 1
  • Topology single
  • Speculative decoding MTP off for context-depth comparison
  • 21.85 tok/s — generation speed (per request) · API server · 1 client
    About 7000 occupied tokens at configured 65536; q4_0 K/V
    This row: speculative decoding: off
  • 8.77 tok/s — generation speed (per request) · API server · 1 client
    About 48000 occupied tokens at configured 65536; q4_0 K/V
    This row: speculative decoding: off
  • 22.15 tok/s — generation speed (per request) · API server · 1 client
    Compared with 18.62 tok/s (MTP on versus off for one long generation); reported difference: speculative decoding
    Separate reported comparison at configured 32K; not the context-depth pair
    This row: configured context: 32,768 tokens · speculative decoding: MTP on

Not reported: Exact GGUF digest, llama.cpp build, Driver revision, CPU and RAM, Physical WDDM residency

  • The configured window is not the occupied input.
  • Earlier author measurements were corrected; use these corrected points only.
  • The default fit setting had silently moved layers to CPU in some runs.
  • MTP worsened the author's multi-step agent wall time; the long-generation comparison is not an agent guarantee.

Setup guide (external)

No approved community measurement is shown here for:

  • two cards

Discussion

Share setups, questions and your own results. Posts are reviewed before they appear. Replies from Bitworks staff are marked.

Log in Create an account to join the discussion.

Bitworks listings for this card

Select the fields to be shown. Others will be hidden. Drag and drop to rearrange the order.
  • Image
  • SKU
  • Rating
  • Price
  • Stock
  • Availability
  • Add to cart
  • Description
  • Content
  • Weight
  • Dimensions
  • Additional information
Click outside to hide the comparison bar
Compare
Compare ×
Let's Compare! Continue shopping