Estimate — sign in for exact rates.

 Trusted GPU & AI infrastructure since 2015 — Huntsville, AL
16 GB GDDR6 · RDNA 4 · gfx1201

AMD Radeon RX 9070 XT 16GB

The RX 9070 XT also advertises 16 GB of GDDR6 memory, but it is not interchangeable with the non-XT RX 9070 or the RDNA 3 RX 7800 XT. The same pinned Qwen3.5-9B Q8 candidate across those pages gives a way to isolate hardware differences in later matched tests, provided driver, prompts, context and active requests are recorded.

This candidate’s 32K context and full-GPU placement are requested settings, not observed fit. Check current free memory and sustained behavior, not only a successful file download or model-load message. Windows and Linux require their own evidence, and no card-specific result is published here.

AMD specifications

Recipes

Windows

Linux

Community reports — not measured or reproduced by Bitworks

Reported by others, linked to the original. Bitworks has not verified these numbers; they can differ on your system.

Reported on Blog by Ataa Aldaghstani / ataadevs — posted 2026-08-17 Original report

Not reproduced by Bitworks

RX 9070 XT Windows Qwen3.8 coding-task MTP comparison

Method (as reported): Same repository-explanation task; q4_0 K/V, FA on, fit off; author averages with timing denominator unspecified.

  • Reported GPU setup RX 9070 XT 16GB
  • GPU count 1
  • Operating system Windows (Windows)
  • Engine llama.cpp
  • Backend ROCm/HIP
  • Model Qwen3.8-27B Q4_K_M and AD-IQ4_XS-IQ3_S (separate arms)
  • Actual input input length not reported
  • Full window tested no
  • Reported context 32,096 Q4 arm; 64,096 mixed-quant arm; occupied input unreported.
  • Simultaneous clients 1
  • Server slots 1
  • Topology single
  • Speculative decoding MTP2
  • CPU Ryzen 7 7800X3D
  • Memory configuration DDR5; capacity unknown
  • Host notes Display on another GPU.
  • 24 tok/s — other (scope not stated) · agent session · 1 client
    Compared with 17 tok/s (Q4_K_M no MTP vs MTP2); reported difference: speculative decoding
    Author average; timing denominator unspecified; Q4_K_M at configured 32,096; repository-explanation task.
    This row: configured context: 32,096 tokens · quantisation: Unsloth Q4_K_M · speculative decoding: MTP2
  • 42 tok/s — other (scope not stated) · agent session · 1 client
    Compared with 27.5 tok/s (AtomicChat mixed quant no MTP vs MTP2); reported difference: speculative decoding
    Author average; timing denominator unspecified; mixed quant at configured 64,096; repository-explanation task.
    This row: configured context: 64,096 tokens · quantisation: AtomicChat AD-IQ4_XS-IQ3_S · speculative decoding: MTP2

Not reported: model file checksum, system memory, Artifact/build digests, Occupied inputs and output counts, Timing denominator, Repetitions/variance, Objective task scoring

  • Q4 arms required mid-task compaction.
  • Mixed-quant MTP quality was judged weaker subjectively.
  • Cross-quant comparison also changes context.
  • Linked setup advice is not adopted or qualified.

Setup guide (external)

Reported on GitHub by 7269827-rgb Original report

Not reproduced by Bitworks

RX 9070 XT custom Qwen3.8 long-context state follow-up

Method (as reported): Post-publication shallow-context empty-/loaded-cache checks; q4_0 K/V and MTP2; state UNCLASSIFIED.

HEAD seen at retrieval: c7abad32e4f2 (not necessarily the commit the author measured)

  • Reported GPU setup RX 9070 XT 16GB
  • GPU count 1
  • Operating system Windows (Windows)
  • Engine patched llama.cpp bf0040e15 lineage; exact archived patch missing
  • Backend Vulkan
  • Model Qwen3.8-27B custom 2.72 BPW tensor mix (7269827-rgb)
  • Configured context 262,144 tokens
  • Actual input input length not reported
  • Full window tested no
  • Reported context 256K is capacity; reported samples were shallow-context, not full-window.
  • Simultaneous clients 1
  • Topology single
  • Speculative decoding MTP2
  • 24 tok/s — generation speed (scope not stated) · 1 client
    Shallow-context empty-cache sample 1/3; state UNCLASSIFIED; 256K configured capacity not occupied.
    This row: configured context: 262,144 tokens · quantisation: custom 2.72 BPW tensor mix · speculative decoding: MTP2
  • 23 tok/s — generation speed (scope not stated) · 1 client
    Shallow-context empty-cache sample 2/3; state UNCLASSIFIED; 256K configured capacity not occupied.
    This row: configured context: 262,144 tokens · quantisation: custom 2.72 BPW tensor mix · speculative decoding: MTP2
  • 25.8 tok/s — generation speed (scope not stated) · 1 client
    Shallow-context empty-cache sample 3/3; state UNCLASSIFIED; 256K configured capacity not occupied.
    This row: configured context: 262,144 tokens · quantisation: custom 2.72 BPW tensor mix · speculative decoding: MTP2
  • 26 tok/s — generation speed (scope not stated) · 1 client
    Shallow-context loaded-cache sample 1/3; state UNCLASSIFIED; 256K configured capacity not occupied.
    This row: configured context: 262,144 tokens · quantisation: custom 2.72 BPW tensor mix · speculative decoding: MTP2
  • 29.3 tok/s — generation speed (scope not stated) · 1 client
    Shallow-context loaded-cache sample 2/3; state UNCLASSIFIED; 256K configured capacity not occupied.
    This row: configured context: 262,144 tokens · quantisation: custom 2.72 BPW tensor mix · speculative decoding: MTP2
  • 26.9 tok/s — generation speed (scope not stated) · 1 client
    Shallow-context loaded-cache sample 3/3; state UNCLASSIFIED; 256K configured capacity not occupied.
    This row: configured context: 262,144 tokens · quantisation: custom 2.72 BPW tensor mix · speculative decoding: MTP2

Not reported: CPU, system memory, Cause of unreproduced fast state, Exact original patch, General quality across tasks, Concurrency

  • Faster campaign headlines did not reproduce reliably.
  • 40K gate speeds used a different binary.
  • 200K sustained figure is a per-problem proxy, not a timed floor.
  • Do not transfer custom quant/attention quality to stock artifacts.

Setup guide (external)

Related link (external)

Related link (external)

Related link (external)

Reported on GitHub by MrLordCat · measured 2026-09-26 Original report

Not reproduced by Bitworks

Dual RX 9070 XT Qwen3.8 Windows custom-fork MTP report

Method (as reported): One cold run per cell; FP8 K/V, FA, batch 8192 / ubatch 1024; no warmup or cache reuse.

HEAD seen at retrieval: e648079e8656 (not necessarily the commit the author measured)

  • Reported GPU setup 2 × RX 9070 XT 16GB
  • GPU count 2
  • Operating system Windows (Windows 11)
  • Engine llama.cpp-rdna-lab fea1c3180 plus uncommitted FA change
  • Backend ROCm/HIP 7.2
  • Model Qwen3.8-27B UD-Q4_K_M
  • Configured context 49,152 tokens
  • Actual input 32,996 tokens
  • Output 256 tokens
  • Full window tested no
  • Reported context One cold run per cell; configured window not fully occupied.
  • Simultaneous clients 1
  • Server slots 1
  • Topology dual-split (ROCm1,ROCm0; layer split 1:1; no P2P; PCIe placement unknown.)
  • Speculative decoding MTP n3
  • CPU Ryzen 7 5800X3D
  • System memory 64 GiB
  • 43.75 tok/s — generation speed (per request) · API server · 1 client · input 32,996 tokens
    Compared with 25.74 tok/s (Same custom fork no MTP vs MTP n3); reported difference: speculative decoding
    Single cold L2 Windows HIP cell; 32,996 input / 256 output; custom fork and uncommitted FA change.
    This row: configured context: 49,152 tokens · speculative decoding: MTP n3
  • 1799.29 tok/s — prompt speed (per request) · API server · 1 client · input 32,996 tokens
    Compared with 1879.06 tok/s (Same custom fork no MTP vs MTP n3); reported difference: speculative decoding
    Same one-run L2 Windows HIP comparison; reported prompt processing declined.
    This row: configured context: 49,152 tokens · speculative decoding: MTP n3

Not reported: model file checksum, Exact executable/artifact hashes, Uncommitted change identity, Windows raw receipts unavailable in pinned tree, Repeated-run variance, Concurrent throughput

  • Specialized fork, not stock FastLLM.
  • Do not infer dual-card scaling without a single-card control.
  • Windows/Linux comparisons also change driver and P2P.
  • W24/W26 kernel gains belong to separate Linux research lanes.

Setup guide (external)

Related link (external)

Reported on GitHub by kr4ckhe4d · measured 2026-10-02 Original report

Not reproduced by Bitworks

Qwen3.8 llama.cpp build comparison on RX 9070 XT

Method (as reported): ABBA build order; two server loads per build, two deep measures/load plus discarded warmup; Q8 KV, batch 2048, microbatch 1024, FA on.

HEAD seen at retrieval: c18f1bc96ce3 (not necessarily the commit the author measured)

  • Reported GPU setup RX 9070 XT
  • GPU count 1
  • Operating system Linux (CachyOS)
  • Engine llama.cpp HIP build 11345 (a868c3e3c)
  • Backend ROCm
  • Model Qwen3.8-27B IQ4_XS-v3
  • Configured context 32,768 tokens
  • Actual input 16,852 tokens
  • Output 700 tokens
  • Full window tested no
  • Simultaneous clients 1
  • Server slots 1
  • Workers 1
  • Topology single
  • Prefix cache off
  • CPU Ryzen 7 9800X3D
  • Memory configuration 2 x 16 GiB DDR5-6000
  • System memory 32 GiB
  • Memory speed 6000 MT/s
  • 28.41 tok/s — generation speed (per request) · 1 client · input 16,852 tokens
    Compared with 26.84 tok/s (baseline build 7c35571e5; reported a868c3e3c; no MTP); reported difference: engine build
    Deep-input build comparison; Q8_0 KV, flash attention on.
    This row: configured context: 32,768 tokens
  • 1193 tok/s — prompt speed (per request) · 1 client · input 16,852 tokens
    Compared with 959 tok/s (build 7c35571e5 vs a868c3e3c; no MTP); reported difference: engine build
    Same source comparison; prompt cache off.
    This row: configured context: 32,768 tokens

Not reported: exact model digest, serving driver identity, concurrent behavior

  • Single Linux host; not Windows or non-XT RX 9070 evidence.
  • The source's Q3 MTP row uses a different quant and is not included.
  • Build changes were not bisected to an individual kernel.

Related link (external)

Reported on GitHub by Niko1221 / Strata maintainers Original report

Not reproduced by Bitworks Project documentation — capability, not a measurement

Strata AMD HIP guide

Method (as reported): Maintainer documents Windows bundled HIP runtime, per-card checks and Linux layer split; no discrete AMD Windows model-load validation.

HEAD seen at retrieval: 6f32ec070f23 (not necessarily the commit the author measured)

  • Reported GPU setup Windows: one GPU per model; Linux: selectable layer split
  • GPU count not reported
  • Operating system Windows and Linux (Windows HIP and Linux HIP)
  • Engine Strata HIP
  • Backend HIP
  • Model Qwen3.8-Flash-Next pack
  • Actual input input length not reported
  • Topology not reported — Windows one GPU per model; Linux selectable layer split

Not reported: model quantisation, configured context, actual input length, simultaneous clients, topology, CPU, system memory, release archive digest and revision, Windows discrete-card load, Windows dual-card execution, quality and memory curves

  • Maintainer says ready-made Windows archive has not run a model on a discrete AMD card in documented validation.
  • Capability documentation, not a measured compatibility or performance result.
  • RX 9060 XT memory variant is unspecified in the guide, so neither variant is mapped.

Setup guide (external)

Discussion

Share setups, questions and your own results. Posts are reviewed before they appear. Replies from Bitworks staff are marked.

Log in Create an account to join the discussion.
Select the fields to be shown. Others will be hidden. Drag and drop to rearrange the order.
  • Image
  • SKU
  • Rating
  • Price
  • Stock
  • Availability
  • Add to cart
  • Description
  • Content
  • Weight
  • Dimensions
  • Additional information
Click outside to hide the comparison bar
Compare
Compare ×
Let's Compare! Continue shopping