Estimate — sign in for exact rates.

 Trusted GPU & AI infrastructure since 2015 — Huntsville, AL
16 GB GDDR6 · RDNA 4 · gfx1200

AMD Radeon RX 9060 XT 16GB

The 16 GB RX 9060 XT is a separate memory variant from the 8 GB card. Its proposed Qwen3.5-9B Q8 artifact and 32K catalog context ask a different question from the 8 GB Q4 profile: whether a larger quantization leaves enough room for context, compute buffers and normal desktop use. The answer requires a card-specific test.

If comparing variants, keep the prompts and output target fixed, then disclose that the quantization differs. Equal model family names do not mean identical memory use or answer behavior. This page has no approved load, fit, quality or speed result, and the Linux candidate is not evidence for Windows or vice versa.

AMD specifications

Recipes

Windows

Linux

Community reports — not measured or reproduced by Bitworks

Reported by others, linked to the original. Bitworks has not verified these numbers; they can differ on your system.

Reported on GitHub by parsapp Original report

Not reproduced by Bitworks

RX 9060 XT 16GB llama.cpp Vulkan and ROCm benchmarks

Method (as reported): Three full matrix runs; llama-bench pp512/tg128 with default repetitions, FA and full offload.

HEAD seen at retrieval: 41cb2ed3b7c1 (not necessarily the commit the author measured)

  • Reported GPU setup RX 9060 XT 16GB
  • GPU count 1
  • Operating system Linux (CachyOS kernel 7.1.3-2)
  • Engine llama.cpp b9957 (c4ae9a88f8)
  • Backend HIP (ROCm 7.2.4) versus RADV Vulkan
  • Driver RADV from Mesa 26.1.4
  • Model Qwen3.5-9B Q4_K_M
  • Actual input input length not reported
  • Reported context Only short synthetic prompt and generation tests were reported.
  • Simultaneous clients 1
  • Workers 1
  • Topology single
  • CPU Intel Core i5-12400F
  • Memory configuration 32 GB DDR4 RAM
  • 1904.16 tok/s ± 1.31 (as reported) — prompt speed (per request) · Synthetic benchmark (not API serving) · 1 client · input 512 tokens
    Compared with 1845.34 tok/s ± 43.26 (as reported) (RADV Vulkan versus HIP, pp512); reported difference: other
    Full offload and FA; author reports ±1.31 versus ±43.26
  • 50.15 tok/s ± 0.39 (as reported) — generation speed (per request) · Synthetic benchmark (not API serving) · 1 client
    Compared with 49.36 tok/s ± 0.06 (as reported) (RADV Vulkan versus HIP, tg128); reported difference: other
    Full offload and FA; author reports ±0.39 versus ±0.06

Not reported: system memory, GGUF digest, Configured context, Long-prompt behavior, API serving, Concurrent clients, Native Windows rate

  • This is the 16GB variant, not the 8GB XT or non-XT card.
  • Short synthetic benchmarks do not establish long-context fit or agent turn time.
  • Author notes asserts-enabled package build could affect absolute speed.
  • Linux RADV performance does not transfer to Windows Vulkan.

Setup guide (external)

Reported on Reddit by L3G10N78 Original report

Not reproduced by Bitworks

RX 9060 XT 16GB Windows Qwen3.8 field report

Method (as reported): Author reports build, GPU/driver, full offload, FA and one slot; exact benchmark harness not stated.

  • Reported GPU setup RX 9060 XT 16GB
  • GPU count 1
  • Operating system Windows (Windows, version not reported)
  • Engine llama.cpp b10167 (ee3d1b54c)
  • Backend Vulkan
  • Driver Adrenalin 26.8.1
  • Model Qwen3.8-27B UD-IQ3_XXS
  • Configured context 32,768 tokens
  • Actual input input length not reported
  • Reported context Author labels additional 64K tests without establishing the changed launch; only 32K row included.
  • Simultaneous clients 1
  • Server slots 1
  • Workers 1
  • Topology single
  • CPU Ryzen 9 5950X
  • Memory configuration 64 GB DDR4-3200 CL14 RAM
  • Board MAG B550 TOMAHAWK MAX WIFI
  • PCIe PCIe 4.0 x16
  • Host notes Separate W5500 used for embedding/reranking, not main model
  • ≈21.4 tok/s — generation speed (per request) · 1 client
    Author's labelled 32K row; actual occupied prompt not stated
    This row: configured context: 32,768 tokens
  • ≈148.9 tok/s — prompt speed (per request) · 1 client
    Author's labelled 32K row; actual occupied prompt not stated
    This row: configured context: 32,768 tokens

Not reported: system memory, GGUF digest, Actual occupied prompt, KV cache types, Repetitions and variance, Physical VRAM residency

  • The source's 64K-labelled rows lack a corresponding documented launch and are excluded.
  • The separate W5500 served embedding and reranking, not dual-GPU model parallelism.
  • The author had not completed an agent-harness comparison.
  • Qwen3.8 IQ3 is not the candidate site's Qwen3.5 Q8 recipe.

Setup guide (external)

No approved community measurement is shown here for:

  • two cards

Discussion

Share setups, questions and your own results. Posts are reviewed before they appear. Replies from Bitworks staff are marked.

Log in Create an account to join the discussion.
Select the fields to be shown. Others will be hidden. Drag and drop to rearrange the order.
  • Image
  • SKU
  • Rating
  • Price
  • Stock
  • Availability
  • Add to cart
  • Description
  • Content
  • Weight
  • Dimensions
  • Additional information
Click outside to hide the comparison bar
Compare
Compare ×
Let's Compare! Continue shopping