Not reproduced by Bitworks
RX 7800 XT Vulkan Qwen3.5-9B field report
- RX 7800 XT 16GB
- 1
- Linux (Fedora 44, Mesa RADV)
- llama.cpp
- Vulkan
- Qwen3.5-9B Q4_K_M
- 102,400 tokens
- input length not reported
- Reported sessions occupied about 20000 to 30000 tokens, not the configured full window.
- 1
- 1
- 1
- single
- MTP2 on server; off in llama-bench
- Ryzen 5 7500F
- 71.95 tok/s ± 0.03 (as reported, n = 20) — generation speed (per request) · Synthetic benchmark (not API serving) · 1 client
- 1700.48 tok/s ± 4.65 (as reported, n = 20) — prompt speed (per request) · Synthetic benchmark (not API serving) · 1 client · input 512 tokens
- ≈86 tok/s — generation speed (per request) · agent session · 1 client
- Configured context was not filled in the reported sessions.
- Synthetic no-MTP and chat with MTP are different workloads.
- Author describes a concurrent slowdown but not a controlled aggregate throughput rate.
- Linux RADV does not qualify native Windows Vulkan.
Discussion
Share setups, questions and your own results. Posts are reviewed before they appear. Replies from Bitworks staff are marked.