Not reproduced by Bitworks
Qwen3.8-27B on dual RX 7900 XTX with tensor split and MTP
- RX 7900 XTX
- 2
- Linux (Docker ROCm)
- Container
- llama.cpp server-rocm b10481 (25ae3a9b3)
- ROCm
- Qwen3.8-27B UD-IQ4_XS
- 131,072 tokens
- 86,675 tokens
- no
- Author also configured 262144; not the occupied depth in these rows.
- 1
- 1
- dual-split (tensor split; PCIe 4.0 x16 CPU plus x4 chipset)
- MTP
- Ryzen 9 9950X
- 96 GB reported; unit not verified
- PCIe 4.0 x16 CPU plus x4 chipset
- 69.25 tok/s — generation speed (per request) · API server · 1 client · input 86,675 tokens
- 361 tok/s — prompt speed (per request) · API server · 1 client · input 86,675 tokens
- UD-IQ4_XS differs from the linked Q8_0 recipe.
- The comparison changes GPU count, split mode, MTP, and configured context; it does not isolate a speedup cause.
- The larger configured window was not filled by the prompt.
Discussion
Share setups, questions and your own results. Posts are reviewed before they appear. Replies from Bitworks staff are marked.