Not reproduced by Bitworks
Paiton Qwen3.8 MXFP4/DFlash2 R9700 release concurrency
- Radeon AI PRO R9700 32GB
- 1
- Linux (Linux x86-64)
- Container
- vLLM 0.29.0 with Paiton plugin 65K ROCm10 release image
- ROCm 10
- Qwen3.8-27B Unsloth NVFP4 checkpoint via MXFP4 path (Unsloth)
- 65,536 tokens
- input length not reported
- no
- Concurrent prompt occupancy unreported; separate prefill sweep reached median 47,016.5.
- 8
- single
- off
- DFlash2
- 115 tok/s — generation speed (aggregate across clients) · API server · 1 client
- 203.2 tok/s — generation speed (aggregate across clients) · API server · 2 clients
- 296.5 tok/s — generation speed (aggregate across clients) · API server · 4 clients
- 400.7 tok/s — generation speed (aggregate across clients) · API server · 8 clients
- Same Paiton project/author as the original register's long-KV4 entry; do not count as independent corroboration.
- Different checkpoint, runtime, sampling and workload from Paiton's earlier 8K matched Radiance comparison; no cross-release speedup percentage follows.
- 65K is configured capacity, not eight simultaneous full-65K conversations; Linux result is not native Windows.
Discussion
Share setups, questions and your own results. Posts are reviewed before they appear. Replies from Bitworks staff are marked.