Qwen3.5-9B Q8_0 on the Radeon RX 9060 XT 16GB (Windows)
Candidate
Candidate Windows single-card configuration for RX 9060 XT 16 GB; no published fit or performance result.
What will be tested
| Card | 1 × AMD Radeon RX 9060 XT 16GB |
|---|---|
| Operating system | Windows |
| Backend | Vulkan |
| Model | Qwen3.5-9B (Apache-2.0) |
| Model file | Unsloth GGUF conversion, Q8_0 |
| Context | 32,768 tokens |
| Concurrent requests | 1 (one request at a time) |
| Placement | Full GPU |
| Reviewed | Not yet |
| Sources | Upstream model · GGUF file page |
What to expect
No published results yet. Measured results appear here only after testing and review.
Install
Candidate — no installer yet.
Overview
This Windows candidate pairs the RX 9060 XT 16 GB with the pinned model and quantization named above at the catalog's 32,768-token context. The 16 GB variant uses a different, larger Q8 9B artifact than the 8 GB recipe. Compare answer behavior and free-memory headroom at the same task, rather than treating the two variants as one card. It is a test plan, not a claim of load, full-GPU placement, answer quality or speed.
Requirements
Use a standard-user 64-bit Windows 11 lab system with an AMD driver exposing Vulkan and the Microsoft Visual C++ x64 runtime. The current control window supervises one loopback server; driver installation and a consumer installer are not included. Check the card's currently reported free GPU memory, not just the number printed on its box. A longer prompt, more active requests and the desktop can change the memory budget. Keep the exact driver and operating-system build with any later test.
The pinned GGUF is 9,527,502,048 bytes (9.53 GB decimal). Keep more free local SSD space than the artifact itself for verified acquisition and cache maintenance; no fixed extra margin or host-RAM minimum has been qualified. Review the exact model license before acquisition. A planned trial must record the actual free-memory budget, selected device, loaded context, host and GPU buffers, delivered output and representative answers. The catalog context is an intended setting, not a demonstrated maximum.
Limitations
- The exact-card load, context, GPU/host buffers, response quality and speed have no published qualification result.
- The supervised Windows path requires separate AMD driver and Visual C++ prerequisites; no consumer installer is released.
- This candidate requests full-GPU placement; it is not a verified physical-residency or all-operations-on-GPU claim.