---
title: Qwen3.5-9B Q4_K_M on the Radeon RX 9060 XT 8GB (Linux)
url: "https://bitworks.io/amd-inference/recipes/rx-9060-xt-8gb-linux-qwen3-5-9b-q4km/"
description: Candidate Linux single-card configuration for RX 9060 XT 8 GB; no published fit or performance result.
---

# Qwen3.5-9B Q4_K_M on the Radeon RX 9060 XT 8GB (Linux)

> Candidate Linux single-card configuration for RX 9060 XT 8 GB; no published fit or performance result.

Human-readable page: https://bitworks.io/amd-inference/recipes/rx-9060-xt-8gb-linux-qwen3-5-9b-q4km/

## Content

- **Status:** In lab testing (candidate)

- **GPU:** 1 × AMD Radeon RX 9060 XT 8GB

- **Operating system:** Linux

- **Backend:** Not yet selected

- **Model:** Qwen3.5-9B Q4_K_M — Unsloth GGUF conversion of Qwen/Qwen3.5-9B (Apache-2.0)

- **Context:** 16384 tokens

- **Concurrency:** 1

- **Placement:** full-gpu

No measured results are published for this recipe.

## Overview

This Linux candidate pairs the RX 9060 XT 8 GB with the pinned model and quantization named above at the catalog's 16,384-token context. The 8 GB variant has much less room for weights and context state than its 16 GB sibling. This Q4 9B profile tests whether the lower-memory card can serve a useful text workload without assuming the same outcome as the 16 GB version. It is a test plan, not a claim of load, full-GPU placement, answer quality or speed.

## Requirements

The current Linux path is a private, supervised source-checkout workflow for one recognized Vulkan GPU and a loopback server. It requires explicit artifact consent. No physical Linux AMD serving outcome is qualified. Check the card's currently reported free GPU memory, not just the number printed on its box. A longer prompt, more active requests and the desktop can change the memory budget. Keep the exact driver and operating-system build with any later test.

The pinned GGUF is 5,680,522,464 bytes (5.68 GB decimal). Keep more free local SSD space than the artifact itself for verified acquisition and cache maintenance; no fixed extra margin or host-RAM minimum has been qualified. Review the exact model license before acquisition. A planned trial must record the actual free-memory budget, selected device, loaded context, host and GPU buffers, delivered output and representative answers. The catalog context is an intended setting, not a demonstrated maximum.

For agents: how to buy here — https://bitworks.io/checkout.md
