---
title: Two RX 7900 XT cards as independent Qwen3.8-27B workers (Windows)
url: "https://bitworks.io/amd-inference/recipes/rx-7900-xt-20gb-x2-windows-qwen3-8-27b-iq4xs/"
description: Candidate Windows two-worker configuration for RX 7900 XT 20 GB; no published fit or performance result.
---

# Two RX 7900 XT cards as independent Qwen3.8-27B workers (Windows)

> Candidate Windows two-worker configuration for RX 7900 XT 20 GB; no published fit or performance result.

Human-readable page: https://bitworks.io/amd-inference/recipes/rx-7900-xt-20gb-x2-windows-qwen3-8-27b-iq4xs/

## Content

- **Status:** In lab testing (candidate)

- **GPU:** 2 × AMD Radeon RX 7900 XT 20GB

- **Operating system:** Windows

- **Backend:** vulkan

- **Model:** Qwen3.8-27B UD-IQ4_XS — Unsloth GGUF conversion of Qwen/Qwen3.8-27B (Apache-2.0)

- **Context:** 16384 tokens

- **Concurrency:** 2

- **Placement:** full-gpu

No measured results are published for this recipe.

## Overview

With two workers, two independent jobs could run at the same time—for example, a coding agent and a review agent working on separate tasks, or two people sharing a workstation—instead of one waiting for the other. Each answer would still use one RX 7900 XT 20 GB, not both cards together. Shared CPU, memory, power and cooling constraints could change each worker's speed. The possible gain is completed work over shared wall-clock time; it has not been measured. This is two independent model copies, not one model split across cards.

The proposed worker on each card uses the pinned IQ4_XS 27B artifact at 16,384 context and one active request. A router would keep each conversation on its worker so cached state is not silently lost. FastLLM does not yet launch or route two isolated workers together on Windows.

## Requirements

A Windows 11 lab system needs two recognized discrete cards, an AMD Vulkan-capable driver and the Microsoft Visual C++ x64 runtime. The current pair selection is heuristic, not a verified PCI/topology binding. There is no qualified dual-card launch or consumer installer. Before attempting two cards, verify the motherboard has two usable slots and record the PCIe lanes and link speed each card actually negotiates. Check the complete system's power-supply connections/capacity and case airflow against both boards; no generic PSU minimum is established here. Identify both cards individually, leave memory reserves on each, and measure sustained thermals. A pair's advertised VRAM is not one contiguous pool.

The pinned GGUF is 14,252,845,984 bytes (14.25 GB decimal). Keep more free local SSD space than the artifact itself for verified acquisition and cache maintenance; no fixed extra margin or host-RAM minimum has been qualified. Review the exact model license before acquisition. One cached file need not be downloaded twice, but each worker would need its own GPU memory, context state and loopback port. Verify the selected card and buffers independently for both workers. No per-card fit or operation-residency result is implied.

### Planned evaluation

Compare one worker with one job, one worker with two waiting jobs, and two workers with the same independent job mix. Record completed jobs, each job's response time, failures and cancellation—not queued work as if it had finished. The separate two-XT Q8 split candidate uses a different quantization at the same 16K catalog context, so its raw rate would not isolate the value of replicas.

For agents: how to buy here — https://bitworks.io/checkout.md
