Two RX 7900 XT cards as independent Qwen3.8-27B workers (Linux)
Candidate
Candidate Linux two-worker configuration for RX 7900 XT 20 GB; no published fit or performance result.
What will be tested
| Card | 2 × AMD Radeon RX 7900 XT 20GB, two workers (one model copy per card) |
|---|---|
| Operating system | Linux |
| Backend | Not yet selected |
| Model | Qwen3.8-27B (Apache-2.0) |
| Model file | Unsloth GGUF conversion, UD-IQ4_XS |
| Context | 16,384 tokens |
| Concurrent requests | 2 in total — 1 active request per worker, 2 workers |
| Placement | Full GPU |
| Reviewed | Not yet |
| Sources | Upstream model · GGUF file page |
What to expect
No published results yet. Measured results appear here only after testing and review.
Install
Candidate — no installer yet.
Overview
With two workers, two independent jobs could run at the same time—for example, a coding agent and a review agent working on separate tasks, or two people sharing a workstation—instead of one waiting for the other. Each answer would still use one RX 7900 XT 20 GB, not both cards together. Shared CPU, memory, power and cooling constraints could change each worker's speed. The possible gain is completed work over shared wall-clock time; it has not been measured. This is two independent model copies, not one model split across cards.
The proposed worker on each card uses the pinned IQ4_XS 27B artifact at 16,384 context and one active request. A router would keep each conversation on its worker so cached state is not silently lost. FastLLM does not yet launch or route two isolated workers together on Linux.
Requirements
The current private Linux guided launcher accepts one Vulkan GPU, so this two-card configuration is a proposed extension, not a runnable FastLLM recipe today. A future launcher must bind both devices and verify the exact model and consent before serving. Before attempting two cards, verify the motherboard has two usable slots and record the PCIe lanes and link speed each card actually negotiates. Check the complete system's power-supply connections/capacity and case airflow against both boards; no generic PSU minimum is established here. Identify both cards individually, leave memory reserves on each, and measure sustained thermals. A pair's advertised VRAM is not one contiguous pool.
The pinned GGUF is 14,252,845,984 bytes (14.25 GB decimal). Keep more free local SSD space than the artifact itself for verified acquisition and cache maintenance; no fixed extra margin or host-RAM minimum has been qualified. Review the exact model license before acquisition. One cached file need not be downloaded twice, but each worker would need its own GPU memory, context state and loopback port. Verify the selected card and buffers independently for both workers. No per-card fit or operation-residency result is implied.
Planned evaluation
Compare one worker with one job, one worker with two waiting jobs, and two workers with the same independent job mix. Record completed jobs, each job's response time, failures and cancellation—not queued work as if it had finished. The separate two-XT Q8 split candidate uses a different quantization at the same 16K catalog context, so its raw rate would not isolate the value of replicas.
Limitations
- Two independent workers and request routing are proposed, not implemented or measured together.
- The current Linux guided launcher supervises one GPU, not two isolated workers.
- Each worker's intended GPU placement and fit require separate verification; no aggregate throughput claim is available.
Bitworks listings for this card
-
In stock
SapphireSapphire PULSE AMD Radeon RX 7900 XT 20GB GDDR6 Gaming Graphics Card
$599.99 Manufacturer refurbished