---
title: Qwen3.8-27B Q8_0 on two Radeon AI PRO R9700 32GB cards (Linux)
url: "https://bitworks.io/amd-inference/recipes/radeon-ai-pro-r9700-32gb-x2-linux-qwen3-8-27b-q8/"
description: Candidate Linux dual-split configuration for Radeon AI PRO R9700 32 GB; no published fit or performance result.
---

# Qwen3.8-27B Q8_0 on two Radeon AI PRO R9700 32GB cards (Linux)

> Candidate Linux dual-split configuration for Radeon AI PRO R9700 32 GB; no published fit or performance result.

Human-readable page: https://bitworks.io/amd-inference/recipes/radeon-ai-pro-r9700-32gb-x2-linux-qwen3-8-27b-q8/

## Content

- **Status:** In lab testing (candidate)

- **GPU:** 2 × AMD Radeon AI PRO R9700 32GB

- **Operating system:** Linux

- **Backend:** Not yet selected

- **Model:** Qwen3.8-27B Q8_0 — Unsloth GGUF conversion of Qwen/Qwen3.8-27B (Apache-2.0)

- **Context:** 16384 tokens

- **Concurrency:** 1

- **Placement:** full-gpu

No measured results are published for this recipe.

## Overview

This Linux candidate proposes splitting one pinned Qwen3.8-27B Q8_0 model across two Radeon AI PRO R9700 32 GB cards with a layer split, one active request and 16,384 context. Each card would hold a portion of the model; the result is not two independent answer workers. A single R9700 remains the simpler option when one card has enough currently free memory for this same model; the second card could then serve a separate worker rather than accelerate one answer. Neither aggregate advertised VRAM nor a requested full-layer count proves fit, residency or faster generation.

The 16K context comes from this exact Q8 catalog profile and its conservative reserve estimate. Some single-card Q4/Q6 profiles specify 32K, but that does not establish that this Q8 pair is limited to 16K or that the longer single-card setting would fit here. Compare matched contexts when testing latency or throughput.

## Requirements

The current private Linux guided launcher accepts one Vulkan GPU, so this two-card configuration is a proposed extension, not a runnable FastLLM recipe today. A future launcher must bind both devices and verify the exact model and consent before serving. Before attempting two cards, verify the motherboard has two usable slots and record the PCIe lanes and link speed each card actually negotiates. Check the complete system's power-supply connections/capacity and case airflow against both boards; no generic PSU minimum is established here. Identify both cards individually, leave memory reserves on each, and measure sustained thermals. A pair's advertised VRAM is not one contiguous pool.

The pinned GGUF is 29,047,086,048 bytes (29.05 GB decimal). Keep more free local SSD space than the artifact itself for verified acquisition and cache maintenance; no fixed extra margin or host-RAM minimum has been qualified. Review the exact model license before acquisition. A trial must record the exact device pair and topology, each card's current free VRAM, the requested split, per-card model and compute buffers, host-memory use, API output and repeated response times. Splitting capacity and two separate replicas answer different buyer needs.

### Separate Flash-Next research

Qwen3.8-Flash-Next is a different model under separate [Qwen Community 1.0 terms](https://huggingface.co/Qwen/Qwen3.8-Flash-Next), not this pinned 27B artifact. A surveyed GSQ-RCO Q2_0 conversion has less weight data than an optimistic 2 × 32 GiB sum, but almost no aggregate headroom for context and compute; the n-gram table can instead stay mapped on host storage. That arithmetic does not establish fit or GPU residency. [Strata's AMD HIP documentation](https://github.com/Niko1221/Strata/blob/main/docs/AMD_HIP.md) describes Linux AMD multi-card layer splitting and host expert caching but one AMD card per model on Windows. A separate exact artifact, runtime, RAM/SSD and output-quality review is needed before any FastLLM recipe.

For agents: how to buy here — https://bitworks.io/checkout.md
