How Bitworks benchmarks AMD inference
How we test model speed, concurrent requests, occupied context, task correctness and whole-system limits on AMD hardware, and how community reports fit in.
What this page means today
This is the shared method behind the AMD inference recipes. It describes the tests a useful result has to cover. It does not claim that every test has run or passed.
FastLLM is Bitworks’ integration package for configuring existing inference engines and models; it is not a new engine. Every recipe is still a candidate. No Bitworks performance result or consumer installer is published yet.
Three kinds of information appear on these pages, and each is labelled where it appears:
- Bitworks measurements: results from our own runs under this method. None is published yet. When one is, it will name its exact configuration and source revision.
- Community reports: attributed results from other people, shown on applicable GPU pages under “Community reports — not measured or reproduced by Bitworks”, with a link to the original. They show what someone reported on their own system. They do not qualify a recipe.
- Proposed settings: the candidate recipe configurations. They describe what we intend to test, not what has passed.
Test the exact model on the whole system
A result belongs to one exact configuration. That means the model file and its checksum, the conversion source, the GPU memory variant and card count, the operating system, the driver, the engine build, the launch settings, the attention and KV-cache choices, and any acceleration or offload. Windows and Linux are separate configurations. So are WSL, containers and virtual machines. A different quantization or backend is a different configuration too.
The record also describes the rest of the computer:
- CPU, and installed and usable memory;
- each memory module’s capacity and configured speed;
- motherboard and BIOS version;
- the PCIe link each card actually negotiated, measured during a stated phase of the test;
- power settings.
A slot’s specification is not a measured link, and two memory modules do not prove memory bandwidth. Anything we could not confirm is recorded as unknown, and later observations are never written back into an earlier run. Where a system difference might explain a result, we separate effects shown in a controlled comparison from possible influences that were not isolated.
Requests, slots and workers are different
| Setting | What it means | What a result must state |
|---|---|---|
| Client requests (C) | How many API requests the test sends at once | Release timing, how far requests actually overlapped, queueing, completions and failures |
| Server slots (S) | How many sequences the server can work on at once | Requested and effective slots, and context per slot versus total context |
| Workers (W) | Separate model copies, each in its own server process | Which GPU each worker uses, how requests are routed, and one shared measurement window |
The default Windows serving path runs one slot and one worker (S1, W1). Sending 2, 4 or 8 requests to it measures queueing under pressure, not parallel decoding. Overlapping HTTP requests do not prove the GPU worked on them at the same time.
Separate multi-slot and two-worker lab paths exist, but neither is an approved automatic customer configuration. A model split across two cards is not two workers, and a split alone does not prove more throughput. Total server context is never presented as the context each concurrent user gets.
Configured context versus the prompt actually occupied
The test ladder uses 4K, 8K, 16K and 32K per-request windows, plus small calibration cases and the recipe’s exact ceiling. Longer 64K or 128K trials need a separately reviewed configuration.
A configured window is a limit, not a measurement. A server configured for 32K that processes a short prompt has not produced a 32K result. Every case records the configured window and the tokens actually used: tokenizer-counted input, the output allowance, and chat-template and tool overhead. Long-context tests fill most of the window, reserve room for the answer, and check for truncation or context shifting. If the requested workload changed, the cell fails rather than being reported as tested.
Per-request and aggregate rates measure different things
Prompt processing is how quickly input is evaluated. Generation rate is output speed during decoding. Time to first text and total response time are what a person actually waits for, and include any queueing. We name each one rather than calling them all “speed”.
For concurrent tests, completed-output throughput is the output tokens from successfully completed requests divided by one common window. That window runs from release until the last request finishes, fails or hits its deadline. Completed-request throughput uses the same window. We never add up separately measured per-request rates, and we never shorten the window by dropping a slow failure.
Aggregate throughput says how much work a system finished, not how fast one user’s answer arrived. Results show per-request latency distributions, success and failure counts and run-to-run variation. A handful of samples cannot support a dependable tail-latency figure: the proposed floor for latency-tail or capacity claims is 100 request observations per cell across at least three fresh serving runs.
Cold, warm and prefix reuse are separate tests
Warmup runs happen outside the measured window, and their number and the model’s warm or cold state are recorded. The baseline does not reuse prompt prefixes.
A shared-prefix agent scenario is its own named test, with its cache hits and reuse policy recorded, because a reused prefix can shorten later turns without making decoding itself faster. Cold loading, warm serving, cached prefixes, speculative decoding and other accelerations are never blended into one number. Comparisons change one setting at a time, and the order is counterbalanced when run-to-run drift matters.
Push the configuration, then show where it stops being useful
After a single-client baseline, the sweep raises concurrent requests through 1, 2, 4 and 8 at each applicable context. Screening uses a few bounded waves. Larger predeclared runs are needed for latency-tail or capacity claims.
Timeouts, memory failures, malformed answers and server exits all stay in the evidence. Escalation stops when a safety or correctness check fails, and a changed model, context, offload setting or output length is a new configuration. The aim is a useful operating envelope, not crashing the computer or rerunning until a flattering number appears.
A finite burst is not a sustained arrival-rate guarantee. Sustained concurrent use needs its own bounded soak, and single-client soaks do not establish concurrent reliability.
Host memory, PCIe topology and offload
Moving part of a model into system memory changes what a result means. Offloaded layers, experts or embeddings depend on host memory capacity and bandwidth, CPU placement and the PCIe path. They can slow a run sharply, and they make it hard to compare with an all-GPU configuration.
Results state what was offloaded, where that is known. A full-GPU layer count or a model-buffer log is a useful clue, but it does not prove that every operation stayed in dedicated GPU memory. For two-card setups, the slot wiring matters too: a card behind a chipset or a narrow link can behave very differently from one on CPU lanes. We record the topology rather than assume it.
Completion is not correctness
Finishing a request is not the same as answering it correctly, so throughput and quality are reported separately.
The independent-task screen runs the same eight versioned tasks one, two and four at a time. Its task definitions and strict graders cover structured extraction, constrained answers, simple arithmetic and retrieval among distractors. Correctly completed tasks are divided by the time of the whole attempted suite, including wrong answers and failures. Planned, attempted, correct, incorrect, inconclusive and unattempted counts are all kept. There are no hidden retries, and a wrong answer is never cleaned up into a pass.
An opt-in assessment mode continues past a wrong answer so that an imperfect model still gets a complete quality count. It still stops on an inconclusive answer, a transport error, truncation or a deadline.
These small screens use non-streamed responses, so time to first text is not measured there. They are not a general intelligence benchmark or an autonomous coding evaluation, and passing tests on mock servers says nothing about a real model’s accuracy. Broader quality results must name their dataset, license, sample count, grader and settings, and the checks are repeated whenever the quantization, KV cache, acceleration, backend or offload changes.
How community reports fill gaps
Where Bitworks has not tested a configuration, attributed community reports can still help a buyer, as long as the gap they close is no wider than what the source supports. Each report keeps:
- the exact card variant and count;
- the operating system, including WSL, VM or container;
- the model generation and quantization;
- the engine and build;
- host and PCIe notes;
- configured context versus the prompt actually occupied;
- client requests versus slots and workers;
- cache state;
- the author’s method.
Unknown details stay unknown. Ranges, approximations, repetitions and separate comparison passes are kept as reported, and each comparison names what changed between its arms. Reports from the same author or host are shown as related, not as independent confirmation.
Community reports never transfer across boundaries:
- Linux results say nothing about native Windows;
- a larger-memory card says nothing about a smaller variant;
- an older model generation says nothing about a newer one;
- total reusable cache is not context for one request.
Failures, spills, instability and version-specific fixes are kept next to successes, because a problem reported on one software version may not apply to a later one. Tuned kernels and engine forks are treated as candidates to reproduce under their exact build, not as recommended defaults. Each recipe page links to its GPU page’s community section when approved reports or scoped gap notes exist there. Those reports may use a different card count, OS, engine, model, quantization or context, so they never confirm the exact recipe.
Reproduce the test and inspect the source
The public source is Bitworks AMD inference integration on GitHub. These links are pinned to source revision a91e83bb:
- benchmark methodology
- single-request benchmark
- concurrent-request collector
- task-quality screen
- catalog-wide test matrix
- strict result comparator
- test suites
The matrix is an offline plan, not proof that every cell runs. The public source includes the concurrent-request collector and its tests. Private Windows/AMD lab screens have exercised concurrent requests, but no physical result bundle has been approved for publication or qualified as a consumer performance claim.
Each future result will link the exact source revision, its reviewed report format and sanitized evidence. No physical result bundle is approved yet. The integration-code license is still being chosen, so public source visibility is not a license grant, a performance qualification or a consumer release.
For examples of transparent reporting, see oMLX’s separate performance and intelligence pages. These are external methodology references, not AMD or Windows measurements, and their scores are not directly comparable.
What the labels mean
“Candidate” identifies a configuration selected for assessment. A reviewed lab result applies only to its exact tested conditions and stated limitations. Qualification needs additional evidence and review, and an installer needs a separately approved release. None of these labels is earned just because a model loads or returns an answer. A community report keeps its “Not reproduced by Bitworks” label whatever it shows.