Estimate — sign in for exact rates.

 Trusted GPU & AI infrastructure since 2015 — Huntsville, AL
AMD Inference

How Bitworks benchmarks AMD inference

Our method for testing model speed, concurrent requests, useful context lengths, answer quality and whole-system limits on AMD hardware.

What this page means today

This is the shared method for evaluating the AMD inference recipes. It describes the tests a useful result should cover, not a claim that every test has already passed. The current recipe pages are candidates. FastLLM is Bitworks’ integration package for configuring existing inference engines and models, not a new engine. It remains in development, and this page does not qualify a consumer installer or publish a card’s performance.

The goal is useful completed work: a responsive conversation, several agents working at once, or a long document processed correctly. A short, single-request speed number cannot answer all three questions.

Test the exact model on the whole system

Every applicable catalog model and quantization belongs in the coverage matrix, not only the headline model. Each result identifies the exact model file and digest, conversion provenance, GPU memory variant and count, operating system, driver, engine version, launch settings, attention/cache choices and any acceleration or offload. Windows and Linux results remain separate. A changed quantization or backend is a different configuration.

The system record also includes CPU, installed and usable RAM, memory-module configuration and configured speed, motherboard, BIOS, storage where relevant, and power settings. PCIe links and card attachment are recorded when verified under a stated workload phase; a motherboard slot specification is not a measured link. Unknowns stay unknown. We distinguish effects demonstrated by controlled tests from possible influences that have not been isolated. Two DIMMs do not prove measured memory bandwidth, and a different Windows patch level does not by itself explain a speed difference.

Concurrent users, processing slots and GPU workers are different

What we vary What it tells you
1, 2, 4 and 8 simultaneous client requests How response time, queueing, completions and total useful throughput change under demand
Server processing slots Whether the configured engine can work on multiple sequences, and how much context each receives
Independent model workers Whether separate GPU-bound model copies can complete jobs concurrently while sharing the rest of the PC
A model split across GPUs Whether distributing one model enables a useful capacity or quality choice; it is not automatically two workers

The current normal Windows service uses one processing slot. Multiple clients against it test contention and queueing, not proven parallel decoding. Multi-slot serving and two-worker orchestration need separate implementation and physical validation. Overlapping HTTP requests alone do not prove overlapping GPU execution, and adding two independently measured card rates is not a concurrent-system result.

For subagent use, the intended comparison runs the same collection of independent tasks sequentially and concurrently. It records correct completed tasks, retries, total completion time and whether a long task makes shorter tasks wait. That task test is separate from simply sending a batch of prompts.

Use context lengths that represent real work

The main context ladder is 4K, 8K, 16K and 32K per request where the recipe permits it, with smaller calibration cases and the exact recipe ceiling also recorded. Longer 64K or 128K trials require a separately reviewed model/backend/memory configuration; they are not promised by the present catalog. A recipe ceiling is a policy limit, not proof of the model’s maximum capability or physical fit.

We distinguish the configured context window from the number of input tokens actually processed. Input, output allowance, chat-template and tool overhead must fit the effective per-request budget. A 32K setting tested with a short prompt is not a long-context result. Long-context tests fill a substantial part of that window, reserve output space and check for truncation, context shifting and answer quality. Server-wide context is not advertised as context available to each concurrent user.

The workload families include short interactive requests, longer document/code inputs, sustained outputs, near-window inputs, equal-length concurrent requests and a declared mix of short and long requests. Results report actual input and output counts. Early termination or a reduced output budget cannot silently become a successful result for the originally requested workload.

Push the configuration, then show where it stops being useful

After a single-client baseline, the planned sweep increases concurrent requests through 1, 2, 4 and 8 at each applicable context. Screening uses bounded repeated request waves. A larger, predeclared run is needed for latency-tail or capacity claims: our proposed floor is 100 request observations per workload/concurrency cell across at least three fresh serving runs, including failures. This is not a guarantee of statistical precision; the sample count and uncertainty still matter.

Timeouts, memory failures, malformed answers and server exits remain in the evidence. An untested or outside-policy cell is labeled, not omitted. Escalation stops when a safety or correctness check fails; a changed model, context, offload policy or output length becomes a new configuration. The purpose is to find a useful operating envelope, not to crash the operating system or repeatedly rerun until a flattering number appears.

A finite burst test is not a sustainable production arrival-rate guarantee. Sustained concurrent use also needs a bounded soak. Single-client soak results do not establish concurrent reliability.

Read speed and latency together

Prompt processing describes how quickly input is evaluated. Generation rate describes output during decoding. Client-observed time to first non-empty streamed text and total response time include additional waiting, including queueing. We name the measurement rather than treating all of these as interchangeable speed scores.

For concurrent requests, completed-output throughput is the total output tokens from successfully completed requests divided by the common elapsed test window, from release until the last request finishes, fails or reaches its deadline. Completed-request throughput uses the same window. We do not sum individual request rates or shorten the denominator by excluding a slow failure. Incomplete output is disclosed separately.

The result should show per-request latency distributions, successful and failed request counts, variation between runs and results for each workload class. A tiny sample does not establish a dependable tail-latency estimate. For agent tasks, useful throughput additionally depends on meeting the declared correctness and response-time criteria.

Keep cache and acceleration comparisons fair

The exact warmup count and model state are recorded. The uncached baseline does not reuse a prompt prefix; a shared-prefix agent scenario is a separate test. Cold loading, warm serving, cached prefixes, speculative decoding and other accelerations are not blended into one number. Inputs, output targets and sampling settings remain fixed when comparing one change. Material run-to-run drift calls for investigation and a counterbalanced comparison, not a causal claim from one pair of runs.

Measurements collected before or after the requests do not establish load-time memory peaks, clocks or temperature. Measurements taken during the workload must disclose coverage and possible collection overhead. A full reported GPU layer count and a model-buffer log are useful clues, but neither proves every operation or allocation stayed in dedicated GPU memory.

Answer quality needs its own evidence

Speed and intelligence are separate panels. A startup response check is not an intelligence benchmark. A quality result needs its dataset and version, license, sample count and denominator, selection method, grading method, exact model/chat template, sampling and reasoning settings, and permitted tools.

The planned quality coverage includes structured output, tool-call validity, retrieval from different positions in long inputs, and task completion under contention. Changing quantization, KV cache, acceleration, backend or offload requires checking quality again. Synthetic token workloads establish speed behavior, not reasoning ability. External community scores stay external and do not become Bitworks measurements.

Reproduce the test and inspect the source

FastLLM’s lab suite contains a single-request API benchmark, a reliability runner, narrow semantic checks and a strict result comparator. The new catalog-wide matrix is an offline plan, not an automated claim that every cell runs. The new concurrent-client collector has mock-process checks but no physical Windows/AMD qualification yet. Mixed-length task orchestration, multi-slot serving, independent workers and broader task-quality coverage remain separate implementation and qualification work. New report formats need review before they can appear as public results.

The public source is Bitworks AMD inference integration on GitHub. Start with the initial source revision’s benchmark methodology, catalog-wide test matrix, single-request benchmark, concurrent-client collector and test suites. These pinned development links are not an installer or a promise that every planned test is implemented.

Each future published result will link to the exact source revision, reviewed report schema and sanitized evidence used for that measurement. No sanitized physical-result bundle is approved here yet. Bitworks’ integration-code license is still being selected; public source visibility should not be read as a license grant, performance qualification or consumer release.

For examples of transparent reporting, see oMLX’s separate performance and intelligence pages and its benchmark suite on GitHub. These are external methodology references, not AMD/Windows measurements or directly comparable scores.

What the labels mean

“Candidate” identifies a configuration selected for assessment. A reviewed lab result applies only to its exact tested conditions and declared limitations. Qualification requires additional evidence and review, and an installer requires a separately approved release. None of those labels is earned merely because a model loads or returns an answer.

Select the fields to be shown. Others will be hidden. Drag and drop to rearrange the order.
  • Image
  • SKU
  • Rating
  • Price
  • Stock
  • Availability
  • Add to cart
  • Description
  • Content
  • Weight
  • Dimensions
  • Additional information
Click outside to hide the comparison bar
Compare
Compare ×
Let's Compare! Continue shopping