Bare-metal GPU for business · 96 GB VRAM · Katowice

TESTS & GUIDES

How we test and report GPU server performance

GPU Server Hub testing policy: the actual measured configuration, tool versions, repetitions, limitations and downloadable result summaries.

Conceptual GPU module with subtle streams of data.
AI illustration.

Our results should help you choose a server for a specific workload. We publish measurements made on our hardware, not estimates derived from TFLOPS or results from other GPUs presented as our own. GPU Server Hub / Bestconnect prepares these materials and maintains this methodology as the supplier of the service, not an independent laboratory.

Hardware configuration

Each test describes the hardware actually measured: the GPU variant, its VRAM and configured power limit, CPU, system RAM and storage. Operating system, driver and tool versions form part of the environment description.

Test configuration: 96GB RAM. Offered configuration: 192GB RAM. System RAM and the GPU's 96GB VRAM are separate resources.

What we record before a measurement

Configurations that do not place the full model on the GPU, use offloading or use another weight format are described separately. We do not imply that different quantisations produce identical answer quality.

Repetitions and variation

For a comparable scenario, our starting procedure is a warm-up and at least three measured repetitions. The publication states the actual number of runs and how they are summarised. If a tool reports a mean and standard deviation, we do not call these a median or a minimum–maximum range.

We do not select only the fastest run. We retain measurements, settings and information about failed attempts. If a scenario needs a different procedure, we document the difference instead of presenting it as the same test.

LLM tests do not all measure the same thing

A compute microbenchmark may measure prompt processing and token generation separately, without HTTP, user queues or a complete application. We label such a result a microbenchmark; it is not automatically the response time of a deployed service.

A request-serving test requires a defined API server, workload and test clients. Only then do we report time to first token, response latency, concurrency and error rates for that scenario. Total server throughput is not the same as generation speed for one user.

A memory-capacity test shows whether a particular configuration ran with the stated settings. Loading model weights alone does not prove that every context length and user count will fit in VRAM.

Rendering, loading and sustained workloads

For rendering, we state the scene, application version, GPU backend and settings. Model-loading time is separate from computation time; we state whether a measurement used a warm cache.

A test lasting a few minutes is not a multi-hour test. When publishing longer-running results, we report the actual duration, temperature and power samples, observed errors and any performance changes. We do not change firmware or power limits to improve a headline result.

We distinguish the harness's stopping threshold from the GPU's factory protections. Reaching that threshold alone does not establish a fault, thermal clock reduction or protective GPU shutdown. We record the actual stopping reason, temperature and T.Limit margin, and sampling interval. “Not Active” describes a particular reading, not a guarantee that no event occurred between samples.

A threshold change creates a separate series. We retain earlier attempts and do not fill gaps using measurements from a different run. The two API attempts with a 90°C threshold are described separately from the earlier 85°C tests; they do not establish completion of four hours of load.

Public data and updates

Where a publication offers downloads, we provide prepared CSV or JSON summaries. These contain results and settings selected for publication, not complete system or administrative log dumps. SHA-256 checksums allow readers to verify the files.

We distinguish the test date from publication and editorial update dates. Corrections to results or methods are described alongside the material. A driver, model or environment change may require a new test; we do not replace an old configuration's name without taking new measurements.

Using the results

A result describes the specified scenario, not a performance guarantee for every application. We do not claim superiority over a particular competitor without comparable measurements of its hardware. We also describe limitations and cases where another configuration is needed.

Browse tests and guides, check the offer or ask about your model and workload.