Bare-metal GPU for business · 96 GB VRAM · Katowice

TESTS & GUIDES

How quickly does an AI server answer one or several requests?

We measured AI response times with one and several simultaneous requests: results for 8B, 32B and 72B models and a four-hour GPU server test.

Conceptual GPU module with subtle streams of data.
AI illustration.

A quick answer for one person and a high total number of completed answers are different goals. This test covers both: one request and four requests sent simultaneously. We compare Qwen 8B, 32B and 72B on the same GPU. For 32B, we also completed a separate four-hour run.

An API is the interface an application uses to send a question and receive an answer. We measure it locally on the server, excluding Internet latency and an additional client application. All current results follow our shared CORE 1.1 methodology.

Reading the table

The first time shows when the measurement client received the first generated token. This is not always the first part of the final answer visible to the user. The second time indicates completion of the whole response. A reader can follow streamed text as it appears; software expecting a complete data object needs to wait until the end.

We test inputs of 2048 and 8192 tokens, the pieces of text used by a model. Every response has a fixed length of 256 tokens. A configuration therefore cannot achieve a better time merely by writing less. This is a controlled computation test, not an assessment of the generated text's meaning.

Aggregate throughput describes the request group as a whole. It is not the speed experienced by each individual user. Time figures are medians: the middle observations after sorting the measurements.

One request and four concurrent requests

The labels 8B, 32B and 72B describe the approximate number of billions of model parameters, not the file size. Q4_K_M is a more compact way of storing those parameters. Our AI and GPU glossary explains these terms with simple examples.

Model and format Input tokens Concurrent requests First token, median Median complete response Aggregate throughput, tokens/s
Qwen3-8B Q4_K_M 2048 1 0.2 s 1.5 s 171.9
Qwen3-8B Q4_K_M 8192 1 0.8 s 2.4 s 108.8
Qwen3-8B Q4_K_M 2048 4 0.7 s 3.3 s 313.0
Qwen3-8B Q4_K_M 8192 4 2.5 s 6.5 s 155.5
Qwen3-32B Q4_K_M 2048 1 0.8 s 5.5 s 46.2
Qwen3-32B Q4_K_M 8192 1 3.4 s 8.3 s 30.8
Qwen3-32B Q4_K_M 2048 4 2.8 s 12.5 s 81.8
Qwen3-32B Q4_K_M 8192 4 10.1 s 24.5 s 41.5
Qwen2.5-72B-Instruct Q4_K_M 2048 1 1.7 s 11.6 s 22.0
Qwen2.5-72B-Instruct Q4_K_M 8192 1 7.1 s 17.5 s 14.7
Qwen2.5-72B-Instruct Q4_K_M 2048 4 6.1 s 26.2 s 38.9
Qwen2.5-72B-Instruct Q4_K_M 8192 4 21.2 s 50.9 s 20.0

Each scenario has a full warmup round at its target concurrency and five measured rounds. Both the single-request and concurrent cases reserve four service slots. We change active request count while preserving the memory allocation. Separate one-slot baseline measurements answer another question and do not replace this control.

The 32B model with the longer input illustrates the trade-off: four simultaneous requests increased aggregate throughput from 30.8 to 41.5 tokens/s, a gain of 34.9%. At the same time, the typical individual response took 24.5 s instead of 8.3 s. The server completed more work per unit of time, but each person waited longer. That may suit a background task queue; for live conversation, acceptable waiting time matters more.

What does this mean for a company assistant?

Four simultaneous requests do not mean a limit of four user accounts. More people can have access without sending a question at exactly the same time. Planning needs concurrent task counts and input lengths, not just employee numbers.

A useful next step is to prepare representative questions and longer documents, then try them under expected traffic. Our practical tasks complement speed measurements with output-format and exact-answer checks.

Sustained operation with 32B

Measured duration Completed responses Aggregate throughput, tokens/s
240.3 min 2304 40.9

Across 576 rounds, the server completed 2304 responses. The median individual response took 24.9 s, and 95% took around 25.1 s or less — the latter is called p95. This describes the distribution in this run, not a guaranteed time for any question. The aggregate 40.9 tokens/s applies to the group of four requests, not to each person separately.

The server completed a four-hour request-handling test. Temperature was stable in the measured configuration. This result describes that run, not continuously full GPU utilisation or an identical speed for every response.

The sustained workload concerns one specific configuration: Qwen3-32B Q4_K_M, the longer input and four concurrent requests. Its outcome is not extended to every model in the table. Measured time starts with the first measured request and ends with the last completed response, excluding loading and warmup.

It measures repeated request handling, not guaranteed uninterrupted full GPU utilization in every second. Our power and energy article also examines the longer workload.

Choosing a configuration

Start with acceptable waiting time, input length and expected simultaneous requests. Then assess the model's answers on your own data. Together, these observations tell you whether the configuration suits the application.

For a first prototype, the private API guide walks through a protected initial deployment. Explore the server offer or describe your application.

Technical details

Measurement comparability

The short CORE 1.1 matrix includes Qwen3-8B Q4_K_M, Qwen3-32B Q4_K_M and Qwen2.5-72B-Instruct Q4_K_M. Two input lengths × two concurrency levels × three models give 12 scenarios. Each has one full warmup round and five measured rounds; a four-request round contains four responses.

Both C1 and C4 reserve four slots of 8704 tokens, or 34816 in total. Other settings remain identical: all GPU layers, F16 KV, Flash Attention, 16 threads, batch 2048, ubatch 512 and mmap, with lazy loading, fitting, prompt caching and context shifting disabled. The environment is Rocky Linux 10.2, NVIDIA 595.91.07 and llama.cpp b11146.

Timing and throughput calculations

Concurrent requests start through a common barrier, and we confirm a common active interval for all responses. Request timing ends at the generation terminal event, not the later HTTP connection closure. A synthetic first token is not the same as the first visible final-answer text from a reasoning model.

Aggregate throughput is generated tokens divided by the union of response windows. Overlapping time is not counted repeatedly. This is neither the historical mean of round rates nor an individual decode rate. For 256 output tokens, the engine's internal timing can cover 255 decode steps; the data preserve both meanings.

Five rounds do not support strong claims about rare delays. In the sustained homogeneous scenario, descriptive p95 becomes eligible only from 200 responses. It describes the observed distribution under these conditions, not an SLA.

Sustained measurement and sources

The plan, executor, analyser and cooling-assessment criteria are pinned before the sustained campaign. We require at least four hours from the first measured request to the last completed response, preserving gaps and telemetry. Numeric results are in the CORE JSON, while thermal assessment and hourly power are in the sustained-run analysis.

All older files remain unchanged in the archive. The historical 8B API test with 32,768 input tokens is not part of the new 2048/8192 matrix. The separate new 32,768-token measurement concerns the 72B computation engine, not this HTTP layer.

Attachments provide runtime and model versions alongside unrounded statistics. Endpoint behaviour is documented in the pinned llama-server reference.

4 concurrent requests · aggregate throughput
  1. Qwen3-8B Q4_K_M155.5155.5 tokens/s
  2. Qwen3-32B Q4_K_M41.541.5 tokens/s
  3. Qwen2.5-72B-Instruct Q4_K_M2020 tokens/s

8192 input tokens, 4 slots, 5 rounds. Total output tokens divided by response-window time, not the speed of an individual response.

Download the data

Public test results and settings, without credentials or administrative identifiers.

Current measurement

Archived measurements