Bare-metal GPU for business · 96 GB VRAM · Katowice

TESTS & GUIDES

Has Qwen caught up with GPT? A tie with Astra in our coding test

Qwen Flash Next and GPT Astra each scored 15/15 in our coding test. We compare completion times, explanation quality and 96 GB GPU requirements.

A GPU and a stack of computing modules linked by mint-coloured data streams against a dark background.
AI illustration.

Can a local model complete a coding task as successfully as a paid service? In our earlier test, Qwen3.8-Flash-Next and GPT-6 Astra each passed 15 out of 15 attempts. We examine what that tie means for someone choosing a model for coding work. The result covers five tasks, each repeated three times. It does not establish that the models are equally capable overall.

What exactly are we testing?

We prepared small tasks inspired by application development: date handling, data export, a chart’s response to mouse movement, invoice code refactoring and tests that detect bugs. Each had a defined scope and acceptance checks. Passing required completing the entire set of checks. Partially correct code was not enough.

Our reference was GPT-6 Astra, accessed through Codex with reasoning effort set to xhigh. The model ran on OpenAI’s service, while file operations and tests ran locally on Linux. Qwen ran on our server with an NVIDIA RTX PRO 6000 Blackwell Max-Q card and 96 GB of GPU memory.

This was an agent test: the model had access to tools. It could read files, change code, run commands and revise its solution after checking it. That differs from generating finished code in a single response.

Each configuration received the same five tasks, with three attempts each and a 20-minute session limit. We retained every scored result. The series ran at different times and with different configurations, so we are comparing working combinations of models and tools. Earlier errors in the materials and grading software have documented corrections, applied consistently to the assessed solutions.

Results for the five tasks

Labels such as 27B and 122B refer to approximate billions of parameters, not gigabytes. Q4, Q5 and Q8 describe weight storage formats. Lower precision saves memory, but its effect on results needs to be tested on the task.

Astra and Flash Next in xhigh mode each scored 15/15. Qwen3.8-27B Q4 with a 128k context came closest, fully passing 13 of 15 attempts. The much larger Qwen3.5-122B Q5 finished with 8/15. The other configurations we tested had lower success rates on this set.

Five coding tasks repeated three times: Astra and Qwen Flash Next passed 15 of 15 trials. Full results are available in the table below.
First series: five distinct tasks, each run three times. A pass required meeting every checked requirement in that trial. Download PNG chart
Chart results as a table
Chart results as a table
Model and configurationPassed / valid trialsPass rate
GPT-6 Astra · xhigh15/15100%
Qwen3.8 Flash Next · IQ4 · xhigh15/15100%
Qwen3.8 27B · Q4 · 128k13/1586.7%
Qwen3.5 122B · Q58/1553.3%
GPT-OSS 120B · MXFP46/1540%
Qwen3 Coder Next · Q84/1428.6%
Qwen3 Coder 30B · Q82/1513.3%
Qwen3 8B · Q40/150%

One Coder Next infrastructure failure was excluded: result 4/14. Historical configurations used different engines and tool interfaces. The small sample does not prove equivalence in other applications.

GPT-OSS-120B scored 6/15, Qwen3-Coder-Next Q8 passed 4/14 valid attempts, and Qwen3-Coder-30B Q8 scored 2/15. Qwen3-8B Q4 did not fully pass any of its 15 attempts. Coder Next’s fifteenth attempt was excluded because of an infrastructure failure; we did not count it as a failure of model ability. Neither the name “Coder” nor a larger parameter count guaranteed an advantage in our workflow.

Flash was also close to Astra on time. Its median session duration was about 4 minutes and 24 seconds, compared with 4 minutes and 16 seconds for Astra. This included edits, commands and reaching completion. An initial “I’ll check the files” message did not mean the solution was ready. With such a small task set, an eight-second difference is not strong evidence of a speed advantage.

The median is the middle of the ordered attempt times; with an even number of results, it is the average of the two middle values. It helps describe a typical duration, but does not guarantee a waiting time for any task.

What matters beyond correct code?

Users need to know what changed, why and which tests actually ran. We assessed explanations separately, using responses, code changes and recorded commands. In the earlier series, Flash averaged 7.7/10 and Astra 10/10. Reviews of Flash noted, among other things, excessive length and claims of broader verification than the evidence supported.

This was a supplementary assessment by AI from the same family as the reference model. Model names were hidden from the reviewer, but there was no independent human review or confirmation of identical reviewer versions across the two series. A score of 10/10 means meeting a particular rubric, not that Astra is infallible. We assessed visible explanations and evidence, without examining hidden reasoning.

How much memory did Flash Next need?

VRAM is the graphics card’s memory, while RAM is the server’s system memory. Our measurements use GiB, a unit that differs from the decimal GB used in specifications.

Flash Next used the IQ4_NL format and a 128k context. In the earlier series, peak GPU memory use reached 72.9 GiB, leaving about 22.7 GiB of headroom relative to the capacity reported by the driver. The measurement covered the model server’s operation, including preparation and testing, with one session at a time.

All transformer layers and the context cache were on the GPU. Some embeddings, tables representing pieces of text, used server RAM. According to the publisher’s model card, Flash Next has a 125-billion-parameter core, about 6 billion parameters active during computation and another 51 billion embedding parameters. The active parameter count does not describe the size of the whole model that must be stored.

What can a local model offer a company?

A local model lets a company keep code, documents, questions and generated results within its own environment, without sending them to an external model API. For internal analysis, this can make it easier to control where information is processed and who has access to it.

The company can set access permissions, logging rules and retention periods. It can also choose the model version and decide when to update it. That helps maintain a consistent setup for recurring tasks and check changes before introducing them into everyday work.

The condition is that the whole workflow stays local. Applications, integrations, telemetry and backups must not send the relevant data outside the company. After the model, software and necessary files are prepared, work can also continue offline, provided no part of the task depends on an external service.

This gives the company more control over data handling; it does not guarantee that leaks cannot occur, and using an external service does not mean that data becomes public.

The company still needs to configure and maintain the system, provide memory and power, and manage access. Our tests did not calculate total ownership costs or a break-even point against a subscription.

Does a higher setting give a better result?

After the first comparison, we prepared six new tasks and another 72 attempts. Flash Next ran with xhigh and medium settings, and Qwen 27B in Q4 and Q8 variants. Each configuration received 18 attempts. Astra did not take part in this series, so its earlier 15/15 is not a result from the same test as these new scores.

Configuration Fully successful attempts Median session duration
Flash Next medium 18/18 2 min 42 s
Flash Next xhigh 18/18 3 min 17 s
Qwen 27B Q4 xhigh 18/18 5 min 14 s
Qwen 27B Q8 xhigh 18/18 6 min 40 s

Flash medium’s median time was about 18% shorter than xhigh’s, with the same number of successful attempts. We changed reasoning effort, not model size. Q8 did not improve the success rate over Q4, and its median was about 27% longer. Q4 and Q8 ran in consecutive blocks, which limits interpretation of the timings. Every configuration passed everything: harder tasks are needed to distinguish their correctness further.

What changes with a larger token limit?

A token is a piece of text: a word, part of a word or a character. The generation limit sets the budget for producing a response. It differs from context, which sets the capacity for material the model can handle.

Separately, we tested code produced in a single response, without letting the model run files or tests. At a 16k-token limit, Flash scored 0/6; at 32k, 2/6; and at 64k, 4/6. Qwen 27B scored 1/6, 3/6 and 4/6 respectively. Results were weaker than when working as an agent with tools.

The limit also included reasoning tokens. Sometimes it ran out before any visible code appeared. A larger budget helped, but truncations and errors remained. We tested only two tasks, and the time limit increased alongside the token budget. The result cannot be attributed solely to token count or generalised to all programming work.

Free model, paid service: a tie in our test

Qwen Flash Next and GPT-6 Astra both scored 15/15 across five tasks repeated three times. That is a tie for fully successful attempts on our task set. Astra scored higher for explanations in the supplementary AI review.

A free-to-download model need not deliver worse results than a paid service. Flash Next’s weights are available without a download fee under the publisher’s license; infrastructure and operation still cost money.

To run your own comparison, explore the GPU Server Hub server offer.

Results source: our own tests conducted from 29 September to 5 October 2026. The data cover the original coding series, the later comparison of four configurations and the generation-budget experiment separately. Main illustration: AI-generated. Chart: based on recorded results.

Download the data

Public test results and settings, without credentials or administrative identifiers.