Which GPU for Xing4.0 and Qwen-Image-2.1

Xing4.0-29B-A4B and Qwen-Image-2.1 do not share a scoreboard. One number is output tokens per second. The other is seconds per still. Rent from the wrong one and you pay for a card that won a different race.

These figures are live Massed Compute runs from 2026-09-29. Rates are that day’s list. Check current GPU pricing before you launch.

Technology Stack

Xing is XingChen-AGI/Xing4.0-29B-A4B. Exact BF16, Apache 2.0. About 31.2 billion parameters on disk, about 4 billion active, 64 experts, four of them used on a token.

Every Xing card ran the same vendor vLLM image and the same flags. Prefix cache off. No speculative decoding. Random prompts, 128 tokens in, 128 out. The published number is output token throughput at 32 concurrent requests.

Qwen-Image is Qwen/Qwen-Image-2.1. Official diffusers weights in BF16, not a smaller fork, under the Qwen Research License. The stack is a 7B diffusion transformer plus a Qwen3-VL text encoder.

Each resolution gets a 2-step warmup, then five timed calls at 40 steps. The clock wraps the pipeline after the GPU syncs. The published size is 1024×1024. The CUDA wheel is cu126 on the A6000 and the L40S, and cu128 on RTX PRO 6000 Blackwell. Same pipeline and step count. Not the same binary.

Xing4.0 at 32 concurrent requests

The A100 is the smallest single GPU in this test that held the BF16 pack, and it leads the published row. $1.35 an hour, 833.3 output tokens per second, 617.2 tokens per second per dollar, $0.450 per million output tokens. Median time to first token is 399.5 ms.

GPU $/hr Output tok/s Median TTFT
A100 $1.35 833.3 399.5 ms
RTX PRO 6000 Blackwell $2.19 745.0 224.2 ms
H100 $2.73 710.6 313.5 ms

Blackwell is the quickest first token in that table, 224.2 ms, at 745.0 tokens per second. The H100 is 710.6 tokens per second and 313.5 ms to the first token, at $2.73 an hour. At this load you pay more per hour than the A100 and get fewer output tokens.

One request is a different race. The H100 does 110.2 output tokens per second. Blackwell does 105.5. The A100 does 88.6. If a single user is the job, look at the H100 before the A100. If the queue is full, the A100 is the row we publish.

An earlier same-day capture on Blackwell stalled at 417.1 tokens per second and 4516.0 ms median time to first token. That run is not the table.

Qwen-Image-2.1, seconds per still

The still below is the Blackwell 1024 frame from that run. Same prompt on every card.

Ceramic cup on a sunlit wooden table, the 1024 still from the Blackwell Qwen-Image-2.1 run

At 1024, Blackwell is the speed card: 9.278 seconds mean, 0.108 images per second, peak VRAM 39.41 GiB. That is 1.78× the L40S and 3.18× the A6000. It is not the value card. The still costs $0.00564.

The L40S takes 16.517 seconds, 0.061 images per second, peak VRAM 39.14 GiB, at $0.97 an hour. Among the three cards in the table, that is the lowest cost per still, $0.00445. The A6000 takes 29.480 seconds, 0.034 images per second, peak VRAM 39.28 GiB, at $0.57 an hour, and the still costs $0.00467 because the card holds the job longer.

GPU $/hr 1024 mean $ per still 2048
A6000 $0.57 29.480 s $0.00467 did not fit
L40S $0.97 16.517 s $0.00445 did not fit
RTX PRO 6000 Blackwell $2.19 9.278 s $0.00564 51.053 s

Full BF16 at 1024 used about 39 GiB, so 32 GB cards were not launched. Two lower A6000 list rates were also not launched: spot at $0.50 an hour, and low-RAM at $0.55 an hour. The L40S figure is the lowest cost per still of the three cards that ran, not a claim about every SKU on the marketplace.

2048 does not fit on the 48 GB A6000 or L40S. Both failed the generation after 1024 had already succeeded. No CPU offload. Blackwell finishes 2048 in 51.053 seconds, peak VRAM 64.44 GiB, $0.0311 per still.

A same-day DGX A100 run, same runner and same 40 steps, is not in the table. At $1.38 an hour its 1024 mean was 14.783 seconds and $0.00567 per still, more than the L40S row. Its 2048 mean was 77.988 seconds and $0.0299 per still, slower than Blackwell and less per still. A same-day L40 at $0.86 an hour took 27.647 seconds at 1024 and $0.00660 per still, more than the A6000 and the L40S, and 2048 did not fit. 1536 fit on Blackwell only, 25.650 seconds, $0.0156 per still.

Which card to rent

Rent the A100 when Xing is under load and the number you need is output tokens per second. $1.35 an hour, 833.3 tokens per second, and the most tokens per dollar in the Xing table.

Look at the H100 when Xing is one request at a time. It leads that slice, 110.2 tokens per second. Look at Blackwell when you care about time to first token on the busy run. 224.2 ms is the lowest median in the Xing table. Neither of those is the published headline.

Rent Blackwell when the Qwen still has to come back soon, or when you need 1536 or 2048. Rent the L40S when the job is 1024 and you want the lowest cost per still of the three cards that ran.

Frequently Asked Questions

01Is the A100 the fastest Xing card at every load?

No. At one request the H100 leads, 110.2 output tokens per second against 88.6 on the A100. At 32 concurrent requests the A100 leads, 833.3 against 710.6.

02Why is the L40S the lowest cost per still if the A6000 hour is less?

The A6000 hour is $0.57 and the L40S hour is $0.97. The A6000 takes 29.480 seconds and the L40S takes 16.517. The still on the L40S costs $0.00445. The still on the A6000 costs $0.00467.

03Will a 48 GB card make a 2048 image with this model?

Not in this test. The A6000 and the L40S ran out of memory. 1024 had already used about 39 GiB. Blackwell completed 2048 in 51.053 seconds.

04Can I treat the Xing row and the Qwen row as one ranking?

No. The weights, the engine, and the clock are different. Xing is output token throughput. Qwen-Image is mean seconds per still.

Launch the card that matches the job

A100 for Xing under load. L40S for the lowest cost per 1024 still in this table. RTX PRO 6000 Blackwell when the still has to come back soon, or when you need 1536 or 2048.

Think it. Build it. Scale it.