Rent the RTX A6000 for packed Spark-X2.5-4B serving and for unattended MiniMax-H3 Turbo clips. It is the least expensive card that actually ran both ladders, and it wins dollars per token and dollars per clip.
That is the renter decision. Blackwell is faster on both models. It is not the better buy unless you are sitting on the wait.
Live Massed Compute list rates on 11 September 2026 still match the 8 September capture. A6000 is $0.57/hr. L40S is $0.97/hr. RTX PRO 6000 Blackwell is $2.19/hr. L40S moved from $0.88 to $0.97 on 8 September 2026; these pages use $0.97. Confirm current GPU pricing before you launch.
The numbers that decide the card
Two jobs. Same three SKUs. Value and speed are not the same GPU.
Spark-X2.5-4B (XHToken/Spark-X2.5-4B BF16, vLLM 0.28.0 plus the spark2_5 plugin, random 128/128 at concurrency 32). A6000 delivers 2825 tok/s per dollar. Blackwell delivers 3422.9 tok/s, 2.1× A6000, and loses on tokens per dollar. L40S is only +20% tok/s for +70% $/hr. Skip it for packed serving.
MiniMax-H3 Turbo (lightx2v/Minimax-h3-Turbo on Comfy-Org MiniMax-H3 INT8, minimax_h3_fl2v_turbo_4step_v1.2_768p_comfyui_bf16.safetensors). A6000 is $0.026 per warm 4-step 768p clip. Blackwell waits 56.6 s, 2.9× A6000 / 1.75× L40S, at $0.034/clip. Peak VRAM is ~40 GiB on every card. Native LightX2V was the wrong engine for this renter path.
Spark-X2.5-4B: packed tokens
Same checkpoint. Same vLLM 0.28.0 plus the spark2_5 plugin. Random 128/128 at concurrency 32. Exact BF16. Capture date 8 September 2026. Native 1M context was not benched. Method and showcases: Spark-X2.5-4B GPU bench (PR).
| SKU | List $/hr | Output tok/s (c32) | TTFT median | tok/s per $ | $/1M output tokens |
|---|---|---|---|---|---|
| RTX A6000 48GB | $0.57 | 1610.4 | 180.4 ms | 2825.2 | $0.098 |
| L40S 48GB | $0.97 | 1934.2 | 128.0 ms | 1994.1 | $0.139 |
| RTX PRO 6000 Blackwell 96GB | $2.19 | 3422.9 | 98.1 ms | 1563.0 | $0.178 |
Cheapest-fit that ran is gpu_1x_a6000. Mid is gpu_1x_l40s. Premium is gpu_1x_pro_6000_blackwell. On the capture date A30 ($0.35) and A5000 ($0.44) had no capacity and were not launched. On 11 September 2026 A30 showed capacity; A5000 still had none. gpu_1x_a6000_low_ram ($0.55) and gpu_1x_a6000_spot ($0.50) were not launched. Those cheaper SKUs are untested for this job.
A6000 is the value card. Blackwell is the speed card. Those two winners are not the same GPU.
What extra money buys on L40S
L40S is 20% faster than A6000 (1934.2 vs 1610.4 tok/s). It costs 70% more per hour ($0.97 vs $0.57). Tokens per dollar fall from 2825 to 1994. Cost per million output tokens rises from $0.098 to $0.139. Skip L40S for packed serving.
What extra money buys on Blackwell
Blackwell is 2.1× A6000 and 1.8× L40S on packed tok/s. The hourly is 3.8× A6000. Tokens per dollar are the worst of the three. First token does improve (180.4 ms → 128.0 ms → 98.1 ms). Pay Blackwell when this job is interactive and 180 ms after send is the problem, or when the queue is tok/s-bound and 3422.9 tok/s is the point. Do not pay it for dollars per token.
The 4.1B BF16 weights are about 8.3 GB. They fit 48 GB. nvidia-smi at serve-ready held 43.38 / 47.99 GiB on A6000, 40.83 / 44.99 GiB on L40S, and 87.86 / 95.59 GiB on Blackwell (MiB / 1024). That reservation is cache headroom from the serve, not a fit requirement.
Stock vllm/vllm-openai could not load Spark2_5ForCausalLM until the plugin. SGLang was not captured this wave.
MiniMax-H3 Turbo: one 5 s clip
Same lock on every SKU: 1344×768, 4 steps, Euler, 124 frames, seed 42. ComfyUI core MiniMax-H3 nodes, not native LightX2V. Default Comfy unloads between jobs, so the second clip is not a hot-weight shortcut. Capture date 8 September 2026. Method and clips: MiniMax-H3 Turbo GPU bench (PR).
| SKU | List $/hr | 1st clip (s) | 2nd clip (s) | Peak / total VRAM | Host RAM | $/clip |
|---|---|---|---|---|---|---|
| RTX A6000 48GB | $0.57 | 166.382 | 164.204 | 39.45 / 48.0 GiB | 48 GiB | $0.026 |
| L40S 48GB | $0.97 | 103.357 | 99.303 | 39.68 / 45.0 GiB | 72 GiB | $0.027 |
| RTX PRO 6000 Blackwell 96GB | $2.19 | 52.486 | 56.618 | 39.88 / 95.6 GiB | 144 GiB | $0.034 |
Cheapest-fit that ran is gpu_1x_a6000. Mid is gpu_1x_l40s. Premium is gpu_1x_pro_6000_blackwell. The cheaper live SKUs named above were not launched for Turbo either.
A6000 is the unattended value card at $0.026/clip. Blackwell is the attended wait card at 56.6 s (2.9× A6000, 1.75× L40S). L40S is $0.027/clip and 99.3 s. It is not the batch card and not the wait card. Skip it.
Peak VRAM is ~40 GiB on every card. Blackwell leaves ~56 GiB idle. You are buying wall time, not H3 memory. A6000’s 48 GiB host RAM loads this ComfyUI path. Native LightX2V needed far more host RAM and skipped A6000 / L40S. Those older numbers are not this table.
One hundred unattended clips: A6000 $2.60 (4.56 h), L40S $2.68 (2.76 h), Blackwell $3.44 (1.57 h). Pay the extra $0.008 per clip on Blackwell when you are sitting on a single render. Do not pay it for a queue. That extra money buys back 107 s on one warm clip, not a cheaper batch.
Where speed and dollars disagree

The left plot is Spark packed tok/s against list dollars per hour. A6000 sits left. Blackwell sits far right. You pay 3.8 times the A6000 hourly to get 2.1 times the tokens.
The right plot is MiniMax-H3 Turbo second-clip wait against dollars per clip. A6000 is slow and least expensive per clip. Blackwell is fast and most expensive per clip. L40S sits in the middle on both charts and wins neither job.
| Workload | Speed winner | Value winner | When premium is justified | When premium is waste |
|---|---|---|---|---|
| Spark-X2.5-4B packed serve | Blackwell, 3422.9 tok/s | A6000, $0.098 / 1M out | Interactive first token, or the queue is tok/s-bound | Dollars per token, batch fill, L40S as a “middle” upgrade |
| MiniMax-H3 Turbo 4-step 768p | Blackwell, 56.6 s | A6000, $0.026/clip | You are waiting on one clip and $0.008 extra is worth 107 s back | Unattended queue, L40S for either goal |
How to launch
- Sign in to the Massed Compute marketplace and open Deploy.
- Pick
gpu_1x_a6000unless you have a wait reason to takegpu_1x_pro_6000_blackwell. Do not pickgpu_1x_l40sfor these two jobs. - Choose an Ubuntu image with NVIDIA drivers.
- Add your SSH key. Launch.
Pay only for the hours you use. Confirm the live rate on GPU pricing at launch. The three SKUs in this article were the ladder we ran on 8 September 2026.
Frequently Asked Questions
01Do I need Blackwell for Spark-X2.5-4B?
No, not for dollars per token. Exact BF16 ran on A6000 at $0.57/hr and 2825 tok/s per dollar. Blackwell is 3422.9 tok/s. It loses on tokens per dollar.
02Do I need Blackwell for MiniMax-H3 Turbo?
Only if you are waiting on one clip. Unattended, A6000 is $0.026/clip vs Blackwell $0.034. Peak VRAM is ~40 GiB on every card.
03Which card is best value?
A6000 on both models. Spark: $0.098 per million output tokens. Turbo: $0.026 per 4-step 768p clip.
04What does L40S buy?
Not enough. Spark: +20% tok/s for +70% $/hr. Turbo: $0.027/clip and 99.3 s, which is neither the batch win nor the wait win.
05When should I rent Blackwell?
Spark: first-token or packed tok/s is the product, and A6000’s 180 ms / 1610 tok/s is the bottleneck. Turbo: one attended clip, and 56.6 s vs 164.2 s is worth $0.008.
06Was Spark a 1M-context bench?
No. Pinned random 128/128 at c32. Do not read these tok/s as long-context speed.
07Did you quantize Spark?
No. Exact BF16 on every SKU. Serving needed the vLLM Spark 2.5 plugin.
08Is this the old LightX2V MiniMax table?
No. Native LightX2V was the wrong engine for this renter path. This table is ComfyUI 4-step Turbo 768p.
Rent the card that matches the wait
Pick gpu_1x_a6000 for packed Spark serving and unattended MiniMax-H3 Turbo clips. Take Blackwell only if you are sitting on the wait.
Think it. Build it. Scale it.











