Cost versus throughput for LFM2.5-VL-3B on A6000, L40S, and RTX PRO 6000 Blackwell.

Do You Need Blackwell for LFM2.5-VL-3B?

Rent the RTX A6000 for LiquidAI LFM2.5-VL-3B. It is the least expensive card that ran the exact BF16 weights, and it wins dollars per token.

GPUNVIDIAA6000L40SBlackwellLFMvLLMInference

That is the renter decision. Blackwell is faster. It is not the better buy for this 3B vision-language model.

Live Massed Compute list rates on 2 September 2026 match the published bench. A6000 is $0.57/hr. L40S is $0.88/hr. RTX PRO 6000 Blackwell is $2.19/hr. Confirm current GPU pricing before you launch.

The numbers that decide the card

Same checkpoint. Same engine. Same serving flags. vLLM v0.27.1 at concurrency 32. Exact BF16. Text-only decode profile. Capture date 25 August 2026. Full method and screenshots are on the LFM2.5-VL-3B GitHub page.

SKU List $/hr Output tok/s (c32) TTFT median tok/s per $ $/1M output tokens
RTX A6000 48GB $0.57 2593.0 253.9 ms 4549.2 $0.061
L40S 48GB $0.88 3112.0 77.7 ms 3536.4 $0.079
RTX PRO 6000 Blackwell 96GB $2.19 5442.2 72.7 ms 2485.0 $0.112

A6000 is the value card. Blackwell is the speed card. Those two winners are not the same GPU. That disagreement is the post.

Cost versus throughput for LFM2.5-VL-3B on A6000, L40S, and RTX PRO 6000 Blackwell at live list rates.

The chart plots list dollars per hour against captured output tokens per second. A6000 sits left and still high enough. Blackwell sits far right. You pay 3.8 times the A6000 hourly to get 2.1 times the tokens.

What extra money buys on L40S

L40S is 20% faster than A6000 (3112 vs 2593 tok/s). It costs 54% more per hour ($0.88 vs $0.57). Tokens per dollar fall from 4549 to 3536. Cost per million output tokens rises from $0.061 to $0.079.

The real purchase is latency. Median time to first token drops from 253.9 ms on A6000 to 77.7 ms on L40S. That is the pause after send. Around 100 ms reads as instant. Past a quarter second, people notice they are waiting.

Pay the L40S premium when this 3B job is interactive and A6000’s 254 ms wait is the problem. Do not pay it for batch throughput. A6000 already wins that on dollars.

What extra money buys on Blackwell

Blackwell is the throughput card. 5442.2 tok/s. 2.1 times A6000. 1.8 times L40S.

The hourly is $2.19. That is 3.8 times A6000 and 2.5 times L40S. Tokens per dollar are the worst of the three at 2485. Cost per million output tokens is $0.112.

Look at first token. Blackwell is 72.7 ms. L40S is already 77.7 ms. You are buying 5 milliseconds. L40S already captured almost all of the latency win.

The 96 GB on Blackwell is unused headroom for this model. The BF16 weights are about 6 GB. They fit 48 GB with room to spare. vLLM reserved most of the card because the serve used --gpu-memory-utilization 0.92, not because the 3B model needs 88 GB. On a later same-profile check, A6000 held about 42 GB of 48 GB, L40S about 39 GB of 45 GB, and Blackwell about 88 GB of 96 GB. That reservation is cache headroom. It is not a fit requirement.

SGLang was attempted on Blackwell. The serving bench could not write its ShareGPT cache and produced no JSON. This article is vLLM only. No quantization. Same weights on every row.

Where latency and dollars disagree

Dollars say A6000. First-token latency says L40S is enough. Raw throughput says Blackwell.

If you are filling a queue, A6000 gives you the most tokens per dollar. If you are waiting on the first word, L40S is the upgrade that actually changes how the app feels. Blackwell is the card you rent when this 3B model is not the point and you already needed 96 GB or maximum tokens per second for some other reason.

You do not need Blackwell for a 3B vision-language model that already runs on 48 GB.

How to launch

  1. Sign in to the Massed Compute marketplace and open Deploy.
  2. Pick gpu_1x_a6000 unless you have a latency reason to take gpu_1x_l40s.
  3. Choose an Ubuntu image with NVIDIA drivers.
  4. Add your SSH key. Launch.

Pay only for the hours you use. Confirm the live rate on GPU pricing at launch. The three SKUs in this article were in stock when we checked rates on 2 September 2026.

Frequently Asked Questions

01Do I need Blackwell for LFM2.5-VL-3B?

No. The exact BF16 weights ran on A6000 at $0.57/hr. Blackwell is faster. It loses on tokens per dollar.

02Which card is best value?

A6000. 4549 tok/s per dollar. Lowest cost per million output tokens at $0.061.

03What does L40S buy that A6000 does not?

Mostly first token. 77.7 ms vs 253.9 ms. Throughput only rises 20% while the hourly rises 54%.

04When should I rent Blackwell for this model?

When the job is latency-bound and L40S is not enough, or when you already need the 96 GB card for another workload on the same box. Not because a 3B model requires it.

05Was this an image-generation bench?

No. It is a vision-language model served with a text-only decode profile. Do not read these tok/s as image or video generation speed.

06Did you quantize the model?

No. Exact BF16 on every SKU. SGLang on Blackwell produced no result, so the comparison is vLLM only.

Run LFM2.5-VL-3B on the card that wins dollars

Launch an NVIDIA RTX A6000 for this 3B vision-language model. Get 48GB VRAM, live list-rate pricing, and per-second billing.

Think it. Build it. Scale it.