How Far Generative AI Has Come

One prompt. One L40S. Six local models, in the order they showed up. The strip is the plan. Disco Diffusion, DALL-E Mini, Stable Diffusion 1.5, Fooocus, Flux. Same ships, same storm, left to right.

GPUNVIDIALLMInference

Five images of the same burning-ships prompt, left to right Disco Diffusion, DALL-E Mini, Stable Diffusion 1.5, Fooocus, and Flux.

The scene is locked. Three 18th-century ships of the line on fire in a storm. We used this prompt on every image and every clip.

three 18th-century wooden ships of the line burning on a violent storm-black ocean, tall masts complex rigging and torn white sails engulfed in roaring orange and yellow fire, flames erupting from the decks and climbing into a sky of thick charcoal smoke and white steam, dark navy waves with crashing white foam, firelight glittering across the water and underlighting the hulls, gun ports carved wood and tattered flags visible through the smoke, wide cinematic view, ships staggered into the distance, epic chaotic age of sail naval disaster, only light is the inferno

No style tags after Disco. Seed 42 where the stack takes a seed. No ChatGPT. No Midjourney. No hosted playground. A 48 GB card is the machine this sequence was built for.

Each rung has a picture and a clip. The clip matches what that era could do. Paint-in and zoom for Disco. A walk across the Mini grid. Generated motion from Stable Diffusion 1.5 on. Native text-to-video at the end.

The order is the history. Disco and Mini first, because that is when local gen was a hobby. Stable Diffusion 1.5 next, because that is when it became a tool. Fooocus as the Massed picture stack. Flux as the current open picture. LTX-2.5 Distilled as the current open video.

Disco Diffusion

Disco is CLIP-guided paint. Fire and masts come through. The picture is storm and flame, not a sharp photo. No watermark on this image.

Disco Diffusion image of three burning wooden ships of the line on a storm sea.

Disco has no video model. The sequence here is the picture forming, then a zoom on the finished frame.

DALL-E Mini

Mini is the local DALL-E-family model. min-dalle Mega, full 4 by 4 grid, same prompt. Low detail. A few tiles find a hull. This is the grid on purpose. It is not OpenAI DALL-E 2 or 3.

Local min-dalle Mega 4 by 4 grid of the same burning-ships prompt.

Mini has no native video. The clip is the sixteen tiles in order.

Stable Diffusion 1.5

Stable Diffusion 1.5 is the break. Hulls, gun ports, fire on the water. Same GPU. Local image gen becomes something you keep.

Stable Diffusion 1.5 image of three burning ships of the line, same prompt.

AnimateDiff is the motion for this rung. Sixteen frames from that checkpoint. Weather moves. The ship holds.

Fooocus

Fooocus is the Massed picture stack. Juggernaut XL, 4-step Lightning LoRA, seed 42. The faint mark in the corner is the usual Massed ghost logo.

Fooocus Lightning image of three burning ships of the line, with a light Massed Compute watermark.

Stable Video Diffusion XT is the motion for this rung. Same picture. Fire and water start to move.

Flux

Flux schnell is the current open picture. Weights are black-forest-labs/FLUX.1-schnell. Apache. Four steps. Same prompt as Disco. Three ships in a line, firelight on the waves, 48 GB.

FLUX.1-schnell image of three burning ships of the line, with a light Massed Compute watermark.

LTX-2.5 Distilled is the motion for this rung. Image-to-video on the Flux frame. Smoke lifts. The water works.

LTX-2.5

The last rung is native video. LTX-2.5 Distilled, same prompt, no starting picture. About five seconds at 1536 by 1024. The frame below is from that clip. Flux stays the picture you would ship. LTX is the newest local video in the set.

Frame from LTX-2.5 Distilled text-to-video of the same burning-ships prompt.

Rent the GPU

The whole sequence ran on one L40S. Sign in at the Massed Compute marketplace, open Deploy, pick a 48 GB card such as gpu_1x_l40s, Ubuntu Server 24.04 with NVIDIA drivers, add your SSH key, launch.

Disco, Mini, Stable Diffusion 1.5, Fooocus, Flux schnell, and LTX Distilled all fit that class of machine for this set. Pay for the hours you use. Check live stock on GPU pricing. The least expensive 48 GB card that is in stock is the place to start.

Frequently Asked Questions

01Did any image come from ChatGPT or Midjourney?

No. Every image and clip was generated on a Massed L40S from open weights.

02Why Mini instead of DALL-E 3?

DALL-E 3 is a closed API. Mini is the DALL-E-family model you can run yourself. The grid is min-dalle Mega.

03Which Flux file is this?

black-forest-labs/FLUX.1-schnell. Four steps. Seed 42.

04Do I need Blackwell for this sequence?

No. The pictures and the LTX Distilled clip ran on one L40S with 48 GB.

05Is the Fooocus picture SDXL-Turbo?

No. Juggernaut XL plus the 4-step Lightning LoRA. That is the Fooocus Lightning stack.

Deploy How Far Generative AI Has Come

Launch NVIDIA L40 instances optimized for OpenAI-compatible inference. Get 48GB VRAM, blazing-fast NVMe storage, and per-second billing.

Think it. Build it. Scale it.