
Same prompt, generated with each model separately — click one to see how it holds up.
Why this works
The phrase “silhouetted human figure tumbling through a dramatic vertical city canyon” creates the central falling action and strong scale contrast, while “slightly wider perspective emphasizing the buildings’ height and depth” keeps the skyscrapers dominant in the frame. “Cool blue-teal atmospheric haze with faint violet accents” establishes the restrained futuristic palette, and “subtle volumetric light beams slicing through the canyon” adds visible depth between the towers. The surreal tone comes from combining “dreamlike composition,” “moody futuristic mood,” and the physically improbable vertical fall rather than from the color grading alone.
FAQ
→How do I make the falling figure more prominent?
Replace “silhouetted human figure” with “large foreground human figure in sharp silhouette, occupying the central third of the frame,” and change “slightly wider perspective” to “medium-wide perspective with the figure close to camera.” Keep “tumbling through” if you want motion, or use “arms and coat trailing in a frozen mid-fall pose” for a clearer outline.
→How do I shift this from a cold futuristic mood to a warmer, more ominous scene?
Replace “cool blue-teal atmospheric haze with faint violet accents” with “smoky amber-orange haze with deep crimson accents,” and change “moody futuristic color grading” to “ominous industrial color grading.” For harsher tension, replace “subtle volumetric light beams” with “narrow red searchlights cutting through dense smoke.”
→How do I build a series with different locations while preserving the same composition?
Keep “silhouetted human figure tumbling” and “slightly wider perspective emphasizing depth and building height,” then replace “towering modern skyscrapers forming a dramatic vertical city canyon” with a specific location such as “massive Art Deco towers forming a narrow dusk canyon,” “rain-soaked neon megacity towers,” or “vertical canyon of brutalist concrete megastructures.” Retain “photorealistic cinematic urban scene” and “no text or symbols” to keep the visual treatment consistent across variations.
Learn the technique behind this
- Why can't AI spell, and how do I get readable text in an image? — Older diffusion models had no character-level representation of text, so they produced letterform-shaped texture instead of words.
- How should a prompt be structured, and does word order matter? — A prompt that behaves predictably names one subject first, then what it is doing, then where, then the light, then the lens or medium, then the style.
- Which aspect ratio should I use, and how does it change the image? — Aspect ratio determines what the model composes, not how it's cropped afterwards.
Related prompts

Grotesque Claymation Gangsters in Gritty Alley

Charcoal coat with tan-handled tote

Woman in Teal Jacket Capturing Twilight Reflection

Red Gummy Bear Monster Attacks
