
Why this works
The phrase “moody low-angle portrait shot from below” makes the close-up feel imposing, while “between towering skyscrapers” supplies vertical scale around the subject. “Cool blue-gray night ambience” and “dramatic cinematic lighting with strong rim highlights, high contrast” establish the navy-black palette and carve the suit and face out of the dark background. The hand “covering part of his face,” combined with “shallow depth of field,” concentrates attention on the gesture, silver rings, and realistic skin texture rather than the surrounding street.
FAQ
→How do I make the subject’s face and jewelry more prominent?
Replace “keep his hand covering part of his face” with “hand lowered beside his cheek, face fully visible,” and change “shallow depth of field” to “tight focus on the eyes, hand, and silver rings.” Keep “fewer, sleeker bands” so the jewelry remains controlled rather than becoming visually busy.
→How do I shift this from moody night fashion to a colder, more futuristic tone?
Replace “cool blue-gray night ambience” with “icy cyan-blue nocturnal ambience with subtle architectural glow,” and change “brown” in the dominant palette toward “black, steel blue, and cyan.” Retain “sharp rim highlights” but add “thin cyan edge light on the suit and rings” for a clearly colder color treatment.
→How do I create a series of variations without losing the original composition?
Keep “moody low-angle portrait shot from below,” “portrait-closeup,” “charcoal tailored suit,” and “hand covering part of his face” unchanged, then swap only the setting phrase “between towering skyscrapers in a dense urban street setting” for variants such as “beneath a rain-soaked elevated train,” “outside a brutalist concrete plaza,” or “in a narrow neon-lit alley.” Preserve “dramatic cinematic lighting,” “strong rim highlights,” and “realistic skin texture” to keep the series visually consistent.
Learn the technique behind this
- Why can't AI spell, and how do I get readable text in an image? — Older diffusion models had no character-level representation of text, so they produced letterform-shaped texture instead of words.
- Why does AI get hands and faces wrong, and how do I fix it? — Hands fail because they're small in frame, extremely variable in pose, and self-occluding — the model has less usable signal per pixel than for any other body part.
- How should a prompt be structured, and does word order matter? — A prompt that behaves predictably names one subject first, then what it is doing, then where, then the light, then the lens or medium, then the style.
Related prompts

Grotesque Claymation Gangsters in Gritty Alley

Charcoal coat with tan-handled tote

Woman in Teal Jacket Capturing Twilight Reflection

Stuffed Bear on Winter Dumpster at Dawn
