
Why this works
The full-body framing is anchored by “full-body view” and “portrait-full-body,” keeping the man’s relaxed stride and hands visible rather than cropping the gesture. The warm, cinematic mood comes specifically from “warm rim light from the setting sun,” “golden sunset glow,” and “calm relaxed mood,” while “soft street bokeh behind him” and “shallow depth of field” separate the charcoal, gray, and brown clothing from the urban background. The restrained palette is reinforced by “charcoal bomber jacket,” “light gray crew-neck t-shirt,” “blue jeans,” and “dark brown leather boots,” giving the image grounded contrast without overpowering the subject.
FAQ
→How do I make the man feel more dramatic and confident?
Replace “calm relaxed mood” with “confident, dramatic mood,” change “walking alone” to “striding purposefully toward the camera,” and replace “warm rim light” with “hard golden backlight with strong shadows.” Keep “full-body view” so the posture and boots remain part of the composition.
→How do I make his face and upper-body styling more prominent?
Replace “full-body view” and “portrait-full-body” with “three-quarter portrait from mid-thigh up,” and change “shallow depth of field” to “very shallow depth of field focused on his eyes.” Retain “short wavy dark hair” and the clothing description, but add “subtle confident expression” to make the facial styling carry more visual weight.
→How do I create a rainy-night variation of this street scene?
Replace “late afternoon,” “setting sun,” and “golden sunset glow” with “rainy blue hour after dark, wet pavement reflecting neon signs.” Change “warm rim light from the setting sun” to “cool neon rim light with reflected streetlights,” and replace “soft street bokeh” with “large blurred neon bokeh in the background.”
Learn the technique behind this
- Why can't AI spell, and how do I get readable text in an image? — Older diffusion models had no character-level representation of text, so they produced letterform-shaped texture instead of words.
- Why does AI get hands and faces wrong, and how do I fix it? — Hands fail because they're small in frame, extremely variable in pose, and self-occluding — the model has less usable signal per pixel than for any other body part.
- How should a prompt be structured, and does word order matter? — A prompt that behaves predictably names one subject first, then what it is doing, then where, then the light, then the lens or medium, then the style.
Related prompts

Grotesque Claymation Gangsters in Gritty Alley

Charcoal coat with tan-handled tote

Woman in Teal Jacket Capturing Twilight Reflection

Stuffed Bear on Winter Dumpster at Dawn
