
Same prompt, generated with each model separately — click one to see how it holds up.
Why this works
The full-body framing comes directly from “candid fashion editorial composition,” while “standing on a narrow urban sidewalk” gives the figure a compressed, vertical city setting rather than an open street scene. The confident urban mood is anchored by “casual confident pose,” “weathered brick building,” and “photorealistic street photography”; “late-afternoon golden hour light” adds warm highlights that keep the beige trousers and gray-charcoal jacket from feeling flat. “Shallow depth of field” separates the man from the “rich brick textures,” preserving the wall’s tactile detail while making the clothing the visual focus.
FAQ
→How do I make the charcoal jacket and outfit more prominent?
Replace “holding a charcoal jacket in one hand” with “wearing the charcoal jacket open over the light blue checkered button-down shirt,” and add “waist-up three-quarter fashion portrait” in place of “portrait-full-body.” Keep “shallow depth of field” to reduce background competition from the weathered brick building.
→How do I shift this from confident street style to a moodier urban portrait?
Replace “late-afternoon golden hour light” with “overcast blue-hour light with deep soft shadows,” and change “casual confident pose” to “quiet, introspective pose with lowered gaze.” Retain “weathered brick building” and “rich brick textures” so the setting still carries the urban character.
→How do I create a series of variations without losing the same editorial identity?
Keep “photorealistic street photography,” “candid fashion editorial composition,” and “shallow depth of field” unchanged, then swap the clothing colors and sidewalk setting in controlled steps. For example, replace “navy baseball cap, light blue checkered button-down shirt, beige trousers” with “olive beanie, cream knit sweater, charcoal trousers,” or replace “beside a weathered brick building” with “beside a graffiti-covered concrete underpass,” while preserving “late-afternoon golden hour light” for visual continuity.
Learn the technique behind this
- Why can't AI spell, and how do I get readable text in an image? — Older diffusion models had no character-level representation of text, so they produced letterform-shaped texture instead of words.
- Which aspect ratio should I use, and how does it change the image? — Aspect ratio determines what the model composes, not how it's cropped afterwards.
- Why does AI get hands and faces wrong, and how do I fix it? — Hands fail because they're small in frame, extremely variable in pose, and self-occluding — the model has less usable signal per pixel than for any other body part.
Related prompts

Grotesque Claymation Gangsters in Gritty Alley

Charcoal coat with tan-handled tote

Woman in Teal Jacket Capturing Twilight Reflection

Stuffed Bear on Winter Dumpster at Dawn
