
Why this works
“Confident man,” “relaxed posture,” and the fitted “navy polo shirt and tan chino shorts” establish a composed lifestyle mood, while “candid street photography” keeps it natural rather than posed. The “outdoor café table on a city sidewalk,” “passing cars,” and “tall street trees” provide urban context, with “portrait-full-body” making the clothing and seated stance central. “Shallow depth of field” separates him from the street, while “golden-hour natural daylight” supplies the warm gold tones that balance the navy, beige, and gray palette.
FAQ
→How do I make the man’s expression and presence more central?
Replace “confident man” with “confident man looking directly into the camera with a subtle half-smile,” and add “medium close-up portrait” while removing “portrait-full-body.” Keep “shallow depth of field” so the face becomes the sharpest visual anchor.
→How do I make the image feel more energetic and less relaxed?
Replace “relaxed posture” with “dynamic seated pose, leaning forward with one arm resting on the café table,” and change “passing cars” to “busy traffic with slight motion blur.” Replace “golden-hour natural daylight” with “crisp late-afternoon sunlight with stronger directional contrast” for a livelier street atmosphere.
→How do I create a related variation in a different urban setting?
Replace “outdoor café table on a city sidewalk” with “rooftop café terrace overlooking a dense downtown skyline,” and change “passing cars and tall street trees” to “glass office towers, rooftop planters, and distant traffic.” Retain “fitted navy polo shirt,” “tan chino shorts,” “candid street photography,” and “shallow depth of field” to preserve the subject and visual identity.
Learn the technique behind this
- Why can't AI spell, and how do I get readable text in an image? — Older diffusion models had no character-level representation of text, so they produced letterform-shaped texture instead of words.
- How should a prompt be structured, and does word order matter? — A prompt that behaves predictably names one subject first, then what it is doing, then where, then the light, then the lens or medium, then the style.
- Why does the model ignore parts of my prompt? — Ignored instructions are almost always conflicts, counts, or spatial relationships — three things current models handle badly — rather than the model failing to read you.
Related prompts

Grotesque Claymation Gangsters in Gritty Alley

Charcoal coat with tan-handled tote

Woman in Teal Jacket Capturing Twilight Reflection

Stuffed Bear on Winter Dumpster at Dawn
