
Why this works
The phrase “overhead framing for an energetic, upbeat, playful mood” makes the full-body portrait feel active, while “centered commercial portrait composition” keeps the calling gesture immediately readable. “Bright teal shirt,” “dark indigo jeans,” and “white low-top sneakers” create a high-contrast teal-blue-white wardrobe against the “soft sage-lavender studio background,” and “vibrant studio lighting with crisp skin texture and sharp fabric detail” preserves the photorealistic commercial finish.
FAQ
→How do I change this prompt for a calmer, more editorial mood?
Replace “cheerful” and “energetic, upbeat, playful mood” with “quiet, contemplative mood,” and change “vibrant studio lighting” to “soft diffused window light with gentle shadows.” Keep “overhead framing” for the unusual viewpoint, or replace it with “eye-level three-quarter framing” for a more conventional portrait.
→How do I make the calling gesture and facial expression more prominent?
Replace “overhead framing” and “portrait-full-body” with “tight waist-up close-up, face and cupped hands dominant in frame.” Strengthen the subject wording to “cheerful Black man loudly calling out, mouth open and hands cupped around his mouth, expressive eyes,” while keeping “centered composition” to preserve the direct visual emphasis.
→How do I turn this into a series with different outfits and poses?
Keep the fixed structure “photorealistic commercial studio portrait,” “clean soft sage-lavender studio background,” and “centered commercial portrait composition,” then swap the wardrobe phrase “bright teal shirt over a black long-sleeve layer, dark indigo jeans and white low-top sneakers” for one outfit per variation. Replace “calling out with hands cupped around his mouth” with controlled actions such as “pointing upward,” “laughing with arms crossed,” or “waving with one hand,” while retaining “vibrant studio lighting” for visual consistency.
Learn the technique behind this
- Why can't AI spell, and how do I get readable text in an image? — Older diffusion models had no character-level representation of text, so they produced letterform-shaped texture instead of words.
- Which aspect ratio should I use, and how does it change the image? — Aspect ratio determines what the model composes, not how it's cropped afterwards.
- Why does AI get hands and faces wrong, and how do I fix it? — Hands fail because they're small in frame, extremely variable in pose, and self-occluding — the model has less usable signal per pixel than for any other body part.
Related prompts

Woman in Teal Light, Shadowed Fashion Portrait

Eerie Elegance Among Marble Busts

Charcoal turtleneck with red eye-beam

Dreamy Macro Portrait in Cool Water
