
Why this works
The phrase “smartphone-in-foreground framing” creates the layered candid composition, making the held phone and the two women’s selfie the visual anchor while “two tall glasses of iced coffee on a light wooden table” supplies a tactile lower foreground. “Late afternoon golden hour” and “amber-gold iced coffee” establish the warm tan and beige palette, while “pink flowering plants” and the green patio foliage add controlled color contrast. “Shallow depth of field” keeps the selfie and glassware readable while softening the umbrellas and flowers, and “crisp reflections” plus “visible condensation” provide photorealistic texture on the clear tumblers.
FAQ
→How do I make the two women more prominent than the phone and coffee glasses?
Replace “smartphone-in-foreground framing” with “women-centered medium close-up, smartphone partially cropped at the lower edge,” and change “two tall iced coffee glasses on a light wooden table” to “small blurred coffee glasses at the bottom edge.” Keep “two young women pose close together” but add “faces sharp and dominant in the frame.”
→How do I make this scene feel more relaxed and less bright or polished?
Replace “late afternoon golden hour sunlight with warm, crisp reflections” with “soft overcast afternoon light with muted reflections,” and change “vibrant lifestyle photography” to “natural candid café documentary photography.” Replace “crisp glass reflections” with “subtle reflections and gentle surface texture” while retaining “visible condensation.”
→How do I build a matching series with different café moments?
Keep the fixed visual language, including “photorealistic,” “amber-gold iced coffee,” “pale straw-colored straws,” “light wooden table,” and “shallow depth of field,” then replace the action “two young women pose close together for a smartphone selfie” with variations such as “two friends clink iced coffee glasses,” “one woman photographs the other,” or “the pair laugh while reviewing a selfie.” Swap “pink flowering plants and café patio umbrellas” for consistent alternate settings such as “striped awning and terracotta planters” or “greenery-covered courtyard and wrought-iron chairs.”
Learn the technique behind this
- Why can't AI spell, and how do I get readable text in an image? — Older diffusion models had no character-level representation of text, so they produced letterform-shaped texture instead of words.
- Why does the model ignore parts of my prompt? — Ignored instructions are almost always conflicts, counts, or spatial relationships — three things current models handle badly — rather than the model failing to read you.
- Do negative prompts work, and what should go in one? — Negative prompts work on models that support a separate negative conditioning channel — mainly Stable Diffusion and FLUX-family models.
Related prompts

Woman's Selfie with Crimson-Powered Anime Guardian

Couple selfie in warm golden light

Young woman with auburn hair holding peach bouquet

Warm indoor flashless couple selfie
