
Same prompt, generated with each model separately — click one to see how it holds up.
Why this works
The wide shot and “view from behind the seated person at desk height” make the workspace the immediate anchor while the floor-to-ceiling windows extend the scene into a larger urban setting. “Many lush green houseplants” and the oak-toned desk establish the brown-and-green color harmony, while “dense traffic” and “tall apartment buildings” add gray urban scale without overturning the calm mood. “Soft natural daylight,” “realistic reflections in the window glass,” and “crisp textures on wood, paper, and leaves” provide the photorealistic editorial finish and keep the desk materials legible.
FAQ
→How do I make the creative desk more central and prominent?
Replace “wide-shot” with “medium-wide shot focused primarily on the desk,” and change “sharp focus on desk and plants” to “sharp focus on the laptop, sketches, open books, paintbrushes, and hands.” Reduce the exterior emphasis by replacing “looking out onto a busy city avenue with dense traffic” with “the city avenue visible as a softly blurred background.”
→How do I change the calm urban mood into a warmer, more intimate evening scene?
Replace “soft natural daylight” and “natural color palette” with “warm late-afternoon window light with amber highlights and gentle interior shadows.” Change “dense traffic” to “city lights beginning to glow below,” while keeping “realistic reflections on the window glass” so the exterior remains convincing.
→How do I build a series of variations without losing the identity of this image?
Keep the fixed anchors “person viewed from behind,” “cluttered oak-toned wooden desk,” “laptop,” “many lush green houseplants,” and “floor-to-ceiling windows,” then swap only the setting phrase. For example, replace “busy city avenue with dense traffic, tall apartment buildings, and patches of urban park greenery” with “rainy residential street with glowing shopfronts,” “sunlit courtyard garden,” or “snow-covered rooftops,” and adjust the lighting phrase to match each location.
Learn the technique behind this
- Why can't AI spell, and how do I get readable text in an image? — Older diffusion models had no character-level representation of text, so they produced letterform-shaped texture instead of words.
- Why does AI get hands and faces wrong, and how do I fix it? — Hands fail because they're small in frame, extremely variable in pose, and self-occluding — the model has less usable signal per pixel than for any other body part.
- How do I keep one consistent style across a whole set of images? — Style consistency is more achievable than character consistency because style lives in describable attributes.
Related prompts

Young woman with two cats

Blonde woman hugging beige pillow on off-white bed

Cozy Winter Living Room in Sage

Young woman holding teal corded phone on pastel bed
