
Same prompt, generated with each model separately — click one to see how it holds up.
Why this works
The phrase 'medium-resolution laptop camera aesthetic' combined with 'slightly soft focus with natural grain' is what sells the intimacy here, it degrades the image just enough to feel like a real webcam capture instead of a studio shot. 'Warm evening lighting' paired with the desk lamp detail in the setting description ties the amber cast to a visible, plausible light source rather than an arbitrary color grade. The 'vertically stacked collage of 3 horizontal frames' plus 'consistent identity across all frames' is the structural instruction doing the heavy lifting on composition, forcing the model to treat this as one continuous session rather than three unrelated portraits. And naming distinct moods per frame, 'contemplative, animated, peaceful, or playful', is what keeps the three panels from collapsing into repetitive near-duplicates of the same expression.
FAQ
→How do I make the lighting feel colder or more clinical instead of cozy?
Swap 'warm evening lighting with soft amber tones' for something like 'cool blue-white monitor glow with harsh overhead lighting,' and drop 'desk lamp' from the setting in favor of 'overhead fluorescent panel.' That kills the amber warmth and pushes the mood toward sterile or isolating rather than intimate.
→How do I get more than 3 frames or change the panel layout?
Change 'vertically stacked collage of 3 horizontal frames' to a number and arrangement like 'a 2x2 grid of 4 square frames' or 'a horizontal filmstrip of 5 frames.' You'll also want to expand the mood list to match, since right now only four moods are listed for three frames, so add one mood per new panel to keep each frame distinct.
→How do I shift this from candid webcam style to something more cinematic?
Replace 'Medium-resolution laptop camera aesthetic' and 'natural candid laptop photography style' with 'shot on 35mm film with shallow depth of field' or 'anamorphic lens with soft bokeh.' You'll lose the low-fi webcam grain, so also drop 'natural grain' if you want a cleaner, more polished look rather than the raw documentary feel.
Learn the technique behind this
- Why can't AI spell, and how do I get readable text in an image? — Older diffusion models had no character-level representation of text, so they produced letterform-shaped texture instead of words.
- How do I keep one consistent style across a whole set of images? — Style consistency is more achievable than character consistency because style lives in describable attributes.
- Why does the model ignore parts of my prompt? — Ignored instructions are almost always conflicts, counts, or spatial relationships — three things current models handle badly — rather than the model failing to read you.
Related prompts

Woman's Selfie with Crimson-Powered Anime Guardian

Blue Tracksuits on Glass Bridge Selfie

Golden-hour café selfie with iced coffee

Warm indoor flashless couple selfie
