If you already know the subject you want to generate but still get flat, inconsistent, or strangely composed results, the missing piece is often visual language. This cheat sheet turns common photography and art direction terms into prompt-ready building blocks you can reuse across models and projects. Instead of treating prompts as one long guess, you will learn how to describe camera angle, lens behavior, lighting setup, composition, texture, and style in a way that is easier to test, revise, and save as part of a repeatable AI art workflow.
Overview
An effective image prompt usually does more than name a subject. It also gives the model instructions about how the image should feel, how the scene is framed, what kind of light is present, what level of realism is expected, and which details matter most. That is why a short prompt like “woman in a cafe” often underperforms compared with a structured version such as “editorial portrait of a woman in a quiet cafe, medium shot, 50mm lens look, soft window light, shallow depth of field, natural skin texture, muted color grading, clean background.”
This article is designed as an evergreen AI image prompt cheat sheet. You can bookmark it and return whenever you need sharper language for text to image prompts, especially if you work across tools that interpret visual terms slightly differently. The goal is not to memorize every token. The goal is to build a small, reliable vocabulary you can combine quickly.
Think of prompt engineering for images as stacking layers:
- Subject: what is in the image
- Medium or style: photo, illustration, anime, poster, cinematic frame, product render
- Camera and lens cues: shot type, angle, focal length, depth of field
- Lighting: soft, hard, backlit, rim light, studio, golden hour
- Composition: centered, rule of thirds, symmetrical, close-up, wide shot
- Surface detail: texture, skin detail, fabric, reflections, grain
- Color and mood: warm palette, desaturated, high contrast, moody
- Constraints: clean background, no text, no watermark, no extra fingers
For many creators, this is the difference between random iteration and controlled iteration. If you need a broader reusable structure, pair this reference with Text-to-Image Prompt Formula: A Reusable Structure for More Consistent AI Images. If your outputs are cluttered or distorted, the companion guide Negative Prompt Guide for AI Art: What to Exclude for Cleaner Image Outputs is the natural next step.
One important note: visual tokens are interpreted differently by different systems. A phrase that works well in one model may need fewer words, stronger weighting, or simpler phrasing in another. So treat this sheet as a practical vocabulary library, not a rigid formula.
Template structure
Use the following base template whenever you need a clean starting point for AI image prompt engineering:
[subject] + [environment] + [shot type] + [camera/lens cues] + [lighting] + [composition] + [style/material detail] + [color/mood] + [quality constraints]
Example:
Skincare product bottle on a stone pedestal, minimal studio set, close-up product shot, 85mm lens look, soft diffused studio light, centered composition, glossy reflections, premium editorial style, warm neutral palette, clean background, high detail
Below is the cheat sheet by category.
Camera and shot terms
These terms help models understand viewpoint and framing. They are especially useful for photorealistic AI prompts and cinematic scenes.
- Close-up: tight framing on a face, object, or detail
- Extreme close-up: very tight crop, often for eyes, hands, textures, product details
- Medium shot: subject framed from waist or torso upward
- Wide shot: shows subject in full and includes environment
- Establishing shot: emphasizes setting and context
- Overhead shot: top-down view, common in food, desk, and layout images
- Low angle: camera looks upward; can make the subject feel dominant
- High angle: camera looks down; can make the subject feel smaller or observational
- Eye-level shot: neutral and natural perspective
- Dutch angle: tilted frame for tension or unease
- Candid: less posed, more natural body language
- Editorial portrait: polished, magazine-style portraiture
- Product shot: clean, commercial framing for objects
- Macro: highly detailed close-up of small subjects
Lens and depth terms
These are common camera terms for AI prompts. They often shape perspective more than many users expect.
- 24mm: wide perspective, more environmental context, potential edge distortion
- 35mm: natural wide look, useful for lifestyle scenes
- 50mm: balanced and familiar, a strong default for portraits and editorial images
- 85mm: flattering portrait look, compressed perspective
- 135mm: stronger compression, isolated subject feel
- Shallow depth of field: blurred background, stronger subject separation
- Deep focus: more of the scene remains sharp
- Bokeh: soft out-of-focus background highlights
- Telephoto compression: background appears closer to the subject
- Wide-angle distortion: can exaggerate scale and perspective
If a model overreacts to specific focal lengths, simplify. For example, replace “shot on 85mm” with “portrait lens look” or “compressed portrait perspective.”
Lighting terms
Lighting is one of the highest-leverage prompt categories. Good lighting prompts for AI art can make a generic subject feel premium.
- Soft light: gentle shadows, flattering skin, common for beauty and editorial work
- Hard light: strong shadows, crisp edges, dramatic contrast
- Diffused light: soft, even illumination with reduced harshness
- Window light: natural indoor light, often directional and calm
- Golden hour: warm low-angle sunlight near sunrise or sunset
- Blue hour: cooler ambient light after sunset or before sunrise
- Backlit: light source behind subject, useful for glow and silhouette effects
- Rim light: edge lighting around subject for separation
- Rembrandt lighting: classic portrait pattern with shaped facial shadow
- Split lighting: half the face lit, half in shadow
- Studio strobe: polished commercial look
- Neon lighting: colored practical light, cyberpunk or nightlife mood
- Volumetric light: visible light beams or atmosphere
- Low-key lighting: darker frame, moody, selective illumination
- High-key lighting: bright, airy, low shadow contrast
Composition terms
Composition language helps reduce the “everything everywhere” problem that many text-to-image systems produce by default.
- Centered composition: subject placed in the middle for clarity and symmetry
- Rule of thirds: more dynamic placement off-center
- Symmetrical composition: balanced, formal, often architectural or minimalist
- Leading lines: visual lines guide the eye toward the subject
- Negative space: open space around subject, useful for ads and thumbnails
- Foreground framing: scene elements frame the subject
- Layered depth: foreground, midground, background are all visible
- Minimal composition: fewer objects, less clutter
- Dynamic composition: stronger diagonals, movement, energy
- Balanced composition: visually stable, even weight distribution
Style and finish terms
This category tells the model what kind of visual language to aim for.
- Photorealistic: realistic textures, lighting, and physical detail
- Cinematic: film-inspired framing and mood
- Editorial: polished magazine aesthetic
- Commercial: clean, market-ready, brand-friendly image style
- Documentary: natural, observational, less stylized
- Concept art: imaginative world-building and design focus
- Anime: stylized linework, cel shading, expressive design
- Painterly: brush-like texture and expressive strokes
- Minimalist poster: reduced forms, graphic clarity, strong layout
- 3D render: polished computer-generated object or environment
Detail and texture terms
- Natural skin texture
- Detailed fabric folds
- Matte finish
- Glossy reflections
- Fine grain
- Crisp edges
- Subtle texture
- Weathered surface
- Clean background
- Sharp focus on subject
These smaller phrases often improve commercial utility because they steer the output away from muddy surfaces and vague materials.
How to customize
The easiest way to use this cheat sheet is to start with one intent, then add only the tokens that directly support that intent. Most prompt problems come from overloading the instruction with competing styles.
1. Start with the use case, not the model
Ask what the image needs to do. A YouTube thumbnail, e-commerce product shot, blog hero image, and fantasy wallpaper all need different composition and detail choices. For example:
- Thumbnail: clear focal point, bold contrast, readable negative space
- Product image: clean background, controlled reflections, centered composition
- Portrait: lens look, skin texture, lighting pattern, expression
- Poster: graphic hierarchy, strong silhouette, stylized finish
2. Choose one dominant visual priority
If you want realism, prioritize lens, light, and material detail. If you want illustration, prioritize style, line quality, and color language. If you want cinematic results, prioritize shot type, atmosphere, and color grading.
3. Add constraints deliberately
Constraints matter as much as descriptive phrases. If the model tends to add clutter, say minimal composition and clean background. If hands or faces are unstable, simplify the scene before adding advanced styling. You can also use negative prompts for AI art where supported to exclude recurring errors.
4. Save successful token clusters
Once you find a combination that works, save it as a reusable module. Examples:
- Portrait cluster: editorial portrait, 85mm lens look, soft window light, shallow depth of field, natural skin texture
- Product cluster: close-up product shot, diffused studio light, centered composition, glossy reflections, clean background
- Cinematic cluster: wide shot, backlit atmosphere, volumetric light, moody color grading, layered depth
This is one of the simplest ways to reduce iteration time in an AI art workflow.
5. Adjust by model behavior
Some models respond better to natural language. Others respond better to compressed phrase stacks. Some heavily stylize; others follow plain instructions more literally. If you are comparing tools, keep the subject constant and change only one variable at a time. That approach makes AI image generator comparison more practical and less subjective.
6. Avoid token conflicts
Try not to ask for all of the following at once: photorealistic, anime, documentary, surreal, studio strobe, golden hour, minimal, crowded, shallow depth of field, and deep focus. Mixed instructions can still work, but only when the tension is intentional. In most cases, fewer aligned terms produce stronger outputs than many disconnected ones.
Examples
Below are prompt examples you can adapt for common creator use cases.
Example 1: Photorealistic portrait
Editorial portrait of a young chef in a modern kitchen, medium close-up, eye-level shot, 50mm lens look, soft window light from the side, shallow depth of field, natural skin texture, muted earth tones, clean background, candid expression
Why it works: It combines subject, environment, framing, lens behavior, lighting, texture, and mood without adding unrelated style terms.
Example 2: Product marketing image
Premium coffee bag on a matte stone surface, close-up product shot, centered composition, diffused studio lighting, soft shadow under product, crisp label area, warm neutral palette, minimal commercial background, high detail packaging texture
Why it works: It gives the model a clear commercial objective and leaves room for clean composition.
Example 3: Cinematic travel scene
Lone traveler standing on a rain-soaked street in Tokyo at night, wide shot, low angle, neon lighting, backlit mist, reflections on pavement, layered depth, cinematic mood, cool magenta and blue palette, dynamic composition
Why it works: The prompt uses atmosphere and lighting to define mood, while the camera cues shape the scene’s dramatic perspective.
Example 4: Food overhead image
Top-down breakfast table with coffee, croissant, orange slices, linen napkin, overhead shot, soft morning window light, balanced composition, subtle texture, natural editorial food photography style, warm highlights, clean arrangement
Why it works: “Overhead shot” and “clean arrangement” are doing a lot of work here. The prompt tells the model exactly how the scene should be organized.
Example 5: Poster-style illustration
Futuristic bicycle poster, side profile view, bold silhouette, symmetrical composition, minimalist poster design, limited color palette, crisp vector-like edges, strong contrast, lots of negative space for headline
Why it works: It is specific about graphic structure rather than pretending a poster is the same as a photo.
Quick swap table
Use these substitutions when a prompt feels close but not quite right:
- If the image feels flat, add: directional light, rim light, layered depth
- If it feels too busy, add: minimal composition, clean background, negative space
- If it feels too synthetic, add: natural skin texture, subtle imperfections, documentary feel
- If it feels too plain, add: cinematic color grading, atmospheric haze, dramatic angle
- If the subject gets lost in the scene, add: close-up, centered composition, shallow depth of field
When to update
This cheat sheet is worth revisiting because prompt language is not static. Models change, interfaces change, and the phrases that once worked reliably may become less necessary or behave differently over time.
Update your personal version of this guide when:
- You switch tools or models. The same prompt may need simpler wording, stronger constraints, or a different ordering of terms.
- Your publishing workflow changes. A shift from blog graphics to thumbnails, ads, or product pages changes what prompt tokens matter most.
- You notice repeated failure patterns. For example, if backgrounds become cluttered or products lose label clarity, add or revise your constraint language.
- You build reusable assets. Once you create a prompt library for portraits, products, or branded scenes, revisit it quarterly and remove tokens that no longer help.
- Commercial standards rise. If you need cleaner results for client-facing or public work, refine texture, lighting, and composition terms rather than simply asking for “more realistic.”
A practical maintenance routine is simple:
- Create five to ten prompt modules you use often.
- Save one best-performing version for each use case.
- Note which tokens had the strongest visual effect.
- Remove decorative words that do not consistently change the output.
- Keep a companion list of negative prompts and failure fixes.
That habit turns a one-off prompt into a durable prompt library. It also makes your workflow easier to scale across collaborators, campaigns, or automated systems.
If you want to go further, build your own internal prompt glossary with sections for portrait, product, environment, poster, and thumbnail generation. Link each cluster to sample outputs and note where each phrase works best. This is one of the most practical ways to improve how to write better prompts over time without starting from zero on every generation.
In short, the best cheat sheet is not the longest one. It is the one you can return to, adapt quickly, and trust under real production pressure. Use the terms in this article as a working vocabulary, save the combinations that consistently perform well, and refine them as text-to-image models continue to evolve.