How to write prompts for AI video, image and audio

A reference for prompting generated video, image and audio, measured from 8,835 real prompts. Which aspects to specify, the vocabulary that works, drop in fragments and what to rule out.

Prompt cheatsheet

Every aspect, value and negative on two pages. Keep it open while you write.

Download the PDF

How to use this

  1. Work down the aspects in order. The order is the recommended order in the prompt too.
  2. One clause per aspect. A prompt is a list of decisions, not a paragraph of atmosphere.
  3. Prefer a number over an adjective wherever a number exists: 116 BPM over upbeat, 85mm over flattering, 15 seconds over short.
  4. Put the negatives last, in one sentence beginning with No.
  5. Reuse the identical phrasing across a batch. Paraphrasing is what makes a set of images look unrelated.

Video prompts

Measured across 4,300 video prompts, mean length 1537 characters.

Download the video sheet as a PDF, one page, or take all three.

Duration25.0% of video prompts

How long the clip runs.

Models hold a shot for a fixed span. Naming the length stops a model compressing a three beat idea into one, or padding a single beat with drift.

Values

6 seconds, 8 seconds, 10 seconds, 15 seconds, 30 seconds, 60 seconds

Also used in production skills, not counted

4-15s, 2-15s, 3-15s, 4, 6, 8, 10, 12, 15, 5, 4 s, ~15s, one to ten whole minutes

Fragments to drop in

  • 15-second cinematic clip.
  • Six seconds, one continuous take, no cuts.
  • 30 seconds in three beats of roughly ten seconds each.

From production skills

  • --duration 12
  • --duration 5
  • --duration 10
  • --duration 4

Model-bound: Seedance 2.0 4-15s (12s is valid); Kling 3.0 and Kling 3.0 Turbo 3-15s; Grok Video 1.5 2-15s; Veo 3.1 accepts only 4, 6 or 8; Marketing Studio integer ≥ 4; explainer blocks are fixed 10-second units with N = duration_minutes × 6; sprite animation is fixed at 4 s; cover reveal ~5s; scroll-scrub single-shot asks for the longest single take the model supports (~15s).

Aspect ratio31.3% of video prompts

Frame shape.

Decides whether the subject can breathe. A vertical crop of a wide composition loses the sides, so the ratio has to be chosen before the composition is described.

Values

16:9, 9:16, 1:1

Also used in production skills, not counted

4:3, 3:4, 21:9, auto, 3:2, 2:3, 9:21, 4:5, 5:4, 1.91:1

Fragments to drop in

  • Framed 16:9 for landscape playback.
  • Vertical 9:16, composed for a phone held upright.
  • Square 1:1 with the subject centred.

From production skills

  • --aspect_ratio 16:9
  • --aspect_ratio 9:16
  • --aspect_ratio auto

Meanings given in prompt-engineering.md: 16:9 landscape, cinematic; 9:16 vertical, social; 1:1 square, profile / icon. Veo 3.1 accepts only 16:9 or 9:16. For sprite work the ratio must be read off the key-pose file and passed explicitly on every call , omitting it falls back to the tool default and crops the subject.

Shot size25.1% of video prompts

How much of the subject fills the frame.

The most reliable lever there is. Naming it removes the model's habit of defaulting to a mid shot for everything.

Values

close-up, wide shot, medium shot, extreme close-up, establishing shot, two-shot, macro shot, over-the-shoulder, closeup

Also used in production skills, not counted

one hero subject, kept centered, clean negative space, center-safe area, cover, full body in frame, empty margin above the head and below the feet

Fragments to drop in

  • Extreme close-up on the texture, filling the frame.
  • Wide establishing shot, subject small in a large space.
  • Medium shot from the waist up, room to move on both sides.

From production skills

  • One hero subject, kept centered, with clean negative space around it , that space is where the chapter copy sits.
  • Full body in frame with empty margin above the head and below the feet.

The video fills the viewport and crops the edges, so essential subjects must stay away from the far left/right edges.

Camera angle6.8% of video prompts

Where the camera sits relative to the subject.

Changes who has power in the frame. Low angle makes a product monumental, overhead flattens it into a diagram.

Values

overhead, low-angle, eye level, three-quarter, top-down, profile, high-angle

Fragments to drop in

  • Low-angle looking up, subject towering over the lens.
  • Directly overhead, top-down on the surface.
  • Eye-level, camera at the subject's own height.

Camera movement19.1% of video prompts

How the camera travels during the shot.

The single most common cause of a shot feeling wrong. If you do not name it the model invents drift, and drift reads as an accident.

Values

push-in, tracking shot, handheld, orbit, locked-off, slow zoom, truck, slow dolly, parallax, pull-back, slow pan, whip pan, crane up, dolly in, steadicam, tilt down, dolly out, pan left, tilt up, dolly forward, dolly back

Also used in production skills, not counted

zooms in, dollies left, sweeping pan, slow push, fast whip, camera slowly pulls back, camera dollies in, slow orbit, rise, fly-through, lateral track, crane, detail push, rack focus soft to sharp, light sweep, Very slow push-in, Subtle camera push-in, slow push-in, drift, scale shock, hard contrast cut, slow forward drift, gentle forward velocity, pull out, descending into, camera locked, no camera movement, no zoom

Fragments to drop in

  • Slow push-in toward the subject, ending tight.
  • Locked-off tripod, no camera movement at all.
  • Smooth orbit around the subject, one quarter turn.
  • Handheld with slight sway, documentary feel.

From production skills

  • camera dollies in
  • subtle product reveal, camera slowly pulls back, ambient motion
  • MOTION: Very slow push-in; the figure erodes grain by grain and particles drift sideways.
  • Subtle camera push-in, no flicker, no extra text, no new objects.
  • no cuts, no camera shake, slow steady motion only, locked exposure, no on-screen text
  • camera locked, no camera movement, no zoom, subject stays fully in frame, plain static background

For image-to-video the movement IS the prompt; for scrubbable footage a single unbroken move is mandatory and cuts are a defect. For sprite sheets the camera must be explicitly locked or frames become unusable.

Lens and focal length3.8% of video prompts

Perspective and compression.

A focal length is a shorthand a model understands: 85mm flatters a face, 24mm bends a room outward, macro turns a surface into a landscape.

Values

35mm, macro lens, anamorphic, 100mm, 50mm, 85mm, 28mm, wide-angle lens, 32mm, 50 mm, fisheye, telephoto, 40mm, 35 mm, 85 mm, 24mm

Fragments to drop in

  • Shot on an 85mm lens, compressed and flattering.
  • Wide 24mm, close to the subject, edges stretching.
  • Macro lens, the subject filling the frame at life size.

Depth of field11.8% of video prompts

What is sharp and what falls away.

Separates subject from background without changing the composition. The corpus reaches for shallow depth of field more than any other optical term.

Values

shallow depth of field, bokeh, rack focus, f/1.4, f/1.8, f/2.0, soft focus, f/2.8

Fragments to drop in

  • Shallow depth of field, background dissolving into bokeh.
  • Deep focus, everything from foreground to horizon sharp.
  • Rack focus from the foreground object to the face behind it.

Lighting7.4% of video prompts

Where the light comes from and how hard it is.

Sets mood more cheaply than any other choice. Naming a setup also stops the flat, evenly lit look models fall back on.

Values

soft lighting, rim lighting, rim light, silhouette, volumetric lighting, golden hour, volumetric light, soft light, side lighting, chiaroscuro, key light, low-key, practical lighting, backlit, high-key, side light, practical lights, backlight, god rays, neon lighting, blue hour, softbox

Fragments to drop in

  • Soft lighting from a large source, shadows barely there.
  • Rim lighting from behind, edge of the subject glowing.
  • Single hard key from the side, deep shadow on the far cheek.
  • Golden hour sun, low and warm, long shadows.

Colour and grade17.0% of video prompts

The palette and contrast of the finished image.

Carries brand more than any other single element, and is easy to get consistent across a set of clips by reusing the same phrase.

Values

vibrant, pastel, film grain, black-and-white, natural color, high contrast, desaturated, crushed blacks, natural colour, low contrast

Also used in production skills, not counted

name the grade and the hexes, one visual grade, colour grading

Fragments to drop in

  • Vibrant saturated grade, colours pushed but not clipping.
  • Muted pastel palette, low contrast, gentle.
  • High contrast with crushed blacks and a cool cast.
  • Fine 35mm film grain over the whole frame.

From production skills

  • Prompt "no cuts, no camera shake, slow steady motion only" and name the grade/hexes.
  • Take only the visual render style and color grading of the input image(s); mix the styles if there is more than one image.

scroll-scrub.md forbids mixing models mid-chain because "their grain, color, and motion signatures create a visible seam even when position matches".

Medium and style38.1% of video prompts

What the output is pretending to be.

The most used word in the whole corpus is cinematic, which tells you models respond to a named medium. Being specific beats being evocative.

Values

cinematic, photorealistic, anime, documentary, 3d render, unreal engine, stop-motion, 35mm film, watercolor, vector, photoreal, claymation, line art, octane

Also used in production skills, not counted

flat 2D vector animation, bold clean outlines, solid vibrant flat fills, no shading, no gradients, hand-inked black marker on off-white paper, solid jet-black fills, thin white scratch highlights, marker grain, strictly monochrome, strict monochrome minimalism, black silhouettes on white void, high contrast, lots of negative space, hand-painted storybook gouache, soft textures, warm muted palette, visible brush strokes, non-photorealistic, illustrated, not a photo, no live-action, no realism

Fragments to drop in

  • Cinematic, shot on 35mm film.
  • Photorealistic, indistinguishable from a real camera.
  • 3D render, clean studio lighting, product visualisation.
  • Hand drawn 2D animation, visible line work.

From production skills

  • STYLE REFERENCE: Match the attached reference image EXACTLY. Replicate its look precisely: {STYLE tokens}. Every element below rendered in that identical style.
  • non-photorealistic, illustrated, not a photo, no live-action, no realism
  • Paste the same STYLE tokens into every block.

Write the render style, palette, line character and finish once, then repeat it byte-identically. scroll-scrub.md calls the equivalent a 'World grammar': one byte-identical style preamble, perspective, palette, light direction, surface finish and background behavior across all scene prompts.

Subject motion10.2% of video prompts

What moves inside the frame, and how fast.

Distinct from camera movement and frequently confused with it. Say what the subject does, or the model will hold it still and move the camera instead.

Values

reveal, float, drift, slow motion, glides, rotates, time-lapse, ripples, unfolds, hovers, cascade

Also used in production skills, not counted

the dancer spins, smoke rises slowly, ambient motion, fur ripples, stars orbit, clouds drift, train lights flicker, object assembling, exploded-view assembly, transformation or morph, a hero object emerging from darkness, smooth idle breathing cycle, subtle weight shift, full walk cycle in place, single sword slash, fast wind-up, sharp strike, follow-through, plain static background, studio black/charcoal, a soft gradient, the subject emerging from darkness, dark, seamless, low-detail, clean background, simple background

Fragments to drop in

  • The lid glides open on hidden runners, slow and even.
  • Fabric billows in slow motion, every fold readable.
  • Steam drifts upward and dissipates before it leaves the frame.

From production skills

  • SCENE: A lone black silhouette slowly dissolves at the edges, crumbling into fine drifting sand that scatters into the white emptiness.
  • The character performs ONLY this action; nothing else happens.
  • Use one clear action per block.
  • A background the copy can survive. Dark, seamless, low-detail (studio black/charcoal, a soft gradient, the subject emerging from darkness) is the reliable choice.

"Exactly one action per video; any second verb in the prompt is a defect" (game-2d-animation.md).

Critical when HTML copy sits over the video: "A bright, busy, full-frame environment behind body copy is the single most common reason a beautiful film reads as unusable."

Mood21.9% of video prompts

The feeling the clip should leave.

A single adjective steers dozens of small decisions the model would otherwise make at random. Cheap to add, and it makes a set of clips feel related.

Values

playful, energetic, whimsical, calm, dramatic, minimal, elegant, uplifting, moody, ominous, ethereal, luxurious, serene, nostalgic, cosy, gritty

Fragments to drop in

  • Calm and unhurried throughout.
  • Playful and energetic, quick and bright.
  • Moody and restrained, close to darkness.

Audio bed11.8% of video prompts

The sound that plays under the picture.

Video prompts in this corpus routinely describe their own soundtrack. Naming instruments and what must not be heard is more reliable than naming a genre.

Values

percussion, bells, ukulele, xylophone, piano, drums, bass, marimba, strings, synth, guitar, choir, handclaps, pads, flute, synthesizer, synthetic, harp, synthase, kalimba

Also used in production skills, not counted

ambient SFX or music only, no voice, dialogue, or narration, Low sustained drone, a soft whisper of falling sand, generate_audio false, sound on/off

Fragments to drop in

  • Minimal deep electronic ambience with soft mechanical clicks.
  • Warm ukulele and marimba, light percussion, no vocals.
  • No music at all, only the room and the mechanism.

From production skills

  • AUDIO: {ambient SFX or music only, NO voice, dialogue, or narration}.
  • AUDIO: Low sustained drone and a soft whisper of falling sand, no voice.
  • Keep clip audio diegetic only. Characters never speak or lip-sync.

In the explainer pipeline narration is a separate audio job; clip audio must never contain speech.

Pacing and cuts

How the time is divided.

Without it a model gives you one unbroken drift. Naming beats is what turns a clip into a sequence.

Also used in production skills, not counted

smooth cinematic easing, slow, steady motion, constant speed, gentle ease only at the very start and end, locked exposure and white balance, no flicker, minimal motion blur, one continuous move, no hard cuts, START state ≠ END state, 0-1.5s , scene alive, no text yet, 1.5-3.5s , text entrance, 3.5-5s , settle, pops/slides/bounces in, pixel type snaps in block by block, chrome bubble letters inflate, condensed uppercase slams down, fades up, slides in with a soft bounce, Micro-motion only

Fragments to drop in

  • One continuous take, no cuts.
  • Three beats: reveal, detail, hero shot.
  • Cut away before the hand touches the product.

From production skills

  • One continuous move , no hard cuts.
  • Slow, steady motion , constant speed, gentle ease only at the very start and end. The scroll supplies the pacing.
  • Locked exposure and white balance, no flicker; minimal motion blur.
  • Resolves at both ends: the first frame reads as the establishing shot, the last as the closing beauty state. START state ≠ END state, or the scrub has no payoff.
  • Premium 5-second motion cover reveal, smooth cinematic easing.
  • All elements ease precisely into their final positions and the video ends exactly on the provided end frame, holding still for the last moments.
  • Then the title "[TITLE TEXT]" [ENTRANCE MOTION matched to its typography], the tagline fades up beneath it, and the small rounded pill button slides in with a soft bounce.

Matters whenever every frame may be held as a still (scroll-scrub) , exposure pumping shimmers and heavy blur smears.

Entrance motion should be matched to the typography of the text being revealed.

Resolution / quality tierfrom production skills

Seedance 2.5 caps at 720p so 1080p/4K work stays on Seedance 2.0; Grok Video 1.5 only 480p or 720p; Marketing Studio video 480p or 720p.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

480p, 720p, 1080p, 4k, basic, high, ultra, pro, std, standard, bitrate_mode standard|high

Fragments to drop in

  • --resolution 4k
  • --resolution 720p
  • --resolution 1080p --mode std

First/last frame anchoringfrom production skills

scroll-scrub.md warns that passing a storyboard as a start frame makes the page open on a static image; pass it in the generic reference role instead.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

--start-image, --end-image, start frame, end frame, start = end, one-shot, loop, end-frame reveal, generic image/reference role , not a start-frame role

Fragments to drop in

  • `--start-image` anchors the first frame. Prompt describes motion.
  • Don't redescribe the static frame , model already has it.
  • Looping actions pass the SAME image as both start frame and end frame. One-shot actions (attack, death, hit, cast) pass only the start frame , forcing them back to the start pose ruins the action.
  • Pass the finished cover as the END frame and describe the buildup

Ad format / mode (branded video)from production skills

Hook text is prepended to the user's prompt and does not replace it; COOKBOOK.md notes "Hooks (the prompt on --prompt) matter more than mode for performance."

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

ugc, ugc_how_to, ugc_unboxing, product_showcase, product_review, tv_spot, wild_card, ugc_virtual_try_on, virtual_try_on, hook, setting, ad reference

Fragments to drop in

  • Looks like a real person filmed on phone
  • Polished broadcast commercial
  • Show the product itself, less presenter
  • Presenter giving an opinion

Edit instruction (video editing workflows)from production skills

draw_to_video takes a source video, an edited/sketched frame, the timestamp for that frame and a short edit instruction.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

make the jacket red, sketch, timestamp

Fragments to drop in

  • --prompt "make the jacket red"

The order to write them in

A prompt is read start to finish, and earlier clauses set the frame the later ones are interpreted against. This order puts the decisions that cannot be undone first.

  1. Duration
  2. Aspect ratio
  3. Medium and style
  4. Shot size
  5. Camera angle
  6. Camera movement
  7. Lens and focal length
  8. Depth of field
  9. Lighting
  10. Colour and grade
  11. Subject motion
  12. Mood
  13. Audio bed
  14. Pacing and cuts
  15. Negatives

How production prompts are laid out

Two dominant shapes. (1) The explainer block, a labelled five-line form in fixed order: `Block N` / `STYLE REFERENCE:` / `SCENE:` / `MOTION:` / `AUDIO:` / `NEGATIVE:` , the STYLE REFERENCE and NEGATIVE lines are pasted byte-identically into every block, only SCENE and MOTION change, and each block holds exactly one clear action (higgsfield-video-explainer/references/prompts.md). (2) The image-to-video single paragraph, where the start image already carries the subject so the prompt is motion only , "Don't redescribe the static frame , model already has it" , optionally followed by a mandatory negative block in three fixed parts: camera lock, prop and action inertia, facing lock (higgsfield-generate/references/prompt-engineering.md, higgsfield-websites/references/game-2d-animation.md). Longer reveal prompts run beat by beat in time order: start state and context motion to text/element entrance to settle, closing with an anchor sentence that ties the last frame to the supplied end frame (higgsfield-websites/references/cover-animator.md). Footage-level constraints (one continuous move, centre-safe hero, survivable background, locked exposure, resolving ends) are restated in the prompt rather than left to the model (higgsfield-websites/references/scroll-scrub.md).

What to rule out

The most common element in the whole corpus. Put these last, in one sentence, and only include what is plausible for the shot: a long list of irrelevant prohibitions wastes the model's attention.

Text and branding

text, logos, subtitles, watermark, watermarks, captions, on-screen text, text on screen

Anatomy and faces

extra limbs, distorted faces, extra characters, hands, extra fingers, extra animals

Content and safety

violence, scary imagery, scary elements, weapons

Audio

music, narration

Other

dialogue, copyrighted characters, talking, cuts, horror, humans, clothes, flickering, frightening imagery, morphing, camera shake, zoom

Fragment to drop in

  • No text, logos, subtitles, watermark, watermarks, captions, extra limbs, distorted faces, extra characters, hands, extra fingers, extra animals, violence, scary imagery, scary elements, weapons, music, narration, dialogue, copyrighted characters, talking, cuts, horror, humans.

Exclusion lines used in production

Copied verbatim from working prompt skills. Useful as whole lines rather than as a vocabulary.

  • color drift, photorealism, 3D render, lip-sync, captions, on-screen text, logos, watermark
  • color, gray midtones, photorealism, 3D render, lip-sync, captions, on-screen text, logos, watermark
  • non-photorealistic, illustrated, not a photo, no live-action, no realism
  • no voice, dialogue, or narration
  • camera locked, no camera movement, no zoom, subject stays fully in frame, plain static background
  • the character performs ONLY this action, nothing else happens
  • the <prop> stays inert and is never used , no firing, no muzzle flash, no swinging, no raising it
  • the subject keeps facing the SAME direction for the entire video , never turns around, never rotates toward or away from the camera, no head turns past the shoulder
  • no cuts, no camera shake, slow steady motion only, locked exposure, no on-screen text
  • Subtle camera push-in, no flicker, no extra text, no new objects.
  • No on-screen text, logos, or watermarks , all type is HTML over the video.
  • Avoid: cuts, fast pans, handheld shake, busy bright backgrounds, and colour/exposure flicker.
  • The solid color frame, corner dots and capsule shape stay perfectly static at all times.
  • no lip-sync, captions, text, logos, or watermark

Image prompts

Measured across 3,597 image prompts, mean length 1063 characters.

Download the image sheet as a PDF, one page, or take all three.

Aspect ratio16.5% of image prompts

Frame shape.

Choose it first. Composition instructions that assume a landscape frame fall apart in a vertical one.

Values

16:9, 9:16, 1:1, 4:5, 21:9, 3:2, 4:3

Fragments to drop in

  • Square 1:1.
  • Vertical 4:5 for a feed post.
  • Wide 16:9 with the subject offset left.

Composition and framing31.3% of image prompts

Where things sit in the frame.

The most common failure in the corpus is a subject cropped or crowded. Explicit composition instructions, including what must stay fully visible, fix it.

Values

close-up, wide shot, medium shot, extreme close-up, two-shot, closeup, over-the-shoulder, establishing shot, macro shot, full shot

Also used in production skills, not counted

one power-third hero, clear scale hierarchy, depth, strong subject/background separation, centered, left-anchored, asymmetric split, one focal visual, decisive negative space, generous negative space, reserve the block side, soft open negative space upper-left, single unified frame, split-screen, diagonal divide, halves/panels, extreme scale contrast, power third, centered isolated presentation

Fragments to drop in

  • Centred on a single vertical axis, symmetrical.
  • Rule of thirds, subject on the right third, negative space left.
  • Everything fully inside the frame, nothing cropped at the edges.
  • Generous negative space above the subject for a headline.

From production skills

  • Composition: one power-third hero, clear scale hierarchy, depth, and strong subject/background separation.
  • Scale + placement: "waist-up, subject fills two thirds of the frame, positioned right-of-center" , and reserve the block side: "soft open negative space upper-left on a warm white wall". Negative space = a real surface (wall, sky, backdrop), NOT a black void.
  • single unified frame , no split-screen, no diagonal divide, everything blends smoothly and organically across the same continuous shot

A split layout is used only when the user asks for split, before/after, versus or side by side; "A topical phrase such as `X vs Y` does not itself require a split."

Shot size11.4% of image prompts

How close the camera is.

Same lever as in video and just as reliable. It also controls how much detail the model has to invent.

Values

close-up, wide shot, medium shot, extreme close-up, two-shot, closeup, over-the-shoulder, establishing shot, macro shot, full shot

Also used in production skills, not counted

chest-up, medium-close, about 40-60% of the frame, waist-up, subject fills two thirds of the frame, 55-75 % of the frame, Full body, Three-quarter, Waist up, Closeup on product area, head shot, head-and-shoulders, full body, large, foreground-dominant, rendered LARGE

Fragments to drop in

  • Extreme close-up, texture filling the frame.
  • Full body shot, head to feet, room above and below.
  • Head and shoulders portrait.

From production skills

  • Subjects: large, foreground-dominant, chest-up or medium-close, filling about 40-60% of the frame. End with `All faces crisply sharp as the anchors of the shot.`
  • waist-up, subject fills two thirds of the frame, positioned right-of-center

app-cover.md kills a candidate whose subject is under ~50 % of frame. Framing is an explicit interview option in product-photoshoot: `[Full body / Three-quarter / Waist up / Closeup on product area]`.

Camera angle4.6% of image prompts

Viewing position.

Three-quarter is the most used angle in the corpus for a reason: it shows form without the stiffness of a straight-on view.

Values

three-quarter, profile, eye-level, low angle, overhead, high angle, top-down, bird's-eye, dutch angle

Also used in production skills, not counted

low, low wide angle, side-view, flat frontal, three-quarter isometric view, front, 3/4 left, 3/4 right, slight up/down, wide establishing view, multi_angle

Fragments to drop in

  • Three-quarter view, turned slightly off axis.
  • Straight-on profile, subject facing frame left.
  • Shot from directly above, flat lay.

From production skills

  • angle (low, overhead)
  • consistent side-view perspective across all assets
  • three-quarter isometric view

game-stylization.md makes the perspective word genre-bound: platformer/runner to side-view; top-down to top-down; puzzle/clicker to flat frontal. For image-to-3D concept images the perspective word is replaced by three-quarter isometric view because it shows top plus two sides.

Lens and focal length5.0% of image prompts

Perspective and compression.

Also the cheapest way to signal photographic realism, because a real photograph was taken at some focal length.

Values

35mm, 85mm, 50mm, 110 mm, 16mm, anamorphic, 24mm, 70mm, 48mm, fisheye, 600mm, 450mm, 40mm, 297 mm, 650mm, 20mm, 135mm, 100mm, wide-angle lens, 105 mm, 220mm, 110mm

Also used in production skills, not counted

low wide angle, close microphone (audio analogue), 35mm film

Fragments to drop in

  • 85mm portrait lens, natural compression.
  • 35mm, mild environmental context around the subject.
  • Macro at life size, the surface reading as landscape.

From production skills

  • **Camera**: lens (35mm, 85mm), angle (low, overhead), motion (dolly in, tracking shot)
  • bright soft key with punchy shadows, low wide angle, commercial editorial grade

app-cover.md notes the lens block is replaced by style words such as "glossy 3D render / claymation diorama / painterly still" on the graphic route. Vector/logo prompts must never contain lens language.

Depth of field6.0% of image prompts

Focus falloff.

Separates a subject from clutter you did not describe, which is most backgrounds.

Values

shallow depth of field, bokeh, f/1.4, f/7.1, soft focus, f/2.8, deep focus, f/2.0, f/4, f/2, f/1.8, f/5.6

Fragments to drop in

  • Shallow depth of field at f/1.8, background soft.
  • Everything sharp front to back, deep focus.
  • Only the eyes critically sharp.

Lighting6.8% of image prompts

Direction, quality and source of light.

Silhouette and rim light are the two most used named setups in the corpus, both of which the model will never choose on its own.

Values

silhouette, rim light, soft light, soft lighting, backlit, key light, golden hour, high-key, rim lighting, backlight, softbox, low-key, volumetric light, bounce light, practical lighting, three-point lighting, fill light, volumetric lighting, practical lights, chiaroscuro, side lighting, god rays

Also used in production skills, not counted

neon glow, moody backlight, strong key light, soft dreamy fill, back light, hair light, bright soft key with punchy shadows, flat ambient light, flat even lighting, soft studio reflections, soft contact shadow, warm hard light, restrained lighting, no shadows, no vignette, soft vignette and edge falloff, warm amber torchlight

Fragments to drop in

  • Large softbox front left, soft shadows falling right.
  • Backlit into silhouette, subject dark against a bright field.
  • High-key, white on white, almost no shadow.
  • Single hard light, sharp shadow edges.

From production skills

  • **Lighting**: rim light, neon glow, moody backlight
  • signature YouTube thumbnail lighting rig , strong key light sculpting the face, soft dreamy fill lifting shadows, and defined back light plus hair light tracing a clean bright rim around hair, shoulders and silhouette.
  • bright soft key with punchy shadows, low wide angle, commercial editorial grade
  • Flat even lighting, no shadows, no vignette.
  • flat even lighting, no perspective, no objects, photographed-flat appearance

In the thumbnail rig "Only the rim may use a colored accent." Tileable textures need flat lighting so the tile has no focal light. Brandkit mockups explicitly ban "Excessive bloom, haze, depth of field, or cinematic lighting".

Colour and palette10.4% of image prompts

Palette, saturation, contrast.

Keeps a set consistent. Reuse the identical phrase across a batch rather than paraphrasing it each time.

Values

vibrant, pastel, film grain, black-and-white, desaturated, high contrast, natural color, natural colour, monochrome, low contrast, sepia

Also used in production skills, not counted

vivid, bright, glossy, poster-punchy, deep blacks, crisp highlights, rich saturated colors, cohesive as one image, commercial editorial grade, duotone, graded atmospheric photo, monochrome / force grayscale, muted cinematic movie still (banned), name 2-3 colors, exact hex values, locked primary tone, locked accent, locked background, background, text, primary, accent, support, environment in deep violet stone with charcoal shadows, hero in warm coral-cream tones contrasting the surroundings, hazards and pickups marked with acid-green glow, deep-navy #243A5E, signal-red #D23B2E, olive #7A7D3C, sky-blue #A9CFF4, peach #F7DDB9, cream #F1EEE6, dusty-rose #E7B7B0, warm-gray #D6D3CE, acid lime #D4FF3F, #FF00FF, #00FF00, #0000FF

Fragments to drop in

  • Vibrant, fully saturated, clean whites.
  • Muted pastel palette, soft contrast.
  • Black and white, deep blacks, bright highlights.

From production skills

  • Grade: vivid, bright, glossy, poster-punchy, deep blacks, crisp highlights, rich saturated colors, cohesive as one image. Restrain it only for an explicit calm/premium/muted brief.
  • for a monochrome look force grayscale in the prompt AND on export
  • Palette, committed and saturated: name 2-3 colors. Avoid lime/acid yellow-green.
  • Use role language such as "locked primary tone," "locked accent," and "locked background."
  • Exact hex values are passed separately through the Recraft request's `colors` and `background_color` params. Never put hex, RGB, Pantone, or other color codes in the prompt.

asset-system.md: a mixed-grade asset kit reads cheaper than no assets; re-generate off-grade pieces "with the hexes named harder".

Two opposite conventions co-exist: website/game prompts NAME the hexes in the prompt; Recraft logo prompts must never contain a colour code and pass hexes as params. Logo palettes use one, two or three colours , "three is a maximum, not a target or default". Game palettes are specified BY ROLE, not as one global gamma.

Medium and style26.6% of image prompts

Photograph, render, illustration or something else.

Ambiguity here produces the uncanny middle ground that looks like neither a photo nor an illustration.

Values

cinematic, photorealistic, documentary, anime, 3d render, line art, watercolor, vector, 16mm, 35mm film, photoreal, kodak vision3, isometric, flat illustration, claymation

Also used in production skills, not counted

oil painting, photograph, soft cel shading, glossy 3D render, claymation diorama, painterly still, flat vector cartoon with soft gradients, chunky pixel art, 32x32 grid feel, soft hand-painted gouache, low-poly faceted, smooth rounded blobby, flat_vector, monoline, vector_gradient, hand_drawn_vector

Fragments to drop in

  • Photorealistic editorial photograph.
  • Clean 3D product render on a seamless backdrop.
  • Flat vector illustration, limited palette, no gradients.
  • Loose watercolour with visible paper texture.

From production skills

  • **Style/medium**: oil painting, watercolor, photograph, anime, 3D render
  • transform into anime style, vibrant colors, soft cel shading

Brandkit treatment vocabulary is a closed set with defined meanings: flat_vector = solid fills, clean SVG paths, no surface effects; monoline = uniform stroke weight, rounded caps, no fills; vector_gradient = vector-safe linear, radial or duotone gradient with the locked stop count; hand_drawn_vector = intentional stroke variation. "Dimensional/3D treatment is forbidden."

Materials and finish8.5% of image prompts

What surfaces are made of.

Materials are where a render either convinces or does not. Naming the finish is more effective than naming the object.

Values

glossy, matte, ceramic, concrete, leather, velvet, satin, brass, linen, frosted, marble, chrome, walnut, oak

Also used in production skills, not counted

print, emboss, debossing, foil, laser engraving, one-color printing, screen printing, embroidery, spot-color, folds, grain, perspective, occlusion, reflections, scale, manufacturing limits, believable materials, light kraft paper, natural cardboard, pale fabric, curved packaging, surface wear

Fragments to drop in

  • Brushed metal with a satin finish, no mirror reflections.
  • Matte ceramic, slightly uneven glaze.
  • Smoked glass over dark walnut.

From production skills

  • [PHYSICAL APPLICATION] Render the logo using <credible print/emboss/foil/engraving/ink behavior>. Respect folds, grain, perspective, occlusion, reflections, scale, and manufacturing limits.
  • Prefer believable materials, restrained lighting, purposeful negative space, specific environments, and one focal branded application.

Use the generator only when branding must interact with fabric folds, curved packaging, embossing/debossing, foil, print texture, reflections, surface wear, occlusion or realistic perspective; otherwise place the logo deterministically.

Background36.9% of image prompts

What sits behind the subject.

Named in 37% of image prompts, the highest of any positive aspect. An undescribed background is where models put clutter, text and logos.

Also used in production skills, not counted

vivid high-contrast color field, soft vignette and edge falloff, 2-4 supporting story props/layers, credible setting, specific environments, no readable screens, clean background, solid uniform bright <KEY COLOR> background, pure flat white, no shadow, no ground, Studio clean, Outdoor natural, Street style, Editorial, Home cozy

Fragments to drop in

  • Seamless mid-grey studio backdrop, no horizon line.
  • Plain white, no props, no shadow on the background.
  • Softly blurred interior, unreadable, no recognisable objects.

From production skills

  • Background: vivid high-contrast color field or environment, soft vignette and edge falloff; do not divide it unless a split layout was explicitly requested.
  • **World + props (story density)**: name 2-4 supporting elements that fill the frame around the subject. Add "no readable screens" whenever screens/props could sprout text.
  • on a solid uniform bright <KEY COLOR> background, no shadows cast on the background, no ground plane, nothing cropped at the edges

app-cover.md: "One person on an empty backdrop reads as AI slop even when bright , stage a world, not a portrait." For image-to-3D the background is instead "pure flat white, no shadow, no ground".

Mood9.9% of image prompts

The feeling.

One adjective, placed early, quietly resolves dozens of choices in a consistent direction.

Values

elegant, playful, minimal, dramatic, calm, energetic, whimsical, moody, luxurious, nostalgic, uplifting, ethereal, cosy, gritty

Also used in production skills, not counted

shock, hype, rage, awe, laugh, fear, smug, charisma, confusion, determination, disgust, neutral, smiling, talking

Fragments to drop in

  • Elegant and restrained.
  • Playful and bright.
  • Moody, close to darkness.

From production skills

  • If the user gives an emotion count without names, use this ladder: shock, hype, rage, awe, laugh, fear, smug, charisma, confusion, determination, disgust.

Emotion is one of the allowed surgical-edit scopes; expression changes are made as an edit, not a full regeneration.

Text handling18.4% of image prompts

Whether any lettering appears.

Models add text nobody asked for. 18% of image prompts in the corpus explicitly forbid it, and that is the single highest value instruction you can add.

Also used in production skills, not counted

No text, no readable UI labels, no watermark, BAKED UI, reading exactly, massive, ultra-legible sans-serif with a clean outline/glow treatment, 2-4 words max, ALL CAPS, no extra text

Fragments to drop in

  • No text, no lettering, no numbers anywhere in the image.
  • No logos, watermarks, captions, borders or frames.
  • Leave the upper third clear for a headline added later.

From production skills

  • Text: default `No text, no readable UI labels, no watermark.`
  • TEXT: bold thumbnail headline text baked into the image, reading exactly "<TEXT>" , massive, ultra-legible sans-serif with a clean outline/glow treatment, placed where it never covers the subject's face. No other text, no watermark.
  • **The image model NEVER renders text. Not the title, not the wordmark, not labels, not UI. It renders ONLY the scene.**
  • Never bake the final logo or exact copy into a generated image.

The dominant convention across the repo is: keep text out of the generation and set type deterministically afterwards. GPT Image 2 is the one model reached for when readable text must be inside the image.

Subject + setting + style (the opening clause)from production skills

"Higgsfield models reward concrete, sensory prompts." app-cover.md wants the product's core verb staged as a physical, photographable moment.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

a red fox curled in a snowy pine forest, golden hour, cinematic, subject + action, concrete and physical

Fragments to drop in

  • a red fox curled in a snowy pine forest, golden hour, cinematic
  • a film editor in a bold red bomber jacket mid-leap, cutting a giant arc of 35mm film with oversized chrome scissors

Typography named in the prompt (when text is generated)from production skills

text-overlay-bake.md gives the MrBeast headline stack as concrete values: font Anton, ALL CAPS, stroke 8-14% of cap size drawn under the fill, hard dark offset shadow, colour white / yellow to orange to red gradient / acid lime #D4FF3F, tracking -0.01...-0.02em, line-height 0.9, cap height 12-18% of frame height.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

Anton, Anton SC, Bebas Neue, Oswald, Archivo Black, Montserrat, Poppins, Roboto Condensed, Inter, Barlow Condensed, Playfair Display, Cormorant Garamond, DM Serif Display, Fraunces, Sacramento, patrick, caveat, marker, anton, clean grotesk (Satoshi-like), refined grotesk (Neue-Montreal-like), expressive display (Cabinet/Clash-like), compressed statement (Monument-like), editorial serif + sans pairing, Swiss rational sans with hard hierarchy

Fragments to drop in

  • The prompt must state: Exact literal copy; Display/body font family names and which text uses each
  • Name the exact display/body families in the prompt; never infer typography from the logo or palette.

Identity lock / subject fidelityfrom production skills

Up to three referenced identities; face photos are passed first in character order, then the logo, and a manifest line must open the prompt when two or more references are attached.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

IDENTITY LOCK, photographic identity match, same bone structure, eye shape, nose, lips, jawline, skin tone, hairline and hair texture, Do not beautify, average, or restyle the face, IMAGE REFERENCES manifest

Fragments to drop in

  • CHARACTER N: the person from attached face reference #K , IDENTITY LOCK: reproduce this exact person with a photographic identity match , same bone structure, eye shape, nose, lips, jawline, skin tone, hairline and hair texture. Do not beautify, average, or restyle the face. Expression: <emotion phrase>.
  • IMAGE REFERENCES: image 1 = CHARACTER 1 face reference; image 2 = brand logo.
  • All faces crisply sharp as the anchors of the shot.

Aesthetic / seasonal presetfrom production skills

These are the labelled options the skill offers for restyle and style/mood questions; they become the short intent line handed to the backend enhancer.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

Clean girl, Cottagecore, Quiet luxury, Dark academia, Y2K, Christmas, Valentine's, Halloween, Black Friday, None, Clean studio, Lifestyle, Conceptual, With a model

Fragments to drop in

  • Christmas version, quiet-luxury aesthetic
  • vertical pin for my candle brand, cottagecore mood

Product-photo genre / modefrom production skills

For both skills the user-facing prompt is a short intent line only , "Backend assembles the final prompt; never freehand it." `--count 3` makes the backend vary preset, lighting, angle and palette across variants.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

product_shot, lifestyle_scene, closeup_product_with_person, moodboard_pin, hero_banner, social_carousel, ad_creative_pack, virtual_model_tryout, conceptual_product, restyle, levitating, floating, splash, frozen motion, surreal, CGI, sculptural, main_image, infographic, multi_angle, detail_shot, lifestyle, whats_in_box

Fragments to drop in

  • bottle of cold-brew on a sunlit kitchen counter, IG feed
  • sparkling peach lemonade can for marketplace listing
  • premium skincare serum, clean clinical marketplace visual system
  • Bold hero shot on marble

Logo / vector mark specificationfrom production skills

One central visual idea per mark, stated in one clause; the exact phrase "no text" must appear in the constraint tail; mark_type routing is evidence-based, defaulting to abstract, never wordmark or combination mark.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

lettermark_monogram, pictorial, abstract, mascot, emblem, one dominant unified silhouette, negative-space device, motif treatment, unexpected locked color-role pairing, symmetry or intentional asymmetry, internal detail limit, small-size scalability, centered isolated presentation, minimal anchor points, SVG-friendly

Fragments to drop in

  • Flat vector design, clean lines, no shadows, no texture, no text. Clean editable vector paths, SVG-friendly, minimal anchor points.
  • Monoline vector design, uniform stroke weight, rounded line caps, no fills, no shadows, no texture, no text. Clean editable vector paths, SVG-friendly, minimal anchor points.
  • Flat vector design with the specified locked vector gradient, no shadows, no texture, no text. Clean editable vector paths, SVG-friendly, minimal anchor points.
  • Intentional hand-drawn vector strokes, approved surface variation, clear scalable silhouette, no shadows, no text.
  • Describe the distinctive element concretely enough that another designer could sketch its structure without guessing.

Logo placement on a mockup surfacefrom production skills

Full-colour marks go on smooth surfaces; black monochrome on light kraft paper, natural cardboard, pale fabric, stamps, dark-ink screen printing, engraving masks and light uncoated stock; white monochrome on dark paper/boxes/fabric, reverse marks, light-ink screen printing and dark signage.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

Target surface, Position, Scale, Orientation, Color, horizontally centered, upper third, optical center aligned to panel, clear-space margin, align to <panel edge/seam/baseline>, follow surface perspective

Fragments to drop in

  • [PLACEMENT LOCK] Target surface: <exact object panel/face/material>. Position: <exact alignment and location, e.g. horizontally centered, upper third, optical center aligned to panel>. Scale: logo occupies <specific proportion> of the target surface while keeping <specific clear-space margin>. Orientation: align to <panel edge/seam/baseline>; follow surface perspective without changing logo proportions. Color: use the supplied <full-color/black/white> variant exactly. State why it contrasts correctly with the material/background.
  • [AUTHORITATIVE LOGO] <ImageN> is the exact approved <full-color/black/white> logo. Preserve its spelling, silhouette, geometry, proportions, internal negative space, and exact color. Do not redraw, simplify, crop, stretch, outline, or add effects.
  • "Place the logo on the bag/box" is insufficient.

Seamless / tileable behaviourfrom production skills

{MATERIAL} in the texture templates is 3-6 concrete words from the reference, e.g. "hand-painted cracked dry earth with small grey stones".

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

seamless tileable game texture tile, uniform pattern density, perfectly seamless edges that wrap horizontally and vertically, no border, no vignette, no single focal object, photographed-flat appearance, seamless repeating pattern

Fragments to drop in

  • seamless tileable game texture tile of <description>, uniform pattern density,
  • , perfectly seamless edges that wrap horizontally and vertically, no border, no vignette, flat even lighting, no single focal object
  • seamless tileable surface texture of <description>,
  • , perfectly seamless edges, flat even lighting, no perspective, no objects, photographed-flat appearance
  • The repeating units must be aligned so rows continue across the tile borders: unit boundaries at the left edge continue into the right edge and across top/bottom.
  • Reduce the visible periodic repetition of distinctive features so the pattern feels organic.

Key colour for transparencyfrom production skills

Models cannot output alpha, so asking for "transparent background" produces a fake checkerboard. Default magenta; switch to green if the subject's own colours are near pink/magenta/purple; blue if both are taken. The chosen colour must appear in both the prompt suffix and the keying script.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

bright magenta #FF00FF, bright green #00FF00, bright blue #0000FF

Fragments to drop in

  • on a solid uniform bright <KEY COLOR> background, crisp edges, no drop shadow outside the element
  • Clean uniform <key color> background.

Reference-image role and what it controlsfrom production skills

For image-to-image the prompt should describe what changes, not redescribe the input: bad "a man with brown hair in a leather jacket holding coffee, made into anime"; good "transform into anime style, vibrant colors, soft cel shading".

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

style donors only, Image0, Image1, Image2, official logo, approved base scene, style-reference, representative-visual, grade/energy only

Fragments to drop in

  • Make an Animated Explainer. Take only the visual render style and color grading of the input image(s); mix the styles if there is more than one image. Never use the characters, inscriptions, etc. from the input image(s) unless the instructions below ask you to. Use only the render style, and follow the user's instructions below:
  • Use references as style donors only. Never copy their people, text, logos, or objects.
  • Label references by role in prompts: `Image 1: official logo`, `Image 2: approved base scene`.
  • State what each reference controls and what it must not control.

Design-board direction vocabulary (web/layout imagery)from production skills

Commit to one option per category before prompting and hold it across all boards; one board per section, landscape 16:9 or 3:2, never one tall full-page image.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

Pristine Light (paper/cream/off-white, dark ink), Deep Dark (charcoal/graphite), Bold Studio Solid (oxblood, royal blue, forest, vermilion, emerald fields), Quiet Premium Neutral (bone, sand, taupe, stone, smoke), technical grid/dot field, solid with soft ambient depth, full-bleed cinematic imagery, tactile paper/material texture, cinematic centered minimalist, asymmetric split, floating polaroid scatter, inline typography behemoth, editorial offset, massive image-first with restrained text, modular bento rhythm, alternating editorial blocks, poster-stacked storytelling, gallery-led cadence, Swiss grid discipline, asymmetric premium flow, artifact/collectible, journey/waypoints, tool/precision instrument, living system/garden, stage/spotlight, archive/dossier

Fragments to drop in

  • "website design mockup, desktop landing page section, [SECTION ROLE], [theme paradigm + exact palette words], [typography character] typography, [hero architecture / composition anchor], [background mode], [narrative spine motif], professional layout, clear hierarchy and spacing, award-winning web design" , plus: "no watermark, no browser chrome".
  • Name real content in the prompt (the actual headline wording you plan) so type sits believably.

Reference photos for identity trainingfrom production skills

5-20 photos, 8-12 the sweet spot, clear face with eyes visible, single person per photo, sharp, ≥1024×1024 ideal.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

front, 3/4 left, 3/4 right, slight up/down, indoor, outdoor, soft, harsh, neutral, smiling, talking, head shot, head-and-shoulders, full body

Fragments to drop in

  • Higher variety = better identity capture.

Image-to-3D input framingfrom production skills

One figure per input image , a 4-view character sheet reconstructs as "a statuette of 4 fused figures".

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

t-pose, a-pose, Full body in frame, nothing cut off, Clean/plain background, Limbs visually separated from the torso, One character, no props overlapping the silhouette, lowpoly, three-quarter isometric view

Fragments to drop in

  • Full body in frame, nothing cut off.
  • Clean/plain background; no shadows, text, watermarks, logos.
  • Limbs visually separated from the torso (arms not pressed to the body) , fused silhouettes produce fused geometry and break auto-rig.

The order to write them in

A prompt is read start to finish, and earlier clauses set the frame the later ones are interpreted against. This order puts the decisions that cannot be undone first.

  1. Aspect ratio
  2. Medium and style
  3. Shot size
  4. Camera angle
  5. Composition and framing
  6. Lens and focal length
  7. Depth of field
  8. Lighting
  9. Colour and palette
  10. Materials and finish
  11. Background
  12. Mood
  13. Text handling
  14. Negatives

How production prompts are laid out

The base order is subject + setting + style to camera (lens, angle, motion) to lighting to style/medium, kept under about 200 tokens because "Models distort with very long prompts" (higgsfield-generate/references/prompt-engineering.md). Specialised skills expand this into numbered block contracts assembled in a fixed order. The thumbnail contract has eleven blocks: frame to scene brief to text to subjects to key elements to logo to location to composition to background to lighting on people to grade, with per-character IDENTITY LOCK paragraphs and an opening IMAGE REFERENCES manifest when two or more references are attached (higgsfield-youtube-thumbnail/SKILL.md). The cover doctrine has six: subject + action to world + props to scale + placement to palette to light + lens to and "Always end with: no text, no letters, no logos, no captions, no UI" (higgsfield-websites/references/app-cover.md). Game assets concatenate exactly four parts in order: `<kind template> + <asset description> + <STYLE FORMULA byte-identical> + <kind suffix>` (higgsfield-websites/references/game-stylization.md). Recraft logo prompts are "exactly one continuous enhanced Recraft prompt string per candidate , no bullet points inside the prompt" ordered mark type to central subject/mechanism to shape logic to style register to palette behavior to composition to constraint tail, and "Every clause must materially affect the drawing" (higgsfield-brandkit/references/logo-prompt-enhancer.md). Brandkit mockups use bracketed labelled sections , [CREATE ONE FINISHED BRANDED MOCKUP], [AUTHORITATIVE LOGO], [PLACEMENT LOCK], [PHYSICAL APPLICATION] , plus a compact [BRAND LOCK , DO NOT DEVIATE] block repeated verbatim across related assets (higgsfield-brandkit/references/mockups.md, higgsfield-brandkit/references/brand-lock.md).

What to rule out

The most common element in the whole corpus. Put these last, in one sentence, and only include what is plausible for the shot: a long list of irrelevant prohibitions wastes the model's attention.

Text and branding

text, logos, captions, watermark, watermarks, logo, labels

Anatomy and faces

hands, arms, extra people, extra characters

Frame artefacts

frame, border, mockup

Content and safety

scary elements, weapons prominently displayed

Style and clutter

people, props

Other

overlays, slogan, letters, sunglasses, dramatic pose, helmet, police department names, cinematic color grading, character redesign, new clothing, new pose, perspective

Fragment to drop in

  • No text, logos, captions, watermark, watermarks, logo, hands, arms, extra people, extra characters, frame, border, mockup, scary elements, weapons prominently displayed, people, props, overlays, slogan, letters, sunglasses, dramatic pose, helmet.

Exclusion lines used in production

Copied verbatim from working prompt skills. Useful as whole lines rather than as a vocabulary.

  • no text, no letters, no logos, no captions, no UI
  • No text, no logos, no watermark.
  • No text, no readable UI labels, no watermark.
  • no text, no logos, no watermark
  • no extra text, no watermark
  • no watermark, no browser chrome
  • no shadows, no texture, no text
  • no fills, no shadows, no texture, no text
  • no border, no vignette
  • no single focal object
  • no characters, no UI elements
  • no drop shadow outside the element
  • no perspective, no objects
  • no shadows cast on the background, no ground plane, nothing cropped at the edges

Audio prompts

Measured across 163 audio prompts, mean length 1253 characters.

Audio is the thinnest part of this corpus at 163 prompts against 4,300 for video, so treat the audio counts as indicative. Video prompts frequently describe their own audio bed, so much of the audio vocabulary was learned there.

Download the audio sheet as a PDF, one page, or take all three.

Duration38.0% of audio prompts

Length of the piece.

A generator will otherwise give you a loop of arbitrary length that has to be cut, which is where the musical ending gets lost.

Values

15 seconds, 30 seconds, 60 seconds, 90 seconds, 120 seconds

Also used in production skills, not counted

--duration 30, --duration 12, --duration 4, --duration 2

Fragments to drop in

  • 90-second bed.
  • 30 seconds with a clean ending.
  • 15-second sting.

From production skills

  • mirelo_text_to_audio --prompt "glass breaking in a large hall" --duration 4
  • sonilo_music --prompt "cinematic synthwave track" --duration 12

Sonilo Music and Mirelo Text to Audio both require an explicit --duration; Seed Audio does not.

Role

What the audio is for.

A bed under narration and a standalone track need opposite decisions about melody and midrange, so state the job before the style.

Also used in production skills, not counted

no music, no voice, no ambience, dry studio recording, close microphone, short isolated, single, sound effects, ambience, foley, impacts, transitions, environmental sounds

Fragments to drop in

  • Instrumental music bed to sit under narration.
  • Standalone track, no voice over it.
  • Short sting to close a scene.

From production skills

  • short isolated sword impact, dry studio recording, no music
  • single heavy wooden door slam, close microphone, no ambience
  • glass breaking in a large hall
  • cinematic rain ambience with distant thunder
  • Describe one sound per SFX prompt. Do not request a complete mixed scene.
  • State `no music`, `no voice`, or `no ambience` when isolation matters.

This is the single richest audio prompt guidance in the repo. One sound per prompt is the hard rule.

Tempo38.7% of audio prompts

Speed, ideally as a number.

39% of audio prompts in the corpus give a BPM. A number is obeyed far more reliably than a word like upbeat.

Values

70 BPM, 90 BPM, 100 BPM, 116 BPM, 128 BPM

Fragments to drop in

  • Buoyant mid-tempo around 116 BPM.
  • Slow, roughly 70 BPM.
  • Driving, 128 BPM, steady four on the floor.

Key and tonality15.3% of audio prompts

Major or minor, and the key if it matters.

The cheapest emotional lever in audio. Major reads as safe and bright, minor as unresolved, and models respond to both.

Values

major key, minor key, modal, pentatonic

Fragments to drop in

  • Bright major key.
  • Minor key, unresolved.
  • Warm, clean major-key resolution in the final eight seconds.

Instrumentation71.2% of audio prompts

Which instruments play.

Named in 71% of audio prompts, the most used aspect of any type. Listing instruments works better than naming a genre, which models interpret loosely.

Values

percussion, ukulele, piano, bass, handclaps, marimba, bells, strings, guitar, xylophone, drums, harp, synthesizer, pads, 808, synth, flute, choir, synthesized

Also used in production skills, not counted

instrumental, no vocals, cozy forest loop, warm marimba and soft strings, cinematic synthwave track, backing tracks, instrumental beds, jingles, musical moods

Fragments to drop in

  • Marimba, ukulele, xylophone, light percussion and soft handclaps.
  • Solo piano with a distant string pad.
  • Deep electronic ambience with soft mechanical clicks.

From production skills

  • instrumental cozy forest loop, warm marimba and soft strings, no vocals
  • cinematic synthwave track
  • Music prompts specify mood, tempo, instrumentation, and `instrumental/no vocals` when lyrics are unwanted.

Tempo is named as a required aspect but the repo gives no BPM values or key/scale vocabulary anywhere , this is genuinely thin.

Mood20.2% of audio prompts

The feeling.

Pairs with tempo and key. Together those three settle most of the composition.

Values

warm, playful, cheerful, buoyant, wholesome, ambient, uplifting, driving, triumphant, minimal, dreamy, suspenseful

Also used in production skills, not counted

material, era, energy, tonal language

Fragments to drop in

  • Cheerful, wholesome and playful.
  • Calm and spacious.
  • Suspenseful and restrained.

From production skills

  • Preserve the game's STYLE FORMULA conceptually through material, era, energy, and tonal language even though audio is not visual.

The visual style contract is carried into audio prompts by translation, not by pasting the formula.

Arrangement and dynamics10.4% of audio prompts

How the piece develops.

Without this you get four bars looped flat. Naming a build and an ending is what makes it usable against picture.

Values

airy, midrange, headroom, no abrupt hits, narration-friendly, smooth transitions, gentle dynamics

Also used in production skills, not counted

-6 dBFS, -10 to -12 dBFS, -18 to -20 dBFS, -3 dBFS true peak, short fades at loop boundaries

Fragments to drop in

  • Airy and narration-friendly, gentle dynamics, no abrupt hits.
  • Build subtle energy throughout, then a celebratory lift at the end.
  • Clear midrange space left for a voice.

From production skills

  • Voice around -6 dBFS.
  • SFX around -10 to -12 dBFS.
  • Music around -18 to -20 dBFS.
  • Final true peak at or below -3 dBFS.
  • Voice stays above SFX; SFX stays above music.
  • Never stack raw model outputs at full gain.

Post-generation mixing, not prompt text, but it is the only mix guidance in the corpus.

Vocals62.6% of audio prompts

Whether anything sings or speaks.

Named in 63% of audio prompts, almost always as a prohibition. Generators add humming and wordless vocals unprompted.

Fragments to drop in

  • No vocals, no lyrics, no spoken words.
  • Wordless humming only, no intelligible words.
  • No prominent lead melody competing with the narration.

Ending

How it stops.

The most common practical complaint about generated audio is that it fades or stops mid-phrase. Ask for a resolution.

Fragments to drop in

  • End on a clean major-key resolution, no fade.
  • Ring out and decay naturally.
  • Hard stop on the final beat.

Voice and spoken linefrom production skills

The prompt IS the line to be spoken; the voice itself is selected by id, never described in prose. Adjust --speech_rate modestly rather than rewriting a take.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

voice_type preset|element, voice_id, speech_rate, one voice per speaking entity, exact line, performance direction

Fragments to drop in

  • The gate is open. Move!
  • same voice, calmer delivery
  • Voice prompts contain the exact line; performance direction belongs in the model-supported prompt, not undocumented flags.
  • Voice: lock one voice per speaking entity and reuse it. Keep lines short enough not to block gameplay.

Narration writing (spoken script)from production skills

Narration is the one thing written in the user's chosen language; every image and video prompt stays English.

Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.

Values

20-24 words, about 8-9 seconds, under roughly 9.5 seconds, plain spoken text only, no timecodes, emotion cues, parentheticals, or stage directions, Spell numbers out, hook through payoff, ~150 words per minute

Fragments to drop in

  • Write one plain line per clip, sized for about 8-9 seconds and normally 20-24 words. Keep every take under roughly 9.5 seconds.
  • Use no timecodes, emotion cues, parentheticals, or stage directions.
  • Spell numbers out.
  • Set tone through word choice and concrete detail.
  • Never say "in this video."
  • For four and a half thousand years, the pyramids of Egypt have stood against the desert, silent and immense.
  • Write scripts for speech, not text: short sentences, natural pauses, ~150 words per minute target. A 60-second video = ~150 words.
  • Don't pad scripts to fit duration. The model paces itself; over-stuffed scripts get rushed delivery.

The order to write them in

A prompt is read start to finish, and earlier clauses set the frame the later ones are interpreted against. This order puts the decisions that cannot be undone first.

  1. Duration
  2. Role
  3. Mood
  4. Key and tonality
  5. Tempo
  6. Instrumentation
  7. Arrangement and dynamics
  8. Vocals
  9. Ending
  10. Negatives

How production prompts are laid out

Thin and single-sentence. One sound per prompt, described as material + action + recording character + an isolation clause, e.g. "short isolated sword impact, dry studio recording, no music". Music prompts name mood, tempo, instrumentation and then `instrumental/no vocals`. Voice prompts contain the exact line to be spoken and nothing else, with the voice chosen by id rather than described. Narration for assembled video is written as one plain labelled line per block (`Block 1` then the line), 20-24 words, about 8-9 seconds, with no timecodes, emotion cues, parentheticals or stage directions and numbers spelled out (higgsfield-websites/references/game-audio.md, higgsfield-video-explainer/references/prompts.md).

What to rule out

The most common element in the whole corpus. Put these last, in one sentence, and only include what is plausible for the shot: a long list of irrelevant prohibitions wastes the model's attention.

Text and branding

spoken words, words, other words, harsh textures, ad-libs that obscure the words

Content and safety

scary sounds

Audio

vocals, lyrics, sound effects, speech, music, spoken narration, voices, melody, singing, ad-libs, background music, vocal-like sounds, prominent lead melody

Other

medical sounds, abrupt peaks, dark, dramatic tension, heavy drums, long instrumental passages, sung parts, dialogue, chanting, tense passages, harsh timbres, abrupt stops

Fragment to drop in

  • No spoken words, words, other words, harsh textures, ad-libs that obscure the words, scary sounds, vocals, lyrics, sound effects, speech, music, spoken narration, medical sounds, abrupt peaks, dark, dramatic tension, heavy drums, long instrumental passages.

Exclusion lines used in production

Copied verbatim from working prompt skills. Useful as whole lines rather than as a vocabulary.

  • no music
  • no voice
  • no ambience
  • no vocals
  • instrumental
  • no voice, dialogue, or narration
  • Never stack raw model outputs at full gain.

How this was measured

Every value on this page was counted in a corpus of 8,835 real prompts, 6,161 of them unique. The number beside each aspect is the share of prompts that specify it, and the values are the terms those prompts actually used, most used first. Nothing here is a list of plausible sounding options.

That distinction matters because the ranking is the useful part. A guide written from memory would lead with camera angles and lenses. In practice the most common element of all is the negative clause, present in 49.3% of video prompts and 73.6% of audio prompts, which is not where most advice starts.

Guidance and fragments are written by hand. Values are not. The process is documented as a repeatable skill so this can be rebuilt from a fresh export.

The cheatsheet, in full

Everything above condensed to one table per content type. This is the same content as the PDF, so you can read it here or take it with you. There is a standalone page for it too if you want to print from the browser.

Video

Order: duration > aspect > medium > shotSize > angle > movement > lens > dof > lighting > grade > motion > mood > audioBed > pacing > negatives

AspectValuesFragment
Duration6 seconds, 8 seconds, 10 seconds, 15 seconds, 30 seconds, 60 seconds15-second cinematic clip.
Aspect ratio16:9, 9:16, 1:1Framed 16:9 for landscape playback.
Shot sizeclose-up, wide shot, medium shot, extreme close-up, establishing shot, two-shot, macro shot, over-the-shoulder, closeupExtreme close-up on the texture, filling the frame.
Camera angleoverhead, low-angle, eye level, three-quarter, top-down, profile, high-angleLow-angle looking up, subject towering over the lens.
Camera movementpush-in, tracking shot, handheld, orbit, locked-off, slow zoom, truck, slow dolly, parallax, pull-back, slow pan, whip pan, crane up, dolly inSlow push-in toward the subject, ending tight.
Lens and focal length35mm, macro lens, anamorphic, 100mm, 50mm, 85mm, 28mm, wide-angle lens, 32mm, 50 mm, fisheye, telephoto, 40mm, 35 mmShot on an 85mm lens, compressed and flattering.
Depth of fieldshallow depth of field, bokeh, rack focus, f/1.4, f/1.8, f/2.0, soft focus, f/2.8Shallow depth of field, background dissolving into bokeh.
Lightingsoft lighting, rim lighting, rim light, silhouette, volumetric lighting, golden hour, volumetric light, soft light, side lighting, chiaroscuro, key light, low-key, practical lighting, backlitSoft lighting from a large source, shadows barely there.
Colour and gradevibrant, pastel, film grain, black-and-white, natural color, high contrast, desaturated, crushed blacks, natural colour, low contrastVibrant saturated grade, colours pushed but not clipping.
Medium and stylecinematic, photorealistic, anime, documentary, 3d render, unreal engine, stop-motion, 35mm film, watercolor, vector, photoreal, claymation, line art, octaneCinematic, shot on 35mm film.
Subject motionreveal, float, drift, slow motion, glides, rotates, time-lapse, ripples, unfolds, hovers, cascadeThe lid glides open on hidden runners, slow and even.
Moodplayful, energetic, whimsical, calm, dramatic, minimal, elegant, uplifting, moody, ominous, ethereal, luxurious, serene, nostalgicCalm and unhurried throughout.
Audio bedpercussion, bells, ukulele, xylophone, piano, drums, bass, marimba, strings, synth, guitar, choir, handclaps, padsMinimal deep electronic ambience with soft mechanical clicks.
Pacing and cutsOne continuous take, no cuts.
Resolution / quality tier480p, 720p, 1080p, 4k, basic, high, ultra, pro, std, standard, bitrate_mode standard|high--resolution 4k
First/last frame anchoring--start-image, --end-image, start frame, end frame, start = end, one-shot, loop, end-frame reveal, generic image/reference role , not a start-frame role`--start-image` anchors the first frame. Prompt describes motion.
Ad format / mode (branded video)ugc, ugc_how_to, ugc_unboxing, product_showcase, product_review, tv_spot, wild_card, ugc_virtual_try_on, virtual_try_on, hook, setting, ad referenceLooks like a real person filmed on phone
Edit instruction (video editing workflows)make the jacket red, sketch, timestamp--prompt "make the jacket red"

Rule out. Text and branding: text, logos, subtitles, watermark, watermarks, captions, on-screen text, text on screen; Anatomy and faces: extra limbs, distorted faces, extra characters, hands, extra fingers, extra animals; Content and safety: violence, scary imagery, scary elements, weapons; Audio: music, narration; Other: dialogue, copyrighted characters, talking, cuts, horror, humans, clothes, flickering

Image

Order: aspect > medium > shotSize > angle > composition > lens > dof > lighting > grade > material > background > mood > text > negatives

AspectValuesFragment
Aspect ratio16:9, 9:16, 1:1, 4:5, 21:9, 3:2, 4:3Square 1:1.
Composition and framingclose-up, wide shot, medium shot, extreme close-up, two-shot, closeup, over-the-shoulder, establishing shot, macro shot, full shotCentred on a single vertical axis, symmetrical.
Shot sizeclose-up, wide shot, medium shot, extreme close-up, two-shot, closeup, over-the-shoulder, establishing shot, macro shot, full shotExtreme close-up, texture filling the frame.
Camera anglethree-quarter, profile, eye-level, low angle, overhead, high angle, top-down, bird's-eye, dutch angleThree-quarter view, turned slightly off axis.
Lens and focal length35mm, 85mm, 50mm, 110 mm, 16mm, anamorphic, 24mm, 70mm, 48mm, fisheye, 600mm, 450mm, 40mm, 297 mm85mm portrait lens, natural compression.
Depth of fieldshallow depth of field, bokeh, f/1.4, f/7.1, soft focus, f/2.8, deep focus, f/2.0, f/4, f/2, f/1.8, f/5.6Shallow depth of field at f/1.8, background soft.
Lightingsilhouette, rim light, soft light, soft lighting, backlit, key light, golden hour, high-key, rim lighting, backlight, softbox, low-key, volumetric light, bounce lightLarge softbox front left, soft shadows falling right.
Colour and palettevibrant, pastel, film grain, black-and-white, desaturated, high contrast, natural color, natural colour, monochrome, low contrast, sepiaVibrant, fully saturated, clean whites.
Medium and stylecinematic, photorealistic, documentary, anime, 3d render, line art, watercolor, vector, 16mm, 35mm film, photoreal, kodak vision3, isometric, flat illustrationPhotorealistic editorial photograph.
Materials and finishglossy, matte, ceramic, concrete, leather, velvet, satin, brass, linen, frosted, marble, chrome, walnut, oakBrushed metal with a satin finish, no mirror reflections.
BackgroundSeamless mid-grey studio backdrop, no horizon line.
Moodelegant, playful, minimal, dramatic, calm, energetic, whimsical, moody, luxurious, nostalgic, uplifting, ethereal, cosy, grittyElegant and restrained.
Text handlingNo text, no lettering, no numbers anywhere in the image.
Subject + setting + style (the opening clause)a red fox curled in a snowy pine forest, golden hour, cinematic, subject + action, concrete and physicala red fox curled in a snowy pine forest, golden hour, cinematic
Typography named in the prompt (when text is generated)Anton, Anton SC, Bebas Neue, Oswald, Archivo Black, Montserrat, Poppins, Roboto Condensed, Inter, Barlow Condensed, Playfair Display, Cormorant Garamond, DM Serif Display, FrauncesThe prompt must state: Exact literal copy; Display/body font family names and which text uses each
Identity lock / subject fidelityIDENTITY LOCK, photographic identity match, same bone structure, eye shape, nose, lips, jawline, skin tone, hairline and hair texture, Do not beautify, average, or restyle the face, IMAGE REFERENCES manifestCHARACTER N: the person from attached face reference #K , IDENTITY LOCK: reproduce this exact person with a photographic identity match , same bone structure, eye shape, nose, lips, jawline, skin tone, hairline and hair texture. Do not beautify, average, or restyle the face. Expression: <emotion phrase>.
Aesthetic / seasonal presetClean girl, Cottagecore, Quiet luxury, Dark academia, Y2K, Christmas, Valentine's, Halloween, Black Friday, None, Clean studio, Lifestyle, Conceptual, With a modelChristmas version, quiet-luxury aesthetic
Product-photo genre / modeproduct_shot, lifestyle_scene, closeup_product_with_person, moodboard_pin, hero_banner, social_carousel, ad_creative_pack, virtual_model_tryout, conceptual_product, restyle, levitating, floating, splash, frozen motionbottle of cold-brew on a sunlit kitchen counter, IG feed
Logo / vector mark specificationlettermark_monogram, pictorial, abstract, mascot, emblem, one dominant unified silhouette, negative-space device, motif treatment, unexpected locked color-role pairing, symmetry or intentional asymmetry, internal detail limit, small-size scalability, centered isolated presentation, minimal anchor pointsFlat vector design, clean lines, no shadows, no texture, no text. Clean editable vector paths, SVG-friendly, minimal anchor points.
Logo placement on a mockup surfaceTarget surface, Position, Scale, Orientation, Color, horizontally centered, upper third, optical center aligned to panel, clear-space margin, align to <panel edge/seam/baseline>, follow surface perspective[PLACEMENT LOCK] Target surface: <exact object panel/face/material>. Position: <exact alignment and location, e.g. horizontally centered, upper third, optical center aligned to panel>. Scale: logo occupies <specific proportion> of the target surface while keeping <specific clear-space margin>. Orientation: align to <panel edge/seam/baseline>; follow surface perspective without changing logo proportions. Color: use the supplied <full-color/black/white> variant exactly. State why it contrasts correctly with the material/background.
Seamless / tileable behaviourseamless tileable game texture tile, uniform pattern density, perfectly seamless edges that wrap horizontally and vertically, no border, no vignette, no single focal object, photographed-flat appearance, seamless repeating patternseamless tileable game texture tile of <description>, uniform pattern density,
Key colour for transparencybright magenta #FF00FF, bright green #00FF00, bright blue #0000FFon a solid uniform bright <KEY COLOR> background, crisp edges, no drop shadow outside the element
Reference-image role and what it controlsstyle donors only, Image0, Image1, Image2, official logo, approved base scene, style-reference, representative-visual, grade/energy onlyMake an Animated Explainer. Take only the visual render style and color grading of the input image(s); mix the styles if there is more than one image. Never use the characters, inscriptions, etc. from the input image(s) unless the instructions below ask you to. Use only the render style, and follow the user's instructions below:
Design-board direction vocabulary (web/layout imagery)Pristine Light (paper/cream/off-white, dark ink), Deep Dark (charcoal/graphite), Bold Studio Solid (oxblood, royal blue, forest, vermilion, emerald fields), Quiet Premium Neutral (bone, sand, taupe, stone, smoke), technical grid/dot field, solid with soft ambient depth, full-bleed cinematic imagery, tactile paper/material texture, cinematic centered minimalist, asymmetric split, floating polaroid scatter, inline typography behemoth, editorial offset, massive image-first with restrained text"website design mockup, desktop landing page section, [SECTION ROLE], [theme paradigm + exact palette words], [typography character] typography, [hero architecture / composition anchor], [background mode], [narrative spine motif], professional layout, clear hierarchy and spacing, award-winning web design" , plus: "no watermark, no browser chrome".
Reference photos for identity trainingfront, 3/4 left, 3/4 right, slight up/down, indoor, outdoor, soft, harsh, neutral, smiling, talking, head shot, head-and-shoulders, full bodyHigher variety = better identity capture.
Image-to-3D input framingt-pose, a-pose, Full body in frame, nothing cut off, Clean/plain background, Limbs visually separated from the torso, One character, no props overlapping the silhouette, lowpoly, three-quarter isometric viewFull body in frame, nothing cut off.

Rule out. Text and branding: text, logos, captions, watermark, watermarks, logo, labels; Anatomy and faces: hands, arms, extra people, extra characters; Frame artefacts: frame, border, mockup; Content and safety: scary elements, weapons prominently displayed; Style and clutter: people, props; Other: overlays, slogan, letters, sunglasses, dramatic pose, helmet, police department names, cinematic color grading

Audio

Order: duration > role > mood > key > tempo > instrumentation > arrangement > vocals > ending > negatives

AspectValuesFragment
Duration15 seconds, 30 seconds, 60 seconds, 90 seconds, 120 seconds90-second bed.
RoleInstrumental music bed to sit under narration.
Tempo70 BPM, 90 BPM, 100 BPM, 116 BPM, 128 BPMBuoyant mid-tempo around 116 BPM.
Key and tonalitymajor key, minor key, modal, pentatonicBright major key.
Instrumentationpercussion, ukulele, piano, bass, handclaps, marimba, bells, strings, guitar, xylophone, drums, harp, synthesizer, padsMarimba, ukulele, xylophone, light percussion and soft handclaps.
Moodwarm, playful, cheerful, buoyant, wholesome, ambient, uplifting, driving, triumphant, minimal, dreamy, suspensefulCheerful, wholesome and playful.
Arrangement and dynamicsairy, midrange, headroom, no abrupt hits, narration-friendly, smooth transitions, gentle dynamicsAiry and narration-friendly, gentle dynamics, no abrupt hits.
VocalsNo vocals, no lyrics, no spoken words.
EndingEnd on a clean major-key resolution, no fade.
Voice and spoken linevoice_type preset|element, voice_id, speech_rate, one voice per speaking entity, exact line, performance directionThe gate is open. Move!
Narration writing (spoken script)20-24 words, about 8-9 seconds, under roughly 9.5 seconds, plain spoken text only, no timecodes, emotion cues, parentheticals, or stage directions, Spell numbers out, hook through payoff, ~150 words per minuteWrite one plain line per clip, sized for about 8-9 seconds and normally 20-24 words. Keep every take under roughly 9.5 seconds.

Rule out. Text and branding: spoken words, words, other words, harsh textures, ad-libs that obscure the words; Content and safety: scary sounds; Audio: vocals, lyrics, sound effects, speech, music, spoken narration, voices, melody; Other: medical sounds, abrupt peaks, dark, dramatic tension, heavy drums, long instrumental passages, sung parts, dialogue

Say what you want, not what you do not

Most image and video models do not expose a separate negative prompt field, so a prohibition inside the prompt can act as a mention and pull the thing you banned into the frame. Where that is a risk, state the positive instead.

  • Most models don't expose a `negative_prompt`. Phrase positively:
  • Instead of "no blur" to "tack sharp"
  • Instead of "no people" to "uninhabited landscape"
  • NEGATIVE: color drift, photorealism, 3D render, lip-sync, captions, on-screen text, logos, watermark{, plus style-specific bans}.
  • NEGATIVE: <style drift and realism bans; no lip-sync, captions, text, logos, or watermark>
  • Regenerate with "no text, no letters, no logos, no captions" appended.
  • Never: <forbidden treatments>
  • Change ONLY the person's facial expression to: <phrase>. Keep identity, face structure, hair, pose, body, clothing, logo, background, lighting and composition EXACTLY unchanged, pixel-faithful.
  • Honor every `forbidden_elements` entry literally. Never replace one forbidden cliché with another generic symbol.
  • Known failure mode: false-positive `nsfw` flags on innocuous product prompts , the fix is removing mood/atmosphere words ("seamless loop feel", "ambient", "intimate") and re-describing the shot as a plain product film/photo.
  • Models reject prompts with `nsfw` or `ip_detected` terminal status. Avoid: Real public figures; Sexual content; Trademarks / branded characters

Which model for which job

Recorded from production prompt skills, not measured here. Model names move fast, so treat this as a snapshot of practice at the time it was written.

ModelBest forPrompt conventions
GPT Image 2 (gpt_image_2)Default high-fidelity image generation; graphic design, UI, banners, typography, and any brief with on-image text; generated product concept / packaging / can / bottle with brand name or label text; the controlled second stage that adds readable text or exact graphic details onto an existing scene; seamless-texture edits; character animation sheets as a 5x5 grid in one image.The model to reach for when literal copy must be inside the image , state exact literal copy, display/body font family names and which text uses each, logo placement/scale/clear space/colour variant, text hierarchy/alignment/line breaks/contrast, palette roles and aspect ratio. In brandkit it is used only for that controlled stage and must preserve Image0's camera, crop, objects, lighting, material, folds, shadows, perspective and background exactly. Does not currently expose 4:5 , generate 3:4 with a centred 4:5 safe area and crop. Texture prompts run at aspect 1:1, quality high, with the reference attached as a media input.
Nano Banana 2 (nano_banana_2)Fast everyday default for character, cartoon, stylized and reference-driven image work; edits; game sprites, tiles, backgrounds and textures; the style-key image in the explainer pipeline; marketplace cards under the hood.Reference-driven , pass references with repeated --image. For game assets use resolution 1k, AR 1:1 for sprite/tile/texture/ui and 16:9 for backgrounds, count 1 each. Keeping one primary generator across asset kinds is itself a cohesion win.
Nano Banana 2 Lite (nano_banana_2_lite)Fast reference-driven generation and edits when the brief is simple or speed/cost matters more than Pro fidelity.Supports up to 14 image references via image_references; aspect_ratio=auto requires at least one image reference.
Nano Banana Pro (nano_banana_pro)Harder briefs where Nano Banana 2 falls short; the main thumbnail render; photoreal and reference-driven website imagery; deriving multiple angles of one approved product so every view is recognizably the SAME product.For thumbnails, run at explicit 4K with the long multi-block prompt written to a file and piped on stdin; face photos first in character order, then the logo.
Higgsfield Soul 2.0 (text2image_soul_v2)Aesthetic UGC, fashion editorial, lifestyle character stills; Soul Character reference work.Accepts a Soul Character reference via --soul-id; quality tier 1.5k or 2k; aspect ratios 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3.
Soul Cinema / soul_cinematicCinematic stills and film-grade lighting; the pick when the user asks for "cinematic" or wants concept-art mood; longer talking-head-style Soul work.Same ratios as Soul 2.0 plus 21:9; quality 1.5k or 2k; used for the PHOTO route on covers with 1-2 hosted scene references as --image.
Soul CastDistinctive, characterful personas , a creative, expressive character rather than a photoreal default.Text-only; prompt-only model that rejects reference media.
Soul Location (soul_location)Best-in-class environments and locations; pure scene and place generation with no person in frame.Prompt-only, no media. No quality/resolution selector , dimensions are fixed by aspect ratio.
Seedream (seedream_v5_pro, seedream_v4_5, Seedream 5.0 Lite)Primary photoreal mockup generator in brandkit; face-anchored edits into a complex new scene without heavy filters; surgical thumbnail tweaks.Mockup prompts use the bracketed [CREATE ONE FINISHED BRANDED MOCKUP] / [AUTHORITATIVE LOGO] / [PLACEMENT LOCK] / [PHYSICAL APPLICATION] contract with Image0/Image1 roles stated explicitly; never pass an SVG path as an image reference; never ask it to render readable text. For edits, scope narrowly and state that every other pixel-level property is unchanged. No 4:5 ratio.
Recraft V4.1 (recraft_v4_1)Clean graphic and vector-style design assets , logos, icons, flat illustrations, brand marks, controlled-palette visuals; the only model used for new logo marks.Use --model_type vector for vector output, standard for raster-style graphics. One continuous prompt string per candidate, no bullets, ordered mark type to mechanism to shape logic to style register to palette behavior to composition to constraint tail. Never put hex, RGB or Pantone in the prompt , pass colours through the colors and background_color params. Never use lens, camera, lighting, depth of field, photorealistic, cinematic, material rendering, grain, paper texture, shadows, mockup or scene language. The exact phrase "no text" must appear in the constraint tail. All three candidates use identical model, palette, background, aspect and quality parameters. Prompt-only , rejects media inputs.
Flux 2.0 (flux_2)Precise prompt adherence; a creative alternative to the Banana family; key-pose generation in the sprite pipeline; cheap photorealistic drafts.Key-pose prompts carry the STYLE FORMULA verbatim plus one pose instruction, and must demand full body in frame with empty margin above the head and below the feet plus a clean uniform key-colour background absent from the subject's palette.
Flux Kontext MaxContext-aware editing and style transfer; anime, stylized looks, typography remix when defaults feel too generic.Reach for it when default models feel flat.
Z Image (z_image)Fastest in the catalog , speed, drafts and LoRA-driven stylization; cheap prompt iteration.Prompt-only, rejects media. Accepts a prompt on stdin.
Grok ImagineExpressive, high-contrast, bold creative output; worth trying for anime and stylized looks.Also has a video variant with audio support.
Seedance 2.0 (seedance_2_0)Default all-purpose serious video , multi-shot, consistent identity, motion-heavy, cinematic, image-to-video, 4-15s up to 4K; seamless loops; the end-frame cover reveal; dramatic motion from a still.Accepts image, start_image, end_image, video and audio roles. Pass the finished cover as --end-image and describe the buildup so the clip lands pixel-perfect on it; do NOT pass it as --start-image. Audio reference comes through the audio media role, never --generate-audio. Aspect ratios auto, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16; duration 4-15; resolution 480p/720p/1080p/4k.
Seedance 2.5 (seedance_2_5)Reference-driven video, editing an existing video, or extending one.Modes t2v / omni_reference / video_edit / video_extension with image/video/audio reference arrays. Caps at 720p , not a newer Seedance 2.0.
Seedance 1.5 Pro / seedance1_5Cheap clean single-take shots without cuts; the exact model for the 2D sprite animation pipeline.For sprites: fixed 4 s, 720p, aspect ratio equal to the key pose and passed explicitly on every call, start frame always, end frame only for loops, one action in the positive prompt plus the mandatory camera/prop/facing negative block. "Do not silently substitute a model."
Kling 3.0 (kling3_0) and Kling 3.0 TurboCheaper Seedance 2.0 substitute for single-plane scenes without heavy motion; cleanest image-to-video with a start frame; Turbo when the user wants speed or lower cost.Roles start_image and end_image (Turbo start_image only). 16:9, 9:16, 1:1; duration 3-15s; modes pro/std; sound on/off.
Gemini Omni Flash (gemini_omni)Fast multimodal reference-to-video from image references and optionally one video reference; the clip generator in the explainer pipeline.0-7 image references and 0-1 video reference (max 5 images when a video reference is present). In the explainer pipeline every call attaches the same style key via --image and runs 10s at 720p. "Never silently replace gemini_omni."
Google Veo 3.1 / Veo 3.1 Lite / Veo 3Ultra-realistic top-tier cinematic quality; Lite for fast batch and volume work.Format set is constrained , Veo 3.1 accepts only 16:9 or 9:16 and durations 4, 6 or 8, with quality tiers basic/high/ultra. Verify accepted aspect ratio and duration before submitting.
Grok Video 1.5 (grok_video_v15)Bold, anime-like, high-contrast or experimental image-to-video from one starting image.Requires exactly one --start-image or --image; duration 2-15s; resolution 480p or 720p.
Minimax HailuoCheap with strong physics when natural-physics motion matters and audio is not needed.No audio in current variants.
Wan 2.7 / Wan 2.6Synchronized audio with character-consistent video (2.7); stylized, experimental, intentionally artistic cheap work (2.6).Some Wan configs are prompt-only and reject media inputs.
Cinema Studio Video 3.0 / Cinema Studio Image 2.5Top-tier cinema-grade execution and dramatic film look at the highest fidelity; cinematic still frames up to 4K.Prefer Cinema Studio Video 3.0 as the modern cinema default; reach for earlier variants only when the user names them.
Marketing Studio (marketing_studio_video / marketing_studio_image)All advertising and commercial work , UGC, unboxing, TV spot, product showcase, product review, virtual try-on, branded ad images with a presenter avatar and product.The prompt is the hook/brief; a selected hook's text is prepended to it and must not be copied into --prompt unless the user wants the wording reinforced. Mode drives staging (default ugc). hook_id and setting_id are video-only and valid only for ugc, ugc_how_to, ugc_unboxing, product_review and ugc_virtual_try_on; hooks are weak without product context. --generate_audio true is supported here (unlike seedance_2_0). Aspect auto/21:9/16:9/4:3/1:1/3:4/9:16; resolution 480p or 720p; duration integer ≥ 4.
Seed Audio 1.0 (seed_audio)Default audio generation , text-to-audio, sound effects, ambience, foley, impacts, environmental audio, voice-style generations and music-like audio; every narration take in the explainer pipeline.Requires --prompt. Optional audio_references or image_references, mutually exclusive. For narration the prompt is the block's spoken line only, with voice chosen by --voice_type (preset|element) and --voice_id, and --speech_rate adjusted modestly for an overlong take.
Sonilo Music (sonilo_music)Music from text , backing tracks, instrumental beds, jingles, musical moods.Requires --prompt and --duration; takes no media inputs. Prompt names mood, tempo, instrumentation and 'instrumental/no vocals'.
Mirelo Text to Audio (mirelo_text_to_audio)Non-speech audio , sound effects, ambience, foley, impacts, transitions, environmental sounds.Requires --prompt and --duration; no media inputs. One sound per prompt with an isolation clause.
Inworld Text to Speech (inworld_text_to_speech)Explicit TTS with one of its listed voices.The prompt is the exact line to be spoken; pass --voice with an exact voice value from the model contract.
Multi-Image to 3D (multi_image_to_3d)An actual 3D mesh/GLB from 1-4 object or product reference images.No text prompt , the reference images are the input. One figure per image, full body, clean plain background with no shadows/text/watermarks/logos, limbs separated from the torso; concept images use three-quarter isometric view and a pure flat white background. Set pose_mode t-pose or a-pose for characters that will be rigged.
AutoSmart routing when the user's intent is open and you don't want to commit to a specific image model.Picks the best image model from the prompt automatically.
Virality Predictor (brain_activity)Scoring a finished clip's hook, attention, retention, distraction risk and virality potential.Takes --video and needs no prompt at all; returns a text report, not media.
image_background_remover / bytedance_image_upscale / outpaintTransparent subject cutouts, upscaling art below the 1500px floor, extending an image.No prompt. Background removal runs on frames one by one, never on a video, because the video path smears edges into a flickering halo.

Prompt cheatsheet

Every aspect, value and negative on two pages. Keep it open while you write.

Download the PDF