How to write prompts for AI video, image and audio
A reference for prompting generated video, image and audio, measured from 8,835 real prompts. Which aspects to specify, the vocabulary that works, drop in fragments and what to rule out.
Prompt cheatsheet
Every aspect, value and negative on two pages. Keep it open while you write.
How to use this
- Work down the aspects in order. The order is the recommended order in the prompt too.
- One clause per aspect. A prompt is a list of decisions, not a paragraph of atmosphere.
- Prefer a number over an adjective wherever a number exists: 116 BPM over upbeat, 85mm over flattering, 15 seconds over short.
- Put the negatives last, in one sentence beginning with No.
- Reuse the identical phrasing across a batch. Paraphrasing is what makes a set of images look unrelated.
Video prompts
Measured across 4,300 video prompts, mean length 1537 characters.
Download the video sheet as a PDF, one page, or take all three.
Duration25.0% of video prompts
How long the clip runs.
Models hold a shot for a fixed span. Naming the length stops a model compressing a three beat idea into one, or padding a single beat with drift.
Values
6 seconds, 8 seconds, 10 seconds, 15 seconds, 30 seconds, 60 seconds
Also used in production skills, not counted
4-15s, 2-15s, 3-15s, 4, 6, 8, 10, 12, 15, 5, 4 s, ~15s, one to ten whole minutes
Fragments to drop in
15-second cinematic clip.Six seconds, one continuous take, no cuts.30 seconds in three beats of roughly ten seconds each.
From production skills
--duration 12--duration 5--duration 10--duration 4
Model-bound: Seedance 2.0 4-15s (12s is valid); Kling 3.0 and Kling 3.0 Turbo 3-15s; Grok Video 1.5 2-15s; Veo 3.1 accepts only 4, 6 or 8; Marketing Studio integer ≥ 4; explainer blocks are fixed 10-second units with N = duration_minutes × 6; sprite animation is fixed at 4 s; cover reveal ~5s; scroll-scrub single-shot asks for the longest single take the model supports (~15s).
Aspect ratio31.3% of video prompts
Frame shape.
Decides whether the subject can breathe. A vertical crop of a wide composition loses the sides, so the ratio has to be chosen before the composition is described.
Values
16:9, 9:16, 1:1
Also used in production skills, not counted
4:3, 3:4, 21:9, auto, 3:2, 2:3, 9:21, 4:5, 5:4, 1.91:1
Fragments to drop in
Framed 16:9 for landscape playback.Vertical 9:16, composed for a phone held upright.Square 1:1 with the subject centred.
From production skills
--aspect_ratio 16:9--aspect_ratio 9:16--aspect_ratio auto
Meanings given in prompt-engineering.md: 16:9 landscape, cinematic; 9:16 vertical, social; 1:1 square, profile / icon. Veo 3.1 accepts only 16:9 or 9:16. For sprite work the ratio must be read off the key-pose file and passed explicitly on every call , omitting it falls back to the tool default and crops the subject.
Shot size25.1% of video prompts
How much of the subject fills the frame.
The most reliable lever there is. Naming it removes the model's habit of defaulting to a mid shot for everything.
Values
close-up, wide shot, medium shot, extreme close-up, establishing shot, two-shot, macro shot, over-the-shoulder, closeup
Also used in production skills, not counted
one hero subject, kept centered, clean negative space, center-safe area, cover, full body in frame, empty margin above the head and below the feet
Fragments to drop in
Extreme close-up on the texture, filling the frame.Wide establishing shot, subject small in a large space.Medium shot from the waist up, room to move on both sides.
From production skills
One hero subject, kept centered, with clean negative space around it , that space is where the chapter copy sits.Full body in frame with empty margin above the head and below the feet.
The video fills the viewport and crops the edges, so essential subjects must stay away from the far left/right edges.
Camera angle6.8% of video prompts
Where the camera sits relative to the subject.
Changes who has power in the frame. Low angle makes a product monumental, overhead flattens it into a diagram.
Values
overhead, low-angle, eye level, three-quarter, top-down, profile, high-angle
Fragments to drop in
Low-angle looking up, subject towering over the lens.Directly overhead, top-down on the surface.Eye-level, camera at the subject's own height.
Camera movement19.1% of video prompts
How the camera travels during the shot.
The single most common cause of a shot feeling wrong. If you do not name it the model invents drift, and drift reads as an accident.
Values
push-in, tracking shot, handheld, orbit, locked-off, slow zoom, truck, slow dolly, parallax, pull-back, slow pan, whip pan, crane up, dolly in, steadicam, tilt down, dolly out, pan left, tilt up, dolly forward, dolly back
Also used in production skills, not counted
zooms in, dollies left, sweeping pan, slow push, fast whip, camera slowly pulls back, camera dollies in, slow orbit, rise, fly-through, lateral track, crane, detail push, rack focus soft to sharp, light sweep, Very slow push-in, Subtle camera push-in, slow push-in, drift, scale shock, hard contrast cut, slow forward drift, gentle forward velocity, pull out, descending into, camera locked, no camera movement, no zoom
Fragments to drop in
Slow push-in toward the subject, ending tight.Locked-off tripod, no camera movement at all.Smooth orbit around the subject, one quarter turn.Handheld with slight sway, documentary feel.
From production skills
camera dollies insubtle product reveal, camera slowly pulls back, ambient motionMOTION: Very slow push-in; the figure erodes grain by grain and particles drift sideways.Subtle camera push-in, no flicker, no extra text, no new objects.no cuts, no camera shake, slow steady motion only, locked exposure, no on-screen textcamera locked, no camera movement, no zoom, subject stays fully in frame, plain static background
For image-to-video the movement IS the prompt; for scrubbable footage a single unbroken move is mandatory and cuts are a defect. For sprite sheets the camera must be explicitly locked or frames become unusable.
Lens and focal length3.8% of video prompts
Perspective and compression.
A focal length is a shorthand a model understands: 85mm flatters a face, 24mm bends a room outward, macro turns a surface into a landscape.
Values
35mm, macro lens, anamorphic, 100mm, 50mm, 85mm, 28mm, wide-angle lens, 32mm, 50 mm, fisheye, telephoto, 40mm, 35 mm, 85 mm, 24mm
Fragments to drop in
Shot on an 85mm lens, compressed and flattering.Wide 24mm, close to the subject, edges stretching.Macro lens, the subject filling the frame at life size.
Depth of field11.8% of video prompts
What is sharp and what falls away.
Separates subject from background without changing the composition. The corpus reaches for shallow depth of field more than any other optical term.
Values
shallow depth of field, bokeh, rack focus, f/1.4, f/1.8, f/2.0, soft focus, f/2.8
Fragments to drop in
Shallow depth of field, background dissolving into bokeh.Deep focus, everything from foreground to horizon sharp.Rack focus from the foreground object to the face behind it.
Lighting7.4% of video prompts
Where the light comes from and how hard it is.
Sets mood more cheaply than any other choice. Naming a setup also stops the flat, evenly lit look models fall back on.
Values
soft lighting, rim lighting, rim light, silhouette, volumetric lighting, golden hour, volumetric light, soft light, side lighting, chiaroscuro, key light, low-key, practical lighting, backlit, high-key, side light, practical lights, backlight, god rays, neon lighting, blue hour, softbox
Fragments to drop in
Soft lighting from a large source, shadows barely there.Rim lighting from behind, edge of the subject glowing.Single hard key from the side, deep shadow on the far cheek.Golden hour sun, low and warm, long shadows.
Colour and grade17.0% of video prompts
The palette and contrast of the finished image.
Carries brand more than any other single element, and is easy to get consistent across a set of clips by reusing the same phrase.
Values
vibrant, pastel, film grain, black-and-white, natural color, high contrast, desaturated, crushed blacks, natural colour, low contrast
Also used in production skills, not counted
name the grade and the hexes, one visual grade, colour grading
Fragments to drop in
Vibrant saturated grade, colours pushed but not clipping.Muted pastel palette, low contrast, gentle.High contrast with crushed blacks and a cool cast.Fine 35mm film grain over the whole frame.
From production skills
Prompt "no cuts, no camera shake, slow steady motion only" and name the grade/hexes.Take only the visual render style and color grading of the input image(s); mix the styles if there is more than one image.
scroll-scrub.md forbids mixing models mid-chain because "their grain, color, and motion signatures create a visible seam even when position matches".
Medium and style38.1% of video prompts
What the output is pretending to be.
The most used word in the whole corpus is cinematic, which tells you models respond to a named medium. Being specific beats being evocative.
Values
cinematic, photorealistic, anime, documentary, 3d render, unreal engine, stop-motion, 35mm film, watercolor, vector, photoreal, claymation, line art, octane
Also used in production skills, not counted
flat 2D vector animation, bold clean outlines, solid vibrant flat fills, no shading, no gradients, hand-inked black marker on off-white paper, solid jet-black fills, thin white scratch highlights, marker grain, strictly monochrome, strict monochrome minimalism, black silhouettes on white void, high contrast, lots of negative space, hand-painted storybook gouache, soft textures, warm muted palette, visible brush strokes, non-photorealistic, illustrated, not a photo, no live-action, no realism
Fragments to drop in
Cinematic, shot on 35mm film.Photorealistic, indistinguishable from a real camera.3D render, clean studio lighting, product visualisation.Hand drawn 2D animation, visible line work.
From production skills
STYLE REFERENCE: Match the attached reference image EXACTLY. Replicate its look precisely: {STYLE tokens}. Every element below rendered in that identical style.non-photorealistic, illustrated, not a photo, no live-action, no realismPaste the same STYLE tokens into every block.
Write the render style, palette, line character and finish once, then repeat it byte-identically. scroll-scrub.md calls the equivalent a 'World grammar': one byte-identical style preamble, perspective, palette, light direction, surface finish and background behavior across all scene prompts.
Subject motion10.2% of video prompts
What moves inside the frame, and how fast.
Distinct from camera movement and frequently confused with it. Say what the subject does, or the model will hold it still and move the camera instead.
Values
reveal, float, drift, slow motion, glides, rotates, time-lapse, ripples, unfolds, hovers, cascade
Also used in production skills, not counted
the dancer spins, smoke rises slowly, ambient motion, fur ripples, stars orbit, clouds drift, train lights flicker, object assembling, exploded-view assembly, transformation or morph, a hero object emerging from darkness, smooth idle breathing cycle, subtle weight shift, full walk cycle in place, single sword slash, fast wind-up, sharp strike, follow-through, plain static background, studio black/charcoal, a soft gradient, the subject emerging from darkness, dark, seamless, low-detail, clean background, simple background
Fragments to drop in
The lid glides open on hidden runners, slow and even.Fabric billows in slow motion, every fold readable.Steam drifts upward and dissipates before it leaves the frame.
From production skills
SCENE: A lone black silhouette slowly dissolves at the edges, crumbling into fine drifting sand that scatters into the white emptiness.The character performs ONLY this action; nothing else happens.Use one clear action per block.A background the copy can survive. Dark, seamless, low-detail (studio black/charcoal, a soft gradient, the subject emerging from darkness) is the reliable choice.
"Exactly one action per video; any second verb in the prompt is a defect" (game-2d-animation.md).
Critical when HTML copy sits over the video: "A bright, busy, full-frame environment behind body copy is the single most common reason a beautiful film reads as unusable."
Mood21.9% of video prompts
The feeling the clip should leave.
A single adjective steers dozens of small decisions the model would otherwise make at random. Cheap to add, and it makes a set of clips feel related.
Values
playful, energetic, whimsical, calm, dramatic, minimal, elegant, uplifting, moody, ominous, ethereal, luxurious, serene, nostalgic, cosy, gritty
Fragments to drop in
Calm and unhurried throughout.Playful and energetic, quick and bright.Moody and restrained, close to darkness.
Audio bed11.8% of video prompts
The sound that plays under the picture.
Video prompts in this corpus routinely describe their own soundtrack. Naming instruments and what must not be heard is more reliable than naming a genre.
Values
percussion, bells, ukulele, xylophone, piano, drums, bass, marimba, strings, synth, guitar, choir, handclaps, pads, flute, synthesizer, synthetic, harp, synthase, kalimba
Also used in production skills, not counted
ambient SFX or music only, no voice, dialogue, or narration, Low sustained drone, a soft whisper of falling sand, generate_audio false, sound on/off
Fragments to drop in
Minimal deep electronic ambience with soft mechanical clicks.Warm ukulele and marimba, light percussion, no vocals.No music at all, only the room and the mechanism.
From production skills
AUDIO: {ambient SFX or music only, NO voice, dialogue, or narration}.AUDIO: Low sustained drone and a soft whisper of falling sand, no voice.Keep clip audio diegetic only. Characters never speak or lip-sync.
In the explainer pipeline narration is a separate audio job; clip audio must never contain speech.
Pacing and cuts
How the time is divided.
Without it a model gives you one unbroken drift. Naming beats is what turns a clip into a sequence.
Also used in production skills, not counted
smooth cinematic easing, slow, steady motion, constant speed, gentle ease only at the very start and end, locked exposure and white balance, no flicker, minimal motion blur, one continuous move, no hard cuts, START state ≠ END state, 0-1.5s , scene alive, no text yet, 1.5-3.5s , text entrance, 3.5-5s , settle, pops/slides/bounces in, pixel type snaps in block by block, chrome bubble letters inflate, condensed uppercase slams down, fades up, slides in with a soft bounce, Micro-motion only
Fragments to drop in
One continuous take, no cuts.Three beats: reveal, detail, hero shot.Cut away before the hand touches the product.
From production skills
One continuous move , no hard cuts.Slow, steady motion , constant speed, gentle ease only at the very start and end. The scroll supplies the pacing.Locked exposure and white balance, no flicker; minimal motion blur.Resolves at both ends: the first frame reads as the establishing shot, the last as the closing beauty state. START state ≠ END state, or the scrub has no payoff.Premium 5-second motion cover reveal, smooth cinematic easing.All elements ease precisely into their final positions and the video ends exactly on the provided end frame, holding still for the last moments.Then the title "[TITLE TEXT]" [ENTRANCE MOTION matched to its typography], the tagline fades up beneath it, and the small rounded pill button slides in with a soft bounce.
Matters whenever every frame may be held as a still (scroll-scrub) , exposure pumping shimmers and heavy blur smears.
Entrance motion should be matched to the typography of the text being revealed.
Resolution / quality tierfrom production skills
Seedance 2.5 caps at 720p so 1080p/4K work stays on Seedance 2.0; Grok Video 1.5 only 480p or 720p; Marketing Studio video 480p or 720p.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
480p, 720p, 1080p, 4k, basic, high, ultra, pro, std, standard, bitrate_mode standard|high
Fragments to drop in
--resolution 4k--resolution 720p--resolution 1080p --mode std
First/last frame anchoringfrom production skills
scroll-scrub.md warns that passing a storyboard as a start frame makes the page open on a static image; pass it in the generic reference role instead.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
--start-image, --end-image, start frame, end frame, start = end, one-shot, loop, end-frame reveal, generic image/reference role , not a start-frame role
Fragments to drop in
`--start-image` anchors the first frame. Prompt describes motion.Don't redescribe the static frame , model already has it.Looping actions pass the SAME image as both start frame and end frame. One-shot actions (attack, death, hit, cast) pass only the start frame , forcing them back to the start pose ruins the action.Pass the finished cover as the END frame and describe the buildup
Ad format / mode (branded video)from production skills
Hook text is prepended to the user's prompt and does not replace it; COOKBOOK.md notes "Hooks (the prompt on --prompt) matter more than mode for performance."
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
ugc, ugc_how_to, ugc_unboxing, product_showcase, product_review, tv_spot, wild_card, ugc_virtual_try_on, virtual_try_on, hook, setting, ad reference
Fragments to drop in
Looks like a real person filmed on phonePolished broadcast commercialShow the product itself, less presenterPresenter giving an opinion
Edit instruction (video editing workflows)from production skills
draw_to_video takes a source video, an edited/sketched frame, the timestamp for that frame and a short edit instruction.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
make the jacket red, sketch, timestamp
Fragments to drop in
--prompt "make the jacket red"
The order to write them in
A prompt is read start to finish, and earlier clauses set the frame the later ones are interpreted against. This order puts the decisions that cannot be undone first.
- Duration
- Aspect ratio
- Medium and style
- Shot size
- Camera angle
- Camera movement
- Lens and focal length
- Depth of field
- Lighting
- Colour and grade
- Subject motion
- Mood
- Audio bed
- Pacing and cuts
- Negatives
How production prompts are laid out
Two dominant shapes. (1) The explainer block, a labelled five-line form in fixed order: `Block N` / `STYLE REFERENCE:` / `SCENE:` / `MOTION:` / `AUDIO:` / `NEGATIVE:` , the STYLE REFERENCE and NEGATIVE lines are pasted byte-identically into every block, only SCENE and MOTION change, and each block holds exactly one clear action (higgsfield-video-explainer/references/prompts.md). (2) The image-to-video single paragraph, where the start image already carries the subject so the prompt is motion only , "Don't redescribe the static frame , model already has it" , optionally followed by a mandatory negative block in three fixed parts: camera lock, prop and action inertia, facing lock (higgsfield-generate/references/prompt-engineering.md, higgsfield-websites/references/game-2d-animation.md). Longer reveal prompts run beat by beat in time order: start state and context motion to text/element entrance to settle, closing with an anchor sentence that ties the last frame to the supplied end frame (higgsfield-websites/references/cover-animator.md). Footage-level constraints (one continuous move, centre-safe hero, survivable background, locked exposure, resolving ends) are restated in the prompt rather than left to the model (higgsfield-websites/references/scroll-scrub.md).
What to rule out
The most common element in the whole corpus. Put these last, in one sentence, and only include what is plausible for the shot: a long list of irrelevant prohibitions wastes the model's attention.
Text and branding
text, logos, subtitles, watermark, watermarks, captions, on-screen text, text on screen
Anatomy and faces
extra limbs, distorted faces, extra characters, hands, extra fingers, extra animals
Content and safety
violence, scary imagery, scary elements, weapons
Audio
music, narration
Other
dialogue, copyrighted characters, talking, cuts, horror, humans, clothes, flickering, frightening imagery, morphing, camera shake, zoom
Fragment to drop in
No text, logos, subtitles, watermark, watermarks, captions, extra limbs, distorted faces, extra characters, hands, extra fingers, extra animals, violence, scary imagery, scary elements, weapons, music, narration, dialogue, copyrighted characters, talking, cuts, horror, humans.
Exclusion lines used in production
Copied verbatim from working prompt skills. Useful as whole lines rather than as a vocabulary.
color drift, photorealism, 3D render, lip-sync, captions, on-screen text, logos, watermarkcolor, gray midtones, photorealism, 3D render, lip-sync, captions, on-screen text, logos, watermarknon-photorealistic, illustrated, not a photo, no live-action, no realismno voice, dialogue, or narrationcamera locked, no camera movement, no zoom, subject stays fully in frame, plain static backgroundthe character performs ONLY this action, nothing else happensthe <prop> stays inert and is never used , no firing, no muzzle flash, no swinging, no raising itthe subject keeps facing the SAME direction for the entire video , never turns around, never rotates toward or away from the camera, no head turns past the shoulderno cuts, no camera shake, slow steady motion only, locked exposure, no on-screen textSubtle camera push-in, no flicker, no extra text, no new objects.No on-screen text, logos, or watermarks , all type is HTML over the video.Avoid: cuts, fast pans, handheld shake, busy bright backgrounds, and colour/exposure flicker.The solid color frame, corner dots and capsule shape stay perfectly static at all times.no lip-sync, captions, text, logos, or watermark
Image prompts
Measured across 3,597 image prompts, mean length 1063 characters.
Download the image sheet as a PDF, one page, or take all three.
Aspect ratio16.5% of image prompts
Frame shape.
Choose it first. Composition instructions that assume a landscape frame fall apart in a vertical one.
Values
16:9, 9:16, 1:1, 4:5, 21:9, 3:2, 4:3
Fragments to drop in
Square 1:1.Vertical 4:5 for a feed post.Wide 16:9 with the subject offset left.
Composition and framing31.3% of image prompts
Where things sit in the frame.
The most common failure in the corpus is a subject cropped or crowded. Explicit composition instructions, including what must stay fully visible, fix it.
Values
close-up, wide shot, medium shot, extreme close-up, two-shot, closeup, over-the-shoulder, establishing shot, macro shot, full shot
Also used in production skills, not counted
one power-third hero, clear scale hierarchy, depth, strong subject/background separation, centered, left-anchored, asymmetric split, one focal visual, decisive negative space, generous negative space, reserve the block side, soft open negative space upper-left, single unified frame, split-screen, diagonal divide, halves/panels, extreme scale contrast, power third, centered isolated presentation
Fragments to drop in
Centred on a single vertical axis, symmetrical.Rule of thirds, subject on the right third, negative space left.Everything fully inside the frame, nothing cropped at the edges.Generous negative space above the subject for a headline.
From production skills
Composition: one power-third hero, clear scale hierarchy, depth, and strong subject/background separation.Scale + placement: "waist-up, subject fills two thirds of the frame, positioned right-of-center" , and reserve the block side: "soft open negative space upper-left on a warm white wall". Negative space = a real surface (wall, sky, backdrop), NOT a black void.single unified frame , no split-screen, no diagonal divide, everything blends smoothly and organically across the same continuous shot
A split layout is used only when the user asks for split, before/after, versus or side by side; "A topical phrase such as `X vs Y` does not itself require a split."
Shot size11.4% of image prompts
How close the camera is.
Same lever as in video and just as reliable. It also controls how much detail the model has to invent.
Values
close-up, wide shot, medium shot, extreme close-up, two-shot, closeup, over-the-shoulder, establishing shot, macro shot, full shot
Also used in production skills, not counted
chest-up, medium-close, about 40-60% of the frame, waist-up, subject fills two thirds of the frame, 55-75 % of the frame, Full body, Three-quarter, Waist up, Closeup on product area, head shot, head-and-shoulders, full body, large, foreground-dominant, rendered LARGE
Fragments to drop in
Extreme close-up, texture filling the frame.Full body shot, head to feet, room above and below.Head and shoulders portrait.
From production skills
Subjects: large, foreground-dominant, chest-up or medium-close, filling about 40-60% of the frame. End with `All faces crisply sharp as the anchors of the shot.`waist-up, subject fills two thirds of the frame, positioned right-of-center
app-cover.md kills a candidate whose subject is under ~50 % of frame. Framing is an explicit interview option in product-photoshoot: `[Full body / Three-quarter / Waist up / Closeup on product area]`.
Camera angle4.6% of image prompts
Viewing position.
Three-quarter is the most used angle in the corpus for a reason: it shows form without the stiffness of a straight-on view.
Values
three-quarter, profile, eye-level, low angle, overhead, high angle, top-down, bird's-eye, dutch angle
Also used in production skills, not counted
low, low wide angle, side-view, flat frontal, three-quarter isometric view, front, 3/4 left, 3/4 right, slight up/down, wide establishing view, multi_angle
Fragments to drop in
Three-quarter view, turned slightly off axis.Straight-on profile, subject facing frame left.Shot from directly above, flat lay.
From production skills
angle (low, overhead)consistent side-view perspective across all assetsthree-quarter isometric view
game-stylization.md makes the perspective word genre-bound: platformer/runner to side-view; top-down to top-down; puzzle/clicker to flat frontal. For image-to-3D concept images the perspective word is replaced by three-quarter isometric view because it shows top plus two sides.
Lens and focal length5.0% of image prompts
Perspective and compression.
Also the cheapest way to signal photographic realism, because a real photograph was taken at some focal length.
Values
35mm, 85mm, 50mm, 110 mm, 16mm, anamorphic, 24mm, 70mm, 48mm, fisheye, 600mm, 450mm, 40mm, 297 mm, 650mm, 20mm, 135mm, 100mm, wide-angle lens, 105 mm, 220mm, 110mm
Also used in production skills, not counted
low wide angle, close microphone (audio analogue), 35mm film
Fragments to drop in
85mm portrait lens, natural compression.35mm, mild environmental context around the subject.Macro at life size, the surface reading as landscape.
From production skills
**Camera**: lens (35mm, 85mm), angle (low, overhead), motion (dolly in, tracking shot)bright soft key with punchy shadows, low wide angle, commercial editorial grade
app-cover.md notes the lens block is replaced by style words such as "glossy 3D render / claymation diorama / painterly still" on the graphic route. Vector/logo prompts must never contain lens language.
Depth of field6.0% of image prompts
Focus falloff.
Separates a subject from clutter you did not describe, which is most backgrounds.
Values
shallow depth of field, bokeh, f/1.4, f/7.1, soft focus, f/2.8, deep focus, f/2.0, f/4, f/2, f/1.8, f/5.6
Fragments to drop in
Shallow depth of field at f/1.8, background soft.Everything sharp front to back, deep focus.Only the eyes critically sharp.
Lighting6.8% of image prompts
Direction, quality and source of light.
Silhouette and rim light are the two most used named setups in the corpus, both of which the model will never choose on its own.
Values
silhouette, rim light, soft light, soft lighting, backlit, key light, golden hour, high-key, rim lighting, backlight, softbox, low-key, volumetric light, bounce light, practical lighting, three-point lighting, fill light, volumetric lighting, practical lights, chiaroscuro, side lighting, god rays
Also used in production skills, not counted
neon glow, moody backlight, strong key light, soft dreamy fill, back light, hair light, bright soft key with punchy shadows, flat ambient light, flat even lighting, soft studio reflections, soft contact shadow, warm hard light, restrained lighting, no shadows, no vignette, soft vignette and edge falloff, warm amber torchlight
Fragments to drop in
Large softbox front left, soft shadows falling right.Backlit into silhouette, subject dark against a bright field.High-key, white on white, almost no shadow.Single hard light, sharp shadow edges.
From production skills
**Lighting**: rim light, neon glow, moody backlightsignature YouTube thumbnail lighting rig , strong key light sculpting the face, soft dreamy fill lifting shadows, and defined back light plus hair light tracing a clean bright rim around hair, shoulders and silhouette.bright soft key with punchy shadows, low wide angle, commercial editorial gradeFlat even lighting, no shadows, no vignette.flat even lighting, no perspective, no objects, photographed-flat appearance
In the thumbnail rig "Only the rim may use a colored accent." Tileable textures need flat lighting so the tile has no focal light. Brandkit mockups explicitly ban "Excessive bloom, haze, depth of field, or cinematic lighting".
Colour and palette10.4% of image prompts
Palette, saturation, contrast.
Keeps a set consistent. Reuse the identical phrase across a batch rather than paraphrasing it each time.
Values
vibrant, pastel, film grain, black-and-white, desaturated, high contrast, natural color, natural colour, monochrome, low contrast, sepia
Also used in production skills, not counted
vivid, bright, glossy, poster-punchy, deep blacks, crisp highlights, rich saturated colors, cohesive as one image, commercial editorial grade, duotone, graded atmospheric photo, monochrome / force grayscale, muted cinematic movie still (banned), name 2-3 colors, exact hex values, locked primary tone, locked accent, locked background, background, text, primary, accent, support, environment in deep violet stone with charcoal shadows, hero in warm coral-cream tones contrasting the surroundings, hazards and pickups marked with acid-green glow, deep-navy #243A5E, signal-red #D23B2E, olive #7A7D3C, sky-blue #A9CFF4, peach #F7DDB9, cream #F1EEE6, dusty-rose #E7B7B0, warm-gray #D6D3CE, acid lime #D4FF3F, #FF00FF, #00FF00, #0000FF
Fragments to drop in
Vibrant, fully saturated, clean whites.Muted pastel palette, soft contrast.Black and white, deep blacks, bright highlights.
From production skills
Grade: vivid, bright, glossy, poster-punchy, deep blacks, crisp highlights, rich saturated colors, cohesive as one image. Restrain it only for an explicit calm/premium/muted brief.for a monochrome look force grayscale in the prompt AND on exportPalette, committed and saturated: name 2-3 colors. Avoid lime/acid yellow-green.Use role language such as "locked primary tone," "locked accent," and "locked background."Exact hex values are passed separately through the Recraft request's `colors` and `background_color` params. Never put hex, RGB, Pantone, or other color codes in the prompt.
asset-system.md: a mixed-grade asset kit reads cheaper than no assets; re-generate off-grade pieces "with the hexes named harder".
Two opposite conventions co-exist: website/game prompts NAME the hexes in the prompt; Recraft logo prompts must never contain a colour code and pass hexes as params. Logo palettes use one, two or three colours , "three is a maximum, not a target or default". Game palettes are specified BY ROLE, not as one global gamma.
Medium and style26.6% of image prompts
Photograph, render, illustration or something else.
Ambiguity here produces the uncanny middle ground that looks like neither a photo nor an illustration.
Values
cinematic, photorealistic, documentary, anime, 3d render, line art, watercolor, vector, 16mm, 35mm film, photoreal, kodak vision3, isometric, flat illustration, claymation
Also used in production skills, not counted
oil painting, photograph, soft cel shading, glossy 3D render, claymation diorama, painterly still, flat vector cartoon with soft gradients, chunky pixel art, 32x32 grid feel, soft hand-painted gouache, low-poly faceted, smooth rounded blobby, flat_vector, monoline, vector_gradient, hand_drawn_vector
Fragments to drop in
Photorealistic editorial photograph.Clean 3D product render on a seamless backdrop.Flat vector illustration, limited palette, no gradients.Loose watercolour with visible paper texture.
From production skills
**Style/medium**: oil painting, watercolor, photograph, anime, 3D rendertransform into anime style, vibrant colors, soft cel shading
Brandkit treatment vocabulary is a closed set with defined meanings: flat_vector = solid fills, clean SVG paths, no surface effects; monoline = uniform stroke weight, rounded caps, no fills; vector_gradient = vector-safe linear, radial or duotone gradient with the locked stop count; hand_drawn_vector = intentional stroke variation. "Dimensional/3D treatment is forbidden."
Materials and finish8.5% of image prompts
What surfaces are made of.
Materials are where a render either convinces or does not. Naming the finish is more effective than naming the object.
Values
glossy, matte, ceramic, concrete, leather, velvet, satin, brass, linen, frosted, marble, chrome, walnut, oak
Also used in production skills, not counted
print, emboss, debossing, foil, laser engraving, one-color printing, screen printing, embroidery, spot-color, folds, grain, perspective, occlusion, reflections, scale, manufacturing limits, believable materials, light kraft paper, natural cardboard, pale fabric, curved packaging, surface wear
Fragments to drop in
Brushed metal with a satin finish, no mirror reflections.Matte ceramic, slightly uneven glaze.Smoked glass over dark walnut.
From production skills
[PHYSICAL APPLICATION] Render the logo using <credible print/emboss/foil/engraving/ink behavior>. Respect folds, grain, perspective, occlusion, reflections, scale, and manufacturing limits.Prefer believable materials, restrained lighting, purposeful negative space, specific environments, and one focal branded application.
Use the generator only when branding must interact with fabric folds, curved packaging, embossing/debossing, foil, print texture, reflections, surface wear, occlusion or realistic perspective; otherwise place the logo deterministically.
Background36.9% of image prompts
What sits behind the subject.
Named in 37% of image prompts, the highest of any positive aspect. An undescribed background is where models put clutter, text and logos.
Also used in production skills, not counted
vivid high-contrast color field, soft vignette and edge falloff, 2-4 supporting story props/layers, credible setting, specific environments, no readable screens, clean background, solid uniform bright <KEY COLOR> background, pure flat white, no shadow, no ground, Studio clean, Outdoor natural, Street style, Editorial, Home cozy
Fragments to drop in
Seamless mid-grey studio backdrop, no horizon line.Plain white, no props, no shadow on the background.Softly blurred interior, unreadable, no recognisable objects.
From production skills
Background: vivid high-contrast color field or environment, soft vignette and edge falloff; do not divide it unless a split layout was explicitly requested.**World + props (story density)**: name 2-4 supporting elements that fill the frame around the subject. Add "no readable screens" whenever screens/props could sprout text.on a solid uniform bright <KEY COLOR> background, no shadows cast on the background, no ground plane, nothing cropped at the edges
app-cover.md: "One person on an empty backdrop reads as AI slop even when bright , stage a world, not a portrait." For image-to-3D the background is instead "pure flat white, no shadow, no ground".
Mood9.9% of image prompts
The feeling.
One adjective, placed early, quietly resolves dozens of choices in a consistent direction.
Values
elegant, playful, minimal, dramatic, calm, energetic, whimsical, moody, luxurious, nostalgic, uplifting, ethereal, cosy, gritty
Also used in production skills, not counted
shock, hype, rage, awe, laugh, fear, smug, charisma, confusion, determination, disgust, neutral, smiling, talking
Fragments to drop in
Elegant and restrained.Playful and bright.Moody, close to darkness.
From production skills
If the user gives an emotion count without names, use this ladder: shock, hype, rage, awe, laugh, fear, smug, charisma, confusion, determination, disgust.
Emotion is one of the allowed surgical-edit scopes; expression changes are made as an edit, not a full regeneration.
Text handling18.4% of image prompts
Whether any lettering appears.
Models add text nobody asked for. 18% of image prompts in the corpus explicitly forbid it, and that is the single highest value instruction you can add.
Also used in production skills, not counted
No text, no readable UI labels, no watermark, BAKED UI, reading exactly, massive, ultra-legible sans-serif with a clean outline/glow treatment, 2-4 words max, ALL CAPS, no extra text
Fragments to drop in
No text, no lettering, no numbers anywhere in the image.No logos, watermarks, captions, borders or frames.Leave the upper third clear for a headline added later.
From production skills
Text: default `No text, no readable UI labels, no watermark.`TEXT: bold thumbnail headline text baked into the image, reading exactly "<TEXT>" , massive, ultra-legible sans-serif with a clean outline/glow treatment, placed where it never covers the subject's face. No other text, no watermark.**The image model NEVER renders text. Not the title, not the wordmark, not labels, not UI. It renders ONLY the scene.**Never bake the final logo or exact copy into a generated image.
The dominant convention across the repo is: keep text out of the generation and set type deterministically afterwards. GPT Image 2 is the one model reached for when readable text must be inside the image.
Subject + setting + style (the opening clause)from production skills
"Higgsfield models reward concrete, sensory prompts." app-cover.md wants the product's core verb staged as a physical, photographable moment.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
a red fox curled in a snowy pine forest, golden hour, cinematic, subject + action, concrete and physical
Fragments to drop in
a red fox curled in a snowy pine forest, golden hour, cinematica film editor in a bold red bomber jacket mid-leap, cutting a giant arc of 35mm film with oversized chrome scissors
Typography named in the prompt (when text is generated)from production skills
text-overlay-bake.md gives the MrBeast headline stack as concrete values: font Anton, ALL CAPS, stroke 8-14% of cap size drawn under the fill, hard dark offset shadow, colour white / yellow to orange to red gradient / acid lime #D4FF3F, tracking -0.01...-0.02em, line-height 0.9, cap height 12-18% of frame height.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
Anton, Anton SC, Bebas Neue, Oswald, Archivo Black, Montserrat, Poppins, Roboto Condensed, Inter, Barlow Condensed, Playfair Display, Cormorant Garamond, DM Serif Display, Fraunces, Sacramento, patrick, caveat, marker, anton, clean grotesk (Satoshi-like), refined grotesk (Neue-Montreal-like), expressive display (Cabinet/Clash-like), compressed statement (Monument-like), editorial serif + sans pairing, Swiss rational sans with hard hierarchy
Fragments to drop in
The prompt must state: Exact literal copy; Display/body font family names and which text uses eachName the exact display/body families in the prompt; never infer typography from the logo or palette.
Identity lock / subject fidelityfrom production skills
Up to three referenced identities; face photos are passed first in character order, then the logo, and a manifest line must open the prompt when two or more references are attached.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
IDENTITY LOCK, photographic identity match, same bone structure, eye shape, nose, lips, jawline, skin tone, hairline and hair texture, Do not beautify, average, or restyle the face, IMAGE REFERENCES manifest
Fragments to drop in
CHARACTER N: the person from attached face reference #K , IDENTITY LOCK: reproduce this exact person with a photographic identity match , same bone structure, eye shape, nose, lips, jawline, skin tone, hairline and hair texture. Do not beautify, average, or restyle the face. Expression: <emotion phrase>.IMAGE REFERENCES: image 1 = CHARACTER 1 face reference; image 2 = brand logo.All faces crisply sharp as the anchors of the shot.
Aesthetic / seasonal presetfrom production skills
These are the labelled options the skill offers for restyle and style/mood questions; they become the short intent line handed to the backend enhancer.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
Clean girl, Cottagecore, Quiet luxury, Dark academia, Y2K, Christmas, Valentine's, Halloween, Black Friday, None, Clean studio, Lifestyle, Conceptual, With a model
Fragments to drop in
Christmas version, quiet-luxury aestheticvertical pin for my candle brand, cottagecore mood
Product-photo genre / modefrom production skills
For both skills the user-facing prompt is a short intent line only , "Backend assembles the final prompt; never freehand it." `--count 3` makes the backend vary preset, lighting, angle and palette across variants.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
product_shot, lifestyle_scene, closeup_product_with_person, moodboard_pin, hero_banner, social_carousel, ad_creative_pack, virtual_model_tryout, conceptual_product, restyle, levitating, floating, splash, frozen motion, surreal, CGI, sculptural, main_image, infographic, multi_angle, detail_shot, lifestyle, whats_in_box
Fragments to drop in
bottle of cold-brew on a sunlit kitchen counter, IG feedsparkling peach lemonade can for marketplace listingpremium skincare serum, clean clinical marketplace visual systemBold hero shot on marble
Logo / vector mark specificationfrom production skills
One central visual idea per mark, stated in one clause; the exact phrase "no text" must appear in the constraint tail; mark_type routing is evidence-based, defaulting to abstract, never wordmark or combination mark.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
lettermark_monogram, pictorial, abstract, mascot, emblem, one dominant unified silhouette, negative-space device, motif treatment, unexpected locked color-role pairing, symmetry or intentional asymmetry, internal detail limit, small-size scalability, centered isolated presentation, minimal anchor points, SVG-friendly
Fragments to drop in
Flat vector design, clean lines, no shadows, no texture, no text. Clean editable vector paths, SVG-friendly, minimal anchor points.Monoline vector design, uniform stroke weight, rounded line caps, no fills, no shadows, no texture, no text. Clean editable vector paths, SVG-friendly, minimal anchor points.Flat vector design with the specified locked vector gradient, no shadows, no texture, no text. Clean editable vector paths, SVG-friendly, minimal anchor points.Intentional hand-drawn vector strokes, approved surface variation, clear scalable silhouette, no shadows, no text.Describe the distinctive element concretely enough that another designer could sketch its structure without guessing.
Logo placement on a mockup surfacefrom production skills
Full-colour marks go on smooth surfaces; black monochrome on light kraft paper, natural cardboard, pale fabric, stamps, dark-ink screen printing, engraving masks and light uncoated stock; white monochrome on dark paper/boxes/fabric, reverse marks, light-ink screen printing and dark signage.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
Target surface, Position, Scale, Orientation, Color, horizontally centered, upper third, optical center aligned to panel, clear-space margin, align to <panel edge/seam/baseline>, follow surface perspective
Fragments to drop in
[PLACEMENT LOCK] Target surface: <exact object panel/face/material>. Position: <exact alignment and location, e.g. horizontally centered, upper third, optical center aligned to panel>. Scale: logo occupies <specific proportion> of the target surface while keeping <specific clear-space margin>. Orientation: align to <panel edge/seam/baseline>; follow surface perspective without changing logo proportions. Color: use the supplied <full-color/black/white> variant exactly. State why it contrasts correctly with the material/background.[AUTHORITATIVE LOGO] <ImageN> is the exact approved <full-color/black/white> logo. Preserve its spelling, silhouette, geometry, proportions, internal negative space, and exact color. Do not redraw, simplify, crop, stretch, outline, or add effects."Place the logo on the bag/box" is insufficient.
Seamless / tileable behaviourfrom production skills
{MATERIAL} in the texture templates is 3-6 concrete words from the reference, e.g. "hand-painted cracked dry earth with small grey stones".
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
seamless tileable game texture tile, uniform pattern density, perfectly seamless edges that wrap horizontally and vertically, no border, no vignette, no single focal object, photographed-flat appearance, seamless repeating pattern
Fragments to drop in
seamless tileable game texture tile of <description>, uniform pattern density,, perfectly seamless edges that wrap horizontally and vertically, no border, no vignette, flat even lighting, no single focal objectseamless tileable surface texture of <description>,, perfectly seamless edges, flat even lighting, no perspective, no objects, photographed-flat appearanceThe repeating units must be aligned so rows continue across the tile borders: unit boundaries at the left edge continue into the right edge and across top/bottom.Reduce the visible periodic repetition of distinctive features so the pattern feels organic.
Key colour for transparencyfrom production skills
Models cannot output alpha, so asking for "transparent background" produces a fake checkerboard. Default magenta; switch to green if the subject's own colours are near pink/magenta/purple; blue if both are taken. The chosen colour must appear in both the prompt suffix and the keying script.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
bright magenta #FF00FF, bright green #00FF00, bright blue #0000FF
Fragments to drop in
on a solid uniform bright <KEY COLOR> background, crisp edges, no drop shadow outside the elementClean uniform <key color> background.
Reference-image role and what it controlsfrom production skills
For image-to-image the prompt should describe what changes, not redescribe the input: bad "a man with brown hair in a leather jacket holding coffee, made into anime"; good "transform into anime style, vibrant colors, soft cel shading".
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
style donors only, Image0, Image1, Image2, official logo, approved base scene, style-reference, representative-visual, grade/energy only
Fragments to drop in
Make an Animated Explainer. Take only the visual render style and color grading of the input image(s); mix the styles if there is more than one image. Never use the characters, inscriptions, etc. from the input image(s) unless the instructions below ask you to. Use only the render style, and follow the user's instructions below:Use references as style donors only. Never copy their people, text, logos, or objects.Label references by role in prompts: `Image 1: official logo`, `Image 2: approved base scene`.State what each reference controls and what it must not control.
Design-board direction vocabulary (web/layout imagery)from production skills
Commit to one option per category before prompting and hold it across all boards; one board per section, landscape 16:9 or 3:2, never one tall full-page image.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
Pristine Light (paper/cream/off-white, dark ink), Deep Dark (charcoal/graphite), Bold Studio Solid (oxblood, royal blue, forest, vermilion, emerald fields), Quiet Premium Neutral (bone, sand, taupe, stone, smoke), technical grid/dot field, solid with soft ambient depth, full-bleed cinematic imagery, tactile paper/material texture, cinematic centered minimalist, asymmetric split, floating polaroid scatter, inline typography behemoth, editorial offset, massive image-first with restrained text, modular bento rhythm, alternating editorial blocks, poster-stacked storytelling, gallery-led cadence, Swiss grid discipline, asymmetric premium flow, artifact/collectible, journey/waypoints, tool/precision instrument, living system/garden, stage/spotlight, archive/dossier
Fragments to drop in
"website design mockup, desktop landing page section, [SECTION ROLE], [theme paradigm + exact palette words], [typography character] typography, [hero architecture / composition anchor], [background mode], [narrative spine motif], professional layout, clear hierarchy and spacing, award-winning web design" , plus: "no watermark, no browser chrome".Name real content in the prompt (the actual headline wording you plan) so type sits believably.
Reference photos for identity trainingfrom production skills
5-20 photos, 8-12 the sweet spot, clear face with eyes visible, single person per photo, sharp, ≥1024×1024 ideal.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
front, 3/4 left, 3/4 right, slight up/down, indoor, outdoor, soft, harsh, neutral, smiling, talking, head shot, head-and-shoulders, full body
Fragments to drop in
Higher variety = better identity capture.
Image-to-3D input framingfrom production skills
One figure per input image , a 4-view character sheet reconstructs as "a statuette of 4 fused figures".
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
t-pose, a-pose, Full body in frame, nothing cut off, Clean/plain background, Limbs visually separated from the torso, One character, no props overlapping the silhouette, lowpoly, three-quarter isometric view
Fragments to drop in
Full body in frame, nothing cut off.Clean/plain background; no shadows, text, watermarks, logos.Limbs visually separated from the torso (arms not pressed to the body) , fused silhouettes produce fused geometry and break auto-rig.
The order to write them in
A prompt is read start to finish, and earlier clauses set the frame the later ones are interpreted against. This order puts the decisions that cannot be undone first.
- Aspect ratio
- Medium and style
- Shot size
- Camera angle
- Composition and framing
- Lens and focal length
- Depth of field
- Lighting
- Colour and palette
- Materials and finish
- Background
- Mood
- Text handling
- Negatives
How production prompts are laid out
The base order is subject + setting + style to camera (lens, angle, motion) to lighting to style/medium, kept under about 200 tokens because "Models distort with very long prompts" (higgsfield-generate/references/prompt-engineering.md). Specialised skills expand this into numbered block contracts assembled in a fixed order. The thumbnail contract has eleven blocks: frame to scene brief to text to subjects to key elements to logo to location to composition to background to lighting on people to grade, with per-character IDENTITY LOCK paragraphs and an opening IMAGE REFERENCES manifest when two or more references are attached (higgsfield-youtube-thumbnail/SKILL.md). The cover doctrine has six: subject + action to world + props to scale + placement to palette to light + lens to and "Always end with: no text, no letters, no logos, no captions, no UI" (higgsfield-websites/references/app-cover.md). Game assets concatenate exactly four parts in order: `<kind template> + <asset description> + <STYLE FORMULA byte-identical> + <kind suffix>` (higgsfield-websites/references/game-stylization.md). Recraft logo prompts are "exactly one continuous enhanced Recraft prompt string per candidate , no bullet points inside the prompt" ordered mark type to central subject/mechanism to shape logic to style register to palette behavior to composition to constraint tail, and "Every clause must materially affect the drawing" (higgsfield-brandkit/references/logo-prompt-enhancer.md). Brandkit mockups use bracketed labelled sections , [CREATE ONE FINISHED BRANDED MOCKUP], [AUTHORITATIVE LOGO], [PLACEMENT LOCK], [PHYSICAL APPLICATION] , plus a compact [BRAND LOCK , DO NOT DEVIATE] block repeated verbatim across related assets (higgsfield-brandkit/references/mockups.md, higgsfield-brandkit/references/brand-lock.md).
What to rule out
The most common element in the whole corpus. Put these last, in one sentence, and only include what is plausible for the shot: a long list of irrelevant prohibitions wastes the model's attention.
Text and branding
text, logos, captions, watermark, watermarks, logo, labels
Anatomy and faces
hands, arms, extra people, extra characters
Frame artefacts
frame, border, mockup
Content and safety
scary elements, weapons prominently displayed
Style and clutter
people, props
Other
overlays, slogan, letters, sunglasses, dramatic pose, helmet, police department names, cinematic color grading, character redesign, new clothing, new pose, perspective
Fragment to drop in
No text, logos, captions, watermark, watermarks, logo, hands, arms, extra people, extra characters, frame, border, mockup, scary elements, weapons prominently displayed, people, props, overlays, slogan, letters, sunglasses, dramatic pose, helmet.
Exclusion lines used in production
Copied verbatim from working prompt skills. Useful as whole lines rather than as a vocabulary.
no text, no letters, no logos, no captions, no UINo text, no logos, no watermark.No text, no readable UI labels, no watermark.no text, no logos, no watermarkno extra text, no watermarkno watermark, no browser chromeno shadows, no texture, no textno fills, no shadows, no texture, no textno border, no vignetteno single focal objectno characters, no UI elementsno drop shadow outside the elementno perspective, no objectsno shadows cast on the background, no ground plane, nothing cropped at the edges
Audio prompts
Measured across 163 audio prompts, mean length 1253 characters.
Audio is the thinnest part of this corpus at 163 prompts against 4,300 for video, so treat the audio counts as indicative. Video prompts frequently describe their own audio bed, so much of the audio vocabulary was learned there.
Download the audio sheet as a PDF, one page, or take all three.
Duration38.0% of audio prompts
Length of the piece.
A generator will otherwise give you a loop of arbitrary length that has to be cut, which is where the musical ending gets lost.
Values
15 seconds, 30 seconds, 60 seconds, 90 seconds, 120 seconds
Also used in production skills, not counted
--duration 30, --duration 12, --duration 4, --duration 2
Fragments to drop in
90-second bed.30 seconds with a clean ending.15-second sting.
From production skills
mirelo_text_to_audio --prompt "glass breaking in a large hall" --duration 4sonilo_music --prompt "cinematic synthwave track" --duration 12
Sonilo Music and Mirelo Text to Audio both require an explicit --duration; Seed Audio does not.
Role
What the audio is for.
A bed under narration and a standalone track need opposite decisions about melody and midrange, so state the job before the style.
Also used in production skills, not counted
no music, no voice, no ambience, dry studio recording, close microphone, short isolated, single, sound effects, ambience, foley, impacts, transitions, environmental sounds
Fragments to drop in
Instrumental music bed to sit under narration.Standalone track, no voice over it.Short sting to close a scene.
From production skills
short isolated sword impact, dry studio recording, no musicsingle heavy wooden door slam, close microphone, no ambienceglass breaking in a large hallcinematic rain ambience with distant thunderDescribe one sound per SFX prompt. Do not request a complete mixed scene.State `no music`, `no voice`, or `no ambience` when isolation matters.
This is the single richest audio prompt guidance in the repo. One sound per prompt is the hard rule.
Tempo38.7% of audio prompts
Speed, ideally as a number.
39% of audio prompts in the corpus give a BPM. A number is obeyed far more reliably than a word like upbeat.
Values
70 BPM, 90 BPM, 100 BPM, 116 BPM, 128 BPM
Fragments to drop in
Buoyant mid-tempo around 116 BPM.Slow, roughly 70 BPM.Driving, 128 BPM, steady four on the floor.
Key and tonality15.3% of audio prompts
Major or minor, and the key if it matters.
The cheapest emotional lever in audio. Major reads as safe and bright, minor as unresolved, and models respond to both.
Values
major key, minor key, modal, pentatonic
Fragments to drop in
Bright major key.Minor key, unresolved.Warm, clean major-key resolution in the final eight seconds.
Instrumentation71.2% of audio prompts
Which instruments play.
Named in 71% of audio prompts, the most used aspect of any type. Listing instruments works better than naming a genre, which models interpret loosely.
Values
percussion, ukulele, piano, bass, handclaps, marimba, bells, strings, guitar, xylophone, drums, harp, synthesizer, pads, 808, synth, flute, choir, synthesized
Also used in production skills, not counted
instrumental, no vocals, cozy forest loop, warm marimba and soft strings, cinematic synthwave track, backing tracks, instrumental beds, jingles, musical moods
Fragments to drop in
Marimba, ukulele, xylophone, light percussion and soft handclaps.Solo piano with a distant string pad.Deep electronic ambience with soft mechanical clicks.
From production skills
instrumental cozy forest loop, warm marimba and soft strings, no vocalscinematic synthwave trackMusic prompts specify mood, tempo, instrumentation, and `instrumental/no vocals` when lyrics are unwanted.
Tempo is named as a required aspect but the repo gives no BPM values or key/scale vocabulary anywhere , this is genuinely thin.
Mood20.2% of audio prompts
The feeling.
Pairs with tempo and key. Together those three settle most of the composition.
Values
warm, playful, cheerful, buoyant, wholesome, ambient, uplifting, driving, triumphant, minimal, dreamy, suspenseful
Also used in production skills, not counted
material, era, energy, tonal language
Fragments to drop in
Cheerful, wholesome and playful.Calm and spacious.Suspenseful and restrained.
From production skills
Preserve the game's STYLE FORMULA conceptually through material, era, energy, and tonal language even though audio is not visual.
The visual style contract is carried into audio prompts by translation, not by pasting the formula.
Arrangement and dynamics10.4% of audio prompts
How the piece develops.
Without this you get four bars looped flat. Naming a build and an ending is what makes it usable against picture.
Values
airy, midrange, headroom, no abrupt hits, narration-friendly, smooth transitions, gentle dynamics
Also used in production skills, not counted
-6 dBFS, -10 to -12 dBFS, -18 to -20 dBFS, -3 dBFS true peak, short fades at loop boundaries
Fragments to drop in
Airy and narration-friendly, gentle dynamics, no abrupt hits.Build subtle energy throughout, then a celebratory lift at the end.Clear midrange space left for a voice.
From production skills
Voice around -6 dBFS.SFX around -10 to -12 dBFS.Music around -18 to -20 dBFS.Final true peak at or below -3 dBFS.Voice stays above SFX; SFX stays above music.Never stack raw model outputs at full gain.
Post-generation mixing, not prompt text, but it is the only mix guidance in the corpus.
Vocals62.6% of audio prompts
Whether anything sings or speaks.
Named in 63% of audio prompts, almost always as a prohibition. Generators add humming and wordless vocals unprompted.
Fragments to drop in
No vocals, no lyrics, no spoken words.Wordless humming only, no intelligible words.No prominent lead melody competing with the narration.
Ending
How it stops.
The most common practical complaint about generated audio is that it fades or stops mid-phrase. Ask for a resolution.
Fragments to drop in
End on a clean major-key resolution, no fade.Ring out and decay naturally.Hard stop on the final beat.
Voice and spoken linefrom production skills
The prompt IS the line to be spoken; the voice itself is selected by id, never described in prose. Adjust --speech_rate modestly rather than rewriting a take.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
voice_type preset|element, voice_id, speech_rate, one voice per speaking entity, exact line, performance direction
Fragments to drop in
The gate is open. Move!same voice, calmer deliveryVoice prompts contain the exact line; performance direction belongs in the model-supported prompt, not undocumented flags.Voice: lock one voice per speaking entity and reuse it. Keep lines short enough not to block gameplay.
Narration writing (spoken script)from production skills
Narration is the one thing written in the user's chosen language; every image and video prompt stays English.
Documented in production prompt skills rather than measured in the corpus, so there is no frequency for it. Treat it as practice worth copying, not as a count.
Values
20-24 words, about 8-9 seconds, under roughly 9.5 seconds, plain spoken text only, no timecodes, emotion cues, parentheticals, or stage directions, Spell numbers out, hook through payoff, ~150 words per minute
Fragments to drop in
Write one plain line per clip, sized for about 8-9 seconds and normally 20-24 words. Keep every take under roughly 9.5 seconds.Use no timecodes, emotion cues, parentheticals, or stage directions.Spell numbers out.Set tone through word choice and concrete detail.Never say "in this video."For four and a half thousand years, the pyramids of Egypt have stood against the desert, silent and immense.Write scripts for speech, not text: short sentences, natural pauses, ~150 words per minute target. A 60-second video = ~150 words.Don't pad scripts to fit duration. The model paces itself; over-stuffed scripts get rushed delivery.
The order to write them in
A prompt is read start to finish, and earlier clauses set the frame the later ones are interpreted against. This order puts the decisions that cannot be undone first.
- Duration
- Role
- Mood
- Key and tonality
- Tempo
- Instrumentation
- Arrangement and dynamics
- Vocals
- Ending
- Negatives
How production prompts are laid out
Thin and single-sentence. One sound per prompt, described as material + action + recording character + an isolation clause, e.g. "short isolated sword impact, dry studio recording, no music". Music prompts name mood, tempo, instrumentation and then `instrumental/no vocals`. Voice prompts contain the exact line to be spoken and nothing else, with the voice chosen by id rather than described. Narration for assembled video is written as one plain labelled line per block (`Block 1` then the line), 20-24 words, about 8-9 seconds, with no timecodes, emotion cues, parentheticals or stage directions and numbers spelled out (higgsfield-websites/references/game-audio.md, higgsfield-video-explainer/references/prompts.md).
What to rule out
The most common element in the whole corpus. Put these last, in one sentence, and only include what is plausible for the shot: a long list of irrelevant prohibitions wastes the model's attention.
Text and branding
spoken words, words, other words, harsh textures, ad-libs that obscure the words
Content and safety
scary sounds
Audio
vocals, lyrics, sound effects, speech, music, spoken narration, voices, melody, singing, ad-libs, background music, vocal-like sounds, prominent lead melody
Other
medical sounds, abrupt peaks, dark, dramatic tension, heavy drums, long instrumental passages, sung parts, dialogue, chanting, tense passages, harsh timbres, abrupt stops
Fragment to drop in
No spoken words, words, other words, harsh textures, ad-libs that obscure the words, scary sounds, vocals, lyrics, sound effects, speech, music, spoken narration, medical sounds, abrupt peaks, dark, dramatic tension, heavy drums, long instrumental passages.
Exclusion lines used in production
Copied verbatim from working prompt skills. Useful as whole lines rather than as a vocabulary.
no musicno voiceno ambienceno vocalsinstrumentalno voice, dialogue, or narrationNever stack raw model outputs at full gain.
How this was measured
Every value on this page was counted in a corpus of 8,835 real prompts, 6,161 of them unique. The number beside each aspect is the share of prompts that specify it, and the values are the terms those prompts actually used, most used first. Nothing here is a list of plausible sounding options.
That distinction matters because the ranking is the useful part. A guide written from memory would lead with camera angles and lenses. In practice the most common element of all is the negative clause, present in 49.3% of video prompts and 73.6% of audio prompts, which is not where most advice starts.
Guidance and fragments are written by hand. Values are not. The process is documented as a repeatable skill so this can be rebuilt from a fresh export.
The cheatsheet, in full
Everything above condensed to one table per content type. This is the same content as the PDF, so you can read it here or take it with you. There is a standalone page for it too if you want to print from the browser.
Video
Order: duration > aspect > medium > shotSize > angle > movement > lens > dof > lighting > grade > motion > mood > audioBed > pacing > negatives
| Aspect | Values | Fragment |
|---|---|---|
| Duration | 6 seconds, 8 seconds, 10 seconds, 15 seconds, 30 seconds, 60 seconds | 15-second cinematic clip. |
| Aspect ratio | 16:9, 9:16, 1:1 | Framed 16:9 for landscape playback. |
| Shot size | close-up, wide shot, medium shot, extreme close-up, establishing shot, two-shot, macro shot, over-the-shoulder, closeup | Extreme close-up on the texture, filling the frame. |
| Camera angle | overhead, low-angle, eye level, three-quarter, top-down, profile, high-angle | Low-angle looking up, subject towering over the lens. |
| Camera movement | push-in, tracking shot, handheld, orbit, locked-off, slow zoom, truck, slow dolly, parallax, pull-back, slow pan, whip pan, crane up, dolly in | Slow push-in toward the subject, ending tight. |
| Lens and focal length | 35mm, macro lens, anamorphic, 100mm, 50mm, 85mm, 28mm, wide-angle lens, 32mm, 50 mm, fisheye, telephoto, 40mm, 35 mm | Shot on an 85mm lens, compressed and flattering. |
| Depth of field | shallow depth of field, bokeh, rack focus, f/1.4, f/1.8, f/2.0, soft focus, f/2.8 | Shallow depth of field, background dissolving into bokeh. |
| Lighting | soft lighting, rim lighting, rim light, silhouette, volumetric lighting, golden hour, volumetric light, soft light, side lighting, chiaroscuro, key light, low-key, practical lighting, backlit | Soft lighting from a large source, shadows barely there. |
| Colour and grade | vibrant, pastel, film grain, black-and-white, natural color, high contrast, desaturated, crushed blacks, natural colour, low contrast | Vibrant saturated grade, colours pushed but not clipping. |
| Medium and style | cinematic, photorealistic, anime, documentary, 3d render, unreal engine, stop-motion, 35mm film, watercolor, vector, photoreal, claymation, line art, octane | Cinematic, shot on 35mm film. |
| Subject motion | reveal, float, drift, slow motion, glides, rotates, time-lapse, ripples, unfolds, hovers, cascade | The lid glides open on hidden runners, slow and even. |
| Mood | playful, energetic, whimsical, calm, dramatic, minimal, elegant, uplifting, moody, ominous, ethereal, luxurious, serene, nostalgic | Calm and unhurried throughout. |
| Audio bed | percussion, bells, ukulele, xylophone, piano, drums, bass, marimba, strings, synth, guitar, choir, handclaps, pads | Minimal deep electronic ambience with soft mechanical clicks. |
| Pacing and cuts | One continuous take, no cuts. | |
| Resolution / quality tier | 480p, 720p, 1080p, 4k, basic, high, ultra, pro, std, standard, bitrate_mode standard|high | --resolution 4k |
| First/last frame anchoring | --start-image, --end-image, start frame, end frame, start = end, one-shot, loop, end-frame reveal, generic image/reference role , not a start-frame role | `--start-image` anchors the first frame. Prompt describes motion. |
| Ad format / mode (branded video) | ugc, ugc_how_to, ugc_unboxing, product_showcase, product_review, tv_spot, wild_card, ugc_virtual_try_on, virtual_try_on, hook, setting, ad reference | Looks like a real person filmed on phone |
| Edit instruction (video editing workflows) | make the jacket red, sketch, timestamp | --prompt "make the jacket red" |
Rule out. Text and branding: text, logos, subtitles, watermark, watermarks, captions, on-screen text, text on screen; Anatomy and faces: extra limbs, distorted faces, extra characters, hands, extra fingers, extra animals; Content and safety: violence, scary imagery, scary elements, weapons; Audio: music, narration; Other: dialogue, copyrighted characters, talking, cuts, horror, humans, clothes, flickering
Image
Order: aspect > medium > shotSize > angle > composition > lens > dof > lighting > grade > material > background > mood > text > negatives
| Aspect | Values | Fragment |
|---|---|---|
| Aspect ratio | 16:9, 9:16, 1:1, 4:5, 21:9, 3:2, 4:3 | Square 1:1. |
| Composition and framing | close-up, wide shot, medium shot, extreme close-up, two-shot, closeup, over-the-shoulder, establishing shot, macro shot, full shot | Centred on a single vertical axis, symmetrical. |
| Shot size | close-up, wide shot, medium shot, extreme close-up, two-shot, closeup, over-the-shoulder, establishing shot, macro shot, full shot | Extreme close-up, texture filling the frame. |
| Camera angle | three-quarter, profile, eye-level, low angle, overhead, high angle, top-down, bird's-eye, dutch angle | Three-quarter view, turned slightly off axis. |
| Lens and focal length | 35mm, 85mm, 50mm, 110 mm, 16mm, anamorphic, 24mm, 70mm, 48mm, fisheye, 600mm, 450mm, 40mm, 297 mm | 85mm portrait lens, natural compression. |
| Depth of field | shallow depth of field, bokeh, f/1.4, f/7.1, soft focus, f/2.8, deep focus, f/2.0, f/4, f/2, f/1.8, f/5.6 | Shallow depth of field at f/1.8, background soft. |
| Lighting | silhouette, rim light, soft light, soft lighting, backlit, key light, golden hour, high-key, rim lighting, backlight, softbox, low-key, volumetric light, bounce light | Large softbox front left, soft shadows falling right. |
| Colour and palette | vibrant, pastel, film grain, black-and-white, desaturated, high contrast, natural color, natural colour, monochrome, low contrast, sepia | Vibrant, fully saturated, clean whites. |
| Medium and style | cinematic, photorealistic, documentary, anime, 3d render, line art, watercolor, vector, 16mm, 35mm film, photoreal, kodak vision3, isometric, flat illustration | Photorealistic editorial photograph. |
| Materials and finish | glossy, matte, ceramic, concrete, leather, velvet, satin, brass, linen, frosted, marble, chrome, walnut, oak | Brushed metal with a satin finish, no mirror reflections. |
| Background | Seamless mid-grey studio backdrop, no horizon line. | |
| Mood | elegant, playful, minimal, dramatic, calm, energetic, whimsical, moody, luxurious, nostalgic, uplifting, ethereal, cosy, gritty | Elegant and restrained. |
| Text handling | No text, no lettering, no numbers anywhere in the image. | |
| Subject + setting + style (the opening clause) | a red fox curled in a snowy pine forest, golden hour, cinematic, subject + action, concrete and physical | a red fox curled in a snowy pine forest, golden hour, cinematic |
| Typography named in the prompt (when text is generated) | Anton, Anton SC, Bebas Neue, Oswald, Archivo Black, Montserrat, Poppins, Roboto Condensed, Inter, Barlow Condensed, Playfair Display, Cormorant Garamond, DM Serif Display, Fraunces | The prompt must state: Exact literal copy; Display/body font family names and which text uses each |
| Identity lock / subject fidelity | IDENTITY LOCK, photographic identity match, same bone structure, eye shape, nose, lips, jawline, skin tone, hairline and hair texture, Do not beautify, average, or restyle the face, IMAGE REFERENCES manifest | CHARACTER N: the person from attached face reference #K , IDENTITY LOCK: reproduce this exact person with a photographic identity match , same bone structure, eye shape, nose, lips, jawline, skin tone, hairline and hair texture. Do not beautify, average, or restyle the face. Expression: <emotion phrase>. |
| Aesthetic / seasonal preset | Clean girl, Cottagecore, Quiet luxury, Dark academia, Y2K, Christmas, Valentine's, Halloween, Black Friday, None, Clean studio, Lifestyle, Conceptual, With a model | Christmas version, quiet-luxury aesthetic |
| Product-photo genre / mode | product_shot, lifestyle_scene, closeup_product_with_person, moodboard_pin, hero_banner, social_carousel, ad_creative_pack, virtual_model_tryout, conceptual_product, restyle, levitating, floating, splash, frozen motion | bottle of cold-brew on a sunlit kitchen counter, IG feed |
| Logo / vector mark specification | lettermark_monogram, pictorial, abstract, mascot, emblem, one dominant unified silhouette, negative-space device, motif treatment, unexpected locked color-role pairing, symmetry or intentional asymmetry, internal detail limit, small-size scalability, centered isolated presentation, minimal anchor points | Flat vector design, clean lines, no shadows, no texture, no text. Clean editable vector paths, SVG-friendly, minimal anchor points. |
| Logo placement on a mockup surface | Target surface, Position, Scale, Orientation, Color, horizontally centered, upper third, optical center aligned to panel, clear-space margin, align to <panel edge/seam/baseline>, follow surface perspective | [PLACEMENT LOCK]
Target surface: <exact object panel/face/material>.
Position: <exact alignment and location, e.g. horizontally centered, upper third, optical center aligned to panel>.
Scale: logo occupies <specific proportion> of the target surface while keeping <specific clear-space margin>.
Orientation: align to <panel edge/seam/baseline>; follow surface perspective without changing logo proportions.
Color: use the supplied <full-color/black/white> variant exactly. State why it contrasts correctly with the material/background. |
| Seamless / tileable behaviour | seamless tileable game texture tile, uniform pattern density, perfectly seamless edges that wrap horizontally and vertically, no border, no vignette, no single focal object, photographed-flat appearance, seamless repeating pattern | seamless tileable game texture tile of <description>, uniform pattern density, |
| Key colour for transparency | bright magenta #FF00FF, bright green #00FF00, bright blue #0000FF | on a solid uniform bright <KEY COLOR> background, crisp edges, no drop shadow outside the element |
| Reference-image role and what it controls | style donors only, Image0, Image1, Image2, official logo, approved base scene, style-reference, representative-visual, grade/energy only | Make an Animated Explainer. Take only the visual render style and color grading of the input image(s); mix the styles if there is more than one image. Never use the characters, inscriptions, etc. from the input image(s) unless the instructions below ask you to. Use only the render style, and follow the user's instructions below: |
| Design-board direction vocabulary (web/layout imagery) | Pristine Light (paper/cream/off-white, dark ink), Deep Dark (charcoal/graphite), Bold Studio Solid (oxblood, royal blue, forest, vermilion, emerald fields), Quiet Premium Neutral (bone, sand, taupe, stone, smoke), technical grid/dot field, solid with soft ambient depth, full-bleed cinematic imagery, tactile paper/material texture, cinematic centered minimalist, asymmetric split, floating polaroid scatter, inline typography behemoth, editorial offset, massive image-first with restrained text | "website design mockup, desktop landing page section, [SECTION ROLE], [theme paradigm + exact palette words], [typography character] typography, [hero architecture / composition anchor], [background mode], [narrative spine motif], professional layout, clear hierarchy and spacing, award-winning web design" , plus: "no watermark, no browser chrome". |
| Reference photos for identity training | front, 3/4 left, 3/4 right, slight up/down, indoor, outdoor, soft, harsh, neutral, smiling, talking, head shot, head-and-shoulders, full body | Higher variety = better identity capture. |
| Image-to-3D input framing | t-pose, a-pose, Full body in frame, nothing cut off, Clean/plain background, Limbs visually separated from the torso, One character, no props overlapping the silhouette, lowpoly, three-quarter isometric view | Full body in frame, nothing cut off. |
Rule out. Text and branding: text, logos, captions, watermark, watermarks, logo, labels; Anatomy and faces: hands, arms, extra people, extra characters; Frame artefacts: frame, border, mockup; Content and safety: scary elements, weapons prominently displayed; Style and clutter: people, props; Other: overlays, slogan, letters, sunglasses, dramatic pose, helmet, police department names, cinematic color grading
Audio
Order: duration > role > mood > key > tempo > instrumentation > arrangement > vocals > ending > negatives
| Aspect | Values | Fragment |
|---|---|---|
| Duration | 15 seconds, 30 seconds, 60 seconds, 90 seconds, 120 seconds | 90-second bed. |
| Role | Instrumental music bed to sit under narration. | |
| Tempo | 70 BPM, 90 BPM, 100 BPM, 116 BPM, 128 BPM | Buoyant mid-tempo around 116 BPM. |
| Key and tonality | major key, minor key, modal, pentatonic | Bright major key. |
| Instrumentation | percussion, ukulele, piano, bass, handclaps, marimba, bells, strings, guitar, xylophone, drums, harp, synthesizer, pads | Marimba, ukulele, xylophone, light percussion and soft handclaps. |
| Mood | warm, playful, cheerful, buoyant, wholesome, ambient, uplifting, driving, triumphant, minimal, dreamy, suspenseful | Cheerful, wholesome and playful. |
| Arrangement and dynamics | airy, midrange, headroom, no abrupt hits, narration-friendly, smooth transitions, gentle dynamics | Airy and narration-friendly, gentle dynamics, no abrupt hits. |
| Vocals | No vocals, no lyrics, no spoken words. | |
| Ending | End on a clean major-key resolution, no fade. | |
| Voice and spoken line | voice_type preset|element, voice_id, speech_rate, one voice per speaking entity, exact line, performance direction | The gate is open. Move! |
| Narration writing (spoken script) | 20-24 words, about 8-9 seconds, under roughly 9.5 seconds, plain spoken text only, no timecodes, emotion cues, parentheticals, or stage directions, Spell numbers out, hook through payoff, ~150 words per minute | Write one plain line per clip, sized for about 8-9 seconds and normally 20-24 words. Keep every take under roughly 9.5 seconds. |
Rule out. Text and branding: spoken words, words, other words, harsh textures, ad-libs that obscure the words; Content and safety: scary sounds; Audio: vocals, lyrics, sound effects, speech, music, spoken narration, voices, melody; Other: medical sounds, abrupt peaks, dark, dramatic tension, heavy drums, long instrumental passages, sung parts, dialogue
Say what you want, not what you do not
Most image and video models do not expose a separate negative prompt field, so a prohibition inside the prompt can act as a mention and pull the thing you banned into the frame. Where that is a risk, state the positive instead.
Most models don't expose a `negative_prompt`. Phrase positively:Instead of "no blur" to "tack sharp"Instead of "no people" to "uninhabited landscape"NEGATIVE: color drift, photorealism, 3D render, lip-sync, captions, on-screen text, logos, watermark{, plus style-specific bans}.NEGATIVE: <style drift and realism bans; no lip-sync, captions, text, logos, or watermark>Regenerate with "no text, no letters, no logos, no captions" appended.Never: <forbidden treatments>Change ONLY the person's facial expression to: <phrase>. Keep identity, face structure, hair, pose, body, clothing, logo, background, lighting and composition EXACTLY unchanged, pixel-faithful.Honor every `forbidden_elements` entry literally. Never replace one forbidden cliché with another generic symbol.Known failure mode: false-positive `nsfw` flags on innocuous product prompts , the fix is removing mood/atmosphere words ("seamless loop feel", "ambient", "intimate") and re-describing the shot as a plain product film/photo.Models reject prompts with `nsfw` or `ip_detected` terminal status. Avoid: Real public figures; Sexual content; Trademarks / branded characters
Which model for which job
Recorded from production prompt skills, not measured here. Model names move fast, so treat this as a snapshot of practice at the time it was written.
| Model | Best for | Prompt conventions |
|---|---|---|
| GPT Image 2 (gpt_image_2) | Default high-fidelity image generation; graphic design, UI, banners, typography, and any brief with on-image text; generated product concept / packaging / can / bottle with brand name or label text; the controlled second stage that adds readable text or exact graphic details onto an existing scene; seamless-texture edits; character animation sheets as a 5x5 grid in one image. | The model to reach for when literal copy must be inside the image , state exact literal copy, display/body font family names and which text uses each, logo placement/scale/clear space/colour variant, text hierarchy/alignment/line breaks/contrast, palette roles and aspect ratio. In brandkit it is used only for that controlled stage and must preserve Image0's camera, crop, objects, lighting, material, folds, shadows, perspective and background exactly. Does not currently expose 4:5 , generate 3:4 with a centred 4:5 safe area and crop. Texture prompts run at aspect 1:1, quality high, with the reference attached as a media input. |
| Nano Banana 2 (nano_banana_2) | Fast everyday default for character, cartoon, stylized and reference-driven image work; edits; game sprites, tiles, backgrounds and textures; the style-key image in the explainer pipeline; marketplace cards under the hood. | Reference-driven , pass references with repeated --image. For game assets use resolution 1k, AR 1:1 for sprite/tile/texture/ui and 16:9 for backgrounds, count 1 each. Keeping one primary generator across asset kinds is itself a cohesion win. |
| Nano Banana 2 Lite (nano_banana_2_lite) | Fast reference-driven generation and edits when the brief is simple or speed/cost matters more than Pro fidelity. | Supports up to 14 image references via image_references; aspect_ratio=auto requires at least one image reference. |
| Nano Banana Pro (nano_banana_pro) | Harder briefs where Nano Banana 2 falls short; the main thumbnail render; photoreal and reference-driven website imagery; deriving multiple angles of one approved product so every view is recognizably the SAME product. | For thumbnails, run at explicit 4K with the long multi-block prompt written to a file and piped on stdin; face photos first in character order, then the logo. |
| Higgsfield Soul 2.0 (text2image_soul_v2) | Aesthetic UGC, fashion editorial, lifestyle character stills; Soul Character reference work. | Accepts a Soul Character reference via --soul-id; quality tier 1.5k or 2k; aspect ratios 1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3. |
| Soul Cinema / soul_cinematic | Cinematic stills and film-grade lighting; the pick when the user asks for "cinematic" or wants concept-art mood; longer talking-head-style Soul work. | Same ratios as Soul 2.0 plus 21:9; quality 1.5k or 2k; used for the PHOTO route on covers with 1-2 hosted scene references as --image. |
| Soul Cast | Distinctive, characterful personas , a creative, expressive character rather than a photoreal default. | Text-only; prompt-only model that rejects reference media. |
| Soul Location (soul_location) | Best-in-class environments and locations; pure scene and place generation with no person in frame. | Prompt-only, no media. No quality/resolution selector , dimensions are fixed by aspect ratio. |
| Seedream (seedream_v5_pro, seedream_v4_5, Seedream 5.0 Lite) | Primary photoreal mockup generator in brandkit; face-anchored edits into a complex new scene without heavy filters; surgical thumbnail tweaks. | Mockup prompts use the bracketed [CREATE ONE FINISHED BRANDED MOCKUP] / [AUTHORITATIVE LOGO] / [PLACEMENT LOCK] / [PHYSICAL APPLICATION] contract with Image0/Image1 roles stated explicitly; never pass an SVG path as an image reference; never ask it to render readable text. For edits, scope narrowly and state that every other pixel-level property is unchanged. No 4:5 ratio. |
| Recraft V4.1 (recraft_v4_1) | Clean graphic and vector-style design assets , logos, icons, flat illustrations, brand marks, controlled-palette visuals; the only model used for new logo marks. | Use --model_type vector for vector output, standard for raster-style graphics. One continuous prompt string per candidate, no bullets, ordered mark type to mechanism to shape logic to style register to palette behavior to composition to constraint tail. Never put hex, RGB or Pantone in the prompt , pass colours through the colors and background_color params. Never use lens, camera, lighting, depth of field, photorealistic, cinematic, material rendering, grain, paper texture, shadows, mockup or scene language. The exact phrase "no text" must appear in the constraint tail. All three candidates use identical model, palette, background, aspect and quality parameters. Prompt-only , rejects media inputs. |
| Flux 2.0 (flux_2) | Precise prompt adherence; a creative alternative to the Banana family; key-pose generation in the sprite pipeline; cheap photorealistic drafts. | Key-pose prompts carry the STYLE FORMULA verbatim plus one pose instruction, and must demand full body in frame with empty margin above the head and below the feet plus a clean uniform key-colour background absent from the subject's palette. |
| Flux Kontext Max | Context-aware editing and style transfer; anime, stylized looks, typography remix when defaults feel too generic. | Reach for it when default models feel flat. |
| Z Image (z_image) | Fastest in the catalog , speed, drafts and LoRA-driven stylization; cheap prompt iteration. | Prompt-only, rejects media. Accepts a prompt on stdin. |
| Grok Imagine | Expressive, high-contrast, bold creative output; worth trying for anime and stylized looks. | Also has a video variant with audio support. |
| Seedance 2.0 (seedance_2_0) | Default all-purpose serious video , multi-shot, consistent identity, motion-heavy, cinematic, image-to-video, 4-15s up to 4K; seamless loops; the end-frame cover reveal; dramatic motion from a still. | Accepts image, start_image, end_image, video and audio roles. Pass the finished cover as --end-image and describe the buildup so the clip lands pixel-perfect on it; do NOT pass it as --start-image. Audio reference comes through the audio media role, never --generate-audio. Aspect ratios auto, 21:9, 16:9, 4:3, 1:1, 3:4, 9:16; duration 4-15; resolution 480p/720p/1080p/4k. |
| Seedance 2.5 (seedance_2_5) | Reference-driven video, editing an existing video, or extending one. | Modes t2v / omni_reference / video_edit / video_extension with image/video/audio reference arrays. Caps at 720p , not a newer Seedance 2.0. |
| Seedance 1.5 Pro / seedance1_5 | Cheap clean single-take shots without cuts; the exact model for the 2D sprite animation pipeline. | For sprites: fixed 4 s, 720p, aspect ratio equal to the key pose and passed explicitly on every call, start frame always, end frame only for loops, one action in the positive prompt plus the mandatory camera/prop/facing negative block. "Do not silently substitute a model." |
| Kling 3.0 (kling3_0) and Kling 3.0 Turbo | Cheaper Seedance 2.0 substitute for single-plane scenes without heavy motion; cleanest image-to-video with a start frame; Turbo when the user wants speed or lower cost. | Roles start_image and end_image (Turbo start_image only). 16:9, 9:16, 1:1; duration 3-15s; modes pro/std; sound on/off. |
| Gemini Omni Flash (gemini_omni) | Fast multimodal reference-to-video from image references and optionally one video reference; the clip generator in the explainer pipeline. | 0-7 image references and 0-1 video reference (max 5 images when a video reference is present). In the explainer pipeline every call attaches the same style key via --image and runs 10s at 720p. "Never silently replace gemini_omni." |
| Google Veo 3.1 / Veo 3.1 Lite / Veo 3 | Ultra-realistic top-tier cinematic quality; Lite for fast batch and volume work. | Format set is constrained , Veo 3.1 accepts only 16:9 or 9:16 and durations 4, 6 or 8, with quality tiers basic/high/ultra. Verify accepted aspect ratio and duration before submitting. |
| Grok Video 1.5 (grok_video_v15) | Bold, anime-like, high-contrast or experimental image-to-video from one starting image. | Requires exactly one --start-image or --image; duration 2-15s; resolution 480p or 720p. |
| Minimax Hailuo | Cheap with strong physics when natural-physics motion matters and audio is not needed. | No audio in current variants. |
| Wan 2.7 / Wan 2.6 | Synchronized audio with character-consistent video (2.7); stylized, experimental, intentionally artistic cheap work (2.6). | Some Wan configs are prompt-only and reject media inputs. |
| Cinema Studio Video 3.0 / Cinema Studio Image 2.5 | Top-tier cinema-grade execution and dramatic film look at the highest fidelity; cinematic still frames up to 4K. | Prefer Cinema Studio Video 3.0 as the modern cinema default; reach for earlier variants only when the user names them. |
| Marketing Studio (marketing_studio_video / marketing_studio_image) | All advertising and commercial work , UGC, unboxing, TV spot, product showcase, product review, virtual try-on, branded ad images with a presenter avatar and product. | The prompt is the hook/brief; a selected hook's text is prepended to it and must not be copied into --prompt unless the user wants the wording reinforced. Mode drives staging (default ugc). hook_id and setting_id are video-only and valid only for ugc, ugc_how_to, ugc_unboxing, product_review and ugc_virtual_try_on; hooks are weak without product context. --generate_audio true is supported here (unlike seedance_2_0). Aspect auto/21:9/16:9/4:3/1:1/3:4/9:16; resolution 480p or 720p; duration integer ≥ 4. |
| Seed Audio 1.0 (seed_audio) | Default audio generation , text-to-audio, sound effects, ambience, foley, impacts, environmental audio, voice-style generations and music-like audio; every narration take in the explainer pipeline. | Requires --prompt. Optional audio_references or image_references, mutually exclusive. For narration the prompt is the block's spoken line only, with voice chosen by --voice_type (preset|element) and --voice_id, and --speech_rate adjusted modestly for an overlong take. |
| Sonilo Music (sonilo_music) | Music from text , backing tracks, instrumental beds, jingles, musical moods. | Requires --prompt and --duration; takes no media inputs. Prompt names mood, tempo, instrumentation and 'instrumental/no vocals'. |
| Mirelo Text to Audio (mirelo_text_to_audio) | Non-speech audio , sound effects, ambience, foley, impacts, transitions, environmental sounds. | Requires --prompt and --duration; no media inputs. One sound per prompt with an isolation clause. |
| Inworld Text to Speech (inworld_text_to_speech) | Explicit TTS with one of its listed voices. | The prompt is the exact line to be spoken; pass --voice with an exact voice value from the model contract. |
| Multi-Image to 3D (multi_image_to_3d) | An actual 3D mesh/GLB from 1-4 object or product reference images. | No text prompt , the reference images are the input. One figure per image, full body, clean plain background with no shadows/text/watermarks/logos, limbs separated from the torso; concept images use three-quarter isometric view and a pure flat white background. Set pose_mode t-pose or a-pose for characters that will be rigged. |
| Auto | Smart routing when the user's intent is open and you don't want to commit to a specific image model. | Picks the best image model from the prompt automatically. |
| Virality Predictor (brain_activity) | Scoring a finished clip's hook, attention, retention, distraction risk and virality potential. | Takes --video and needs no prompt at all; returns a text report, not media. |
| image_background_remover / bytedance_image_upscale / outpaint | Transparent subject cutouts, upscaling art below the 1500px floor, extending an image. | No prompt. Background removal runs on frames one by one, never on a video, because the video path smears edges into a flickering halo. |
Prompt cheatsheet
Every aspect, value and negative on two pages. Keep it open while you write.