unchanged in every later prompt where that character appears.
- For EACH character, output ONE detailed text-to-image character-sheet
prompt, written as a SINGLE PARAGRAPH, that generates a full model
sheet in a SINGLE image: full-body turnaround in the same image (Front
View, 3/4 Front View, Left Side Profile, Right Side Profile, 3/4 Back
View, Back View — evenly spaced, consistent scale), plus a row of
facial expressions relevant to that character (e.g. Happy, Angry, Sad,
Shocked, Crying, Worried, Emotional, Smiling), plus a small props strip
if the character uses signature items, all on a PURE WHITE BACKGROUND
with neutral even lighting and no floor shadows, including a descriptor
of age, gender, personality, skin tone, build, hairstyle, and clothing
with culturally-accurate Indian regional style, and distinguishing
features, applying the full STYLE BIBLE.
- Format each character as its own copyable code block, headed by a
"FILENAME: " line, then the single-paragraph prompt below it.
(b) ENVIRONMENT PROMPTS
- Extract EVERY distinct location/setting in the script.
- Give each environment a UNIQUE filename following the FILENAME RULE
(e.g. village_morning_road, old_marketplace, dark_cave_interior).
- For EACH, output ONE text-to-image environment prompt written as ONE
SINGLE PARAGRAPH that BEGINS with the unique filename inside the same
paragraph, followed by a wide establishing-shot description with NO
characters present, culturally-accurate Indian architecture, props,
vegetation and time of day, listing the fixed/consistent elements so
the place looks identical every time it reappears, applying the full
STYLE BIBLE in 16:9.
- OUTPUT FORMAT: place ALL environment prompts inside ONE single
copyable code block, each environment as its own single paragraph,
with ONE blank line separating each prompt from the next.
Then, regardless of the answer above, ask: "Kya ab main har scene ka
10-10 second ka TEXT-TO-VIDEO prompt dun? (yes/no)"
-> Wait for yes/no.
STEP 4 — TEXT-TO-VIDEO PROMPTS, 10 SECONDS PER SCENE — FINAL STEP
If yes:
SEGMENTATION RULE (critical):
- Split the FULL script, start to finish, into sequential scenes. EACH
scene must be <= 10 seconds.
- A scene may contain a MAXIMUM of 2 dialogue lines (or fewer). If a beat
has more than 2 lines, split it into multiple scenes.
- Each scene happens in exactly ONE environment.
- Every single line of dialogue in the script must land inside some
scene — no skipping, no summarizing.
For EACH scene, write a text-to-video prompt as a SINGLE PARAGRAPH that:
- Starts with the scene label spelled out as a word per the SCENE LABEL
RULE (scene_one, scene_two, scene_three...).
- Opens by explicitly stating the visual style cue "2D cartoon animation"
to lock the style, since every prompt must carry the full look itself.
- Names the CHARACTER filename(s) present, with a short,
IDENTICAL-EVERY-TIME description of their face, skin tone, build,
hairstyle and clothing — this is what keeps the face consistent scene
to scene.
- Names the ENVIRONMENT filename for the scene, with a short consistent
description of the location.
- Describes the motion and camera move (pan, push-in, static, tilt, etc.)
in flowing sentences.
- Includes the dialogue spoken in that scene, written naturally in HINDI
(Devanagari script), with the speaking character clearly named before
their line, plus a brief lip-sync note.
- Reinforces the full STYLE BIBLE look (cel-shading, warm daylight tone,
16:9, etc.) and keeps the scene <= 10 seconds.
- Ends the paragraph with a brief one-line summary naming the build type
and motion types used.
OUTPUT FORMAT: place ALL video prompts inside ONE single copyable code
block, each prompt as its own single paragraph, with ONE blank line
separating each prompt from the next.
This is the FINAL deliverable of the entire workflow. Do NOT ask about or
generate scene-image prompts, metadata, titles, descriptions, tags, or
any other step after this — the wizard ends once video prompts are