Guides / Filmmaking

A short film with AI: the pipeline from script to final cut

The point of a short is the pipeline, not any one shot. Last Train is twenty five seconds long and has one lead, one busker, one platform and one chrome case. Here is every stage that made it, with the real prompts, the model for each stage, and the places the pipeline pushed back.

  • 13 min read
CastFramesShotsThe cut

What you make, in order. Tap one to jump to its step.

Last Train is a 25 second film made for this guide: Vale runs for the last train, misses it, and laughs. Everything in it is real output from the company's own account, made on the live composers in the order below, and every person in it is fictional. The same set of shots is measured, object by object, in the consistency guide.

Two honest limits first. A saved character carries its wardrobe and its written description into every shot; we have not measured whether it is one person from shot to shot, and this guide does not claim it. And the model we planned to shoot on refused every character reference we sent it, so the film was made on a different model than the plan said. Both are in the steps.

What goes wrong

MistakeWhy it failsFix
Shooting the first clip before the cast existsEvery shot invents its own person, coat and platformSave the character, the location and the prop still first
Sending a character straight to a video modelSome models refuse a photoreal person as a reference; one treats the portrait as the first frame and returns a portrait-shaped clipMake a 16:9 frame in the image composer with the cast attached, then animate that frame
One long shot for a whole beatMotion drifts, props migrate, the set repeats4 to 6 seconds per shot, one action each; cut where it drifts
Describing the prop differently each timeA chrome case becomes a briefcase, then a bagOne sentence for the prop, pasted into every prompt, and a still of it for inserts
Grading each clipFive clips graded five ways look like five filmsGrade the timeline once in Studio

Before you start

  • A script with named beats. Five beats is twenty five seconds.
  • One sentence per character, per place and per prop, written the way a costume or art department would write it. These sentences are pasted, unchanged, into every prompt.
  • A NOLGIA account. Characters, Locations, the composers and Studio come with it.
  • Patience for one redo per stage. Ours needed five, and the guide says where.
The script
LAST TRAIN. Night. An underground platform, magenta neon over white tiles. 1. Vale runs for the last train, a chrome case in one hand. 2. She passes Ozzy, the busker, who turns to watch. 3. The doors close in her face and the train pulls away. 4. The case, set down on the wet platform. 5. On the bench, out of breath, she laughs.

Save the cast, the set and the prop

Open Characters. Type a name and the character's one sentence as the description, open Add reference images, and either upload a photo or use Generate with a portrait prompt and a model. With the portrait selected as the reference, press Create character. We made Vale this way on GPT Image 2.5 Flare, and Ozzy from a still made the same way.

Vale's portrait prompt (Characters page, Generate)
Full-length portrait of a woman in her late twenties, silver-blonde chin-length bob with a straight fringe, pale skin, red lipstick, wearing a glossy red vinyl trench coat belted closed over a black ribbed turtleneck, black wide-leg trousers and chunky white boots, holding a folded transparent bubble umbrella, standing on an underground subway platform with white tiles and a magenta neon strip along the ceiling, wet floor reflections, thin haze. Photoreal, cinematic, editorial fashion lighting, no text, no logos.

Open Locations and do the same for the place: a name, a description, the one sentence as the canonical description, and a still from your Library or an upload. Our platform is a Seedream 5.0 Pro still. The prop is a still too, made in the image composer and kept in the Library; it rides into an insert as a reference image and into every other prompt as one sentence.

Vale: a woman with a silver-blonde bob in a glossy red vinyl trench coat belted over a black turtleneck, black wide-leg trousers and chunky white boots, a folded clear umbrella in hand, on a neon-lit platform
Vale, made on the Characters page with GPT Image 2.5 Flare. The lead.
Ozzy: a man with long dark curls and a light beard in an oversized yellow rain jacket with grey reflective strips over a black hoodie, holding a tenor saxophone on a platform
Ozzy, made with GPT Image 2.5 Flare. The busker.
An empty underground platform: white tiles with a teal stripe, a magenta neon strip along the ceiling, a wet floor, a yellow safety line, a steel bench and a tunnel mouth glowing cyan
Halcyon Street station, made with Seedream 5.0 Pro. Saved as a Location.
A small brushed-chrome hard-shell case with a black rubber handle and two black latches on a dark studio surface
The chrome case, made with Seedream 5.0 Pro. The prop.
The station, as saved on the Location
An empty underground subway platform at night: glossy white tiles with one teal stripe at waist height, a long magenta neon strip along the ceiling, wet floor with mirror reflections, thin haze, a yellow safety line along the platform edge, a single steel bench, the tunnel mouth glowing cyan. No readable text or signs.

Make each shot as a frame first

Open the image composer. Pick GPT Image 2.5 Flare, set the shape to 16:9, then press Characters and pick Vale, and Location and pick Halcyon Street station. The composer shows a chip for each. Write the shot as one moment, name the character with an @ in the prompt, and press Generate images. Five frames, one per shot, cost less than one video take between them and are where the wardrobe, the props and the framing get fixed.

Frame 1, the run, as sent (Vale and the station attached)
@Vale mid-sprint along the empty platform towards the camera, her red vinyl trench coat flying open over the black turtleneck, the folded transparent umbrella in her left hand and a small brushed-chrome hard-shell case with a black rubber handle and two black latches in her right, white boots splashing the wet tiles. Wide low shot, the magenta neon strip above her, thin haze, the tunnel glowing cyan behind. Cinematic film still, anamorphic, fine grain, no text.
Vale sprinting towards the camera along the platform, red coat flying, the chrome case in her right hand and the umbrella in her left
Frame 1, the run. Vale and the station attached; the case by description.
Vale striding past Ozzy, who stands by the steel bench with his saxophone under the magenta neon; she carries the chrome case and the folded umbrella
Frame 2, the busker. Vale, then Ozzy, in the cast, and the station. The second version: the first left the case out of the prompt, and so out of the frame.
Vale reaching for the closing doors of a train at the platform, the chrome case in her other hand, her reflection in the door glass
Frame 3, the doors. Vale and the station.
The chrome case standing alone on the wet platform beside the yellow line, the bench and the cyan tunnel soft behind it
Frame 4, the insert. No character: the location attached, and the case by its sentence.
Vale sitting on the steel bench under the neon, the umbrella across her knees, looking down the platform
Frame 5, the bench. Vale and the station.

Tip:Name who goes where when two people share a frame. Frame 2 says who passes whom, where Ozzy stands, and what she carries. Our first version of it left the case out of the sentence, so the frame had no case. With two characters in the cast, the composer numbers them in the order you picked them, and the prompt's @Vale and @Ozzy resolve to those references.

Animate the frames into shots

Open the video composer, choose Animate an image under Start from, press Add start frame, then Pick an asset and choose the frame from your Library. It becomes the first frame of the clip. Write the action that happens from that frame, press a camera move chip, set the length, and generate. We used Kling v3 for the sprint, the shot with the most movement, and Seedance 1.5 Pro at 720p for the other four, 4 to 6 seconds a shot, and cut two of them shorter in Studio. For a run, a jump or anything with weight, animate the frame on Kling v3 or Seedance 2.5.

Shot 1 from frame 1, as sent (Kling v3, camera chip: Locked off)
She sprints straight towards the camera along the platform, red vinyl coat flying open, the chrome case in her right hand and the folded umbrella in her left, white boots splashing the wet tiles, hair bouncing with each stride, the magenta neon and the cyan tunnel glow behind her. Real sprint physics, cinematic, anamorphic, fine film grain.
Vale sprints towards the camera along the platform, case in one hand, umbrella in the other
Shot 1, made with Kling v3 from frame 1. 5 s, locked off. The case stays in her right hand and the umbrella in her left for the whole sprint.
Vale hurries past Ozzy with the case and the umbrella; he keeps playing and turns his head to watch her go
Shot 2, Seedance 1.5 Pro from frame 2. 6 s, tracking, the third take; cut at 5 s in the edit, where she leaves the frame. She carries the case and the umbrella the whole way past him.
Vale reaches for the closing doors, they shut on her reflection, and the train pulls away
Shot 3, from frame 3. 6 s, handheld.
The chrome case standing still on the wet platform by the yellow line as a train streaks past behind it
Shot 4, the insert, from frame 4. 4 s, rack focus. The case stands still from the first frame; the train streaks past behind it.
Vale on the bench, catching her breath, then laughing and shaking her head
Shot 5, from frame 5. 5 s, push-in.

Watch for this:Why frames first, and not the character straight into the video model. We tried the direct route twice. Seedance 2.0 Fast refused every request that carried Vale's reference, saying the reference looked like a real person, and refunded each one. Seedance 1.5 Pro accepted her, but used her 3:4 portrait as the clip's first frame, so a 16:9 request came back portrait-shaped. A 16:9 frame from the image composer, animated, sidesteps both.

A portrait-shaped video frame of Vale on the bench, taller than it is wide, the result of attaching her portrait directly on Seedance 1.5 Pro
What the composer does with a portrait as the reference on Seedance 1.5 Pro. The clip was asked for at 16:9 and came back 834 by 1112. Kept here as the honest figure.

Score and effects

The film has no dialogue, so sound is two runs in the audio composer: a score from Lyria 3.5 and two effects from ElevenLabs Sound Effects v2, the doors closing and the running boots. Describe the score by tempo and mood, and each effect as one event.

The score
A tense, driving electronic cue at 118 bpm for a short film about missing the last train: pulsing arpeggiated synth, deep sub bass, a ticking hi-hat, a sudden drop to a single warm pad at the end, no vocals, cinematic neon night
The doors
Subway train doors closing: a two-tone chime, a pneumatic hiss, doors thud shut, then the train accelerates away through a tunnel with a rising electric whine and echo, station reverb

Cut, grade and export in Studio

Open Timeline, press New timeline, name the project and the timeline, then open Library in the Media column and add the five shots in order with their + buttons: each press adds the clip at the playhead on a fresh track, and End moves the playhead to the end of what is there. For shot 2, drag the Source Preview to 5 seconds and press Mark Out before the +, so only its first 5 seconds come in. Add the score at the start; it is 2 minutes long, so park the playhead at the end of the picture, press C with nothing selected to split it, click the tail and press Delete. Add the two effects where the boots land and where the doors close. Then open Color from the toolbar with nothing selected and pick one film stock for the whole timeline; ours is CineStill 800T.

Studio on a computer: five clips staggered across five tracks in sequence, the sound effects and the score on audio tracks beneath, Vale mid-run in the program monitor
The cut on a computer. Five shots in sequence, the two effects on A1 and A2, the score on A3, trimmed to the picture.
The Canvas color panel open on the right of Studio with the film stock grid, CineStill 800T selected
The grade. Color with nothing selected grades the canvas, so every clip gets the same stock.
Studio on a phone: a note that the editor is best edited on a desktop, the program monitor at zero, and the NOLGIA Agent chat for edits
On a phone. Studio plays the cut and takes edits through the agent; the timeline wants a bigger screen.
Shot 3 before the grade: the neon and the coat as the model rendered themShot 3 after the timeline gradeBeforeAfter
The same moment of shot 3, 2 seconds in: the clip as generated on the left, the export with the CineStill 800T timeline grade on the right.

Export video renders the MP4 in a minute or two; it lands in your Library at no charge. The arrow beside it opens the Export menu; Premiere Pro or DaVinci Resolve downloads one ZIP with sequence.xml, a media folder with every clip and track, your grade as luts/canvas.cube and a README. Unzip it, then import sequence.xml. Ours came with one warning per video clip: the XML assumes each clip matches the canvas size, so use Scale to Frame Size in Premiere if a clip looks wrong. The export post covers the rest of what carries over.

The Export menu open under the Export video button: Video (MP4), PNG, PNG with transparent background, Premiere Pro or DaVinci Resolve, After Effects
The Export menu. Video, a still, or the cut for another editor.
Last Train: Vale runs, passes the busker, misses the doors, the case on the platform, and the laugh on the bench
Last Train, the export. Five shots, one grade, a score and two effects, 25 seconds.

Which tool for which job

StageWhereModel we used
CastCharacters, GenerateGPT Image 2.5 Flare
Set and prop stillsImage composer, then LocationsSeedream 5.0 Pro
FramesImage composer with Characters and Location attachedGPT Image 2.5 Flare
ShotsVideo composer, Animate an imageKling v3 for the sprint, Seedance 1.5 Pro for the rest
Insert with a propImage composer for a frame of the prop on the set, then Animate an imageGPT Image 2.5 Flare, Seedance 1.5 Pro
Score and effectsAudio composerLyria 3.5, ElevenLabs Sound Effects v2
Cut, grade, exportStudioNo model, no charge

Where you still regenerate

  • A model that refuses your cast. Seedance 2.0 Fast refused Vale's reference on every try. Switch models rather than rewording; the refusal names the checker, not your prompt.
  • A portrait as a start frame. On Seedance 1.5 Pro a character attached directly becomes the first frame. Go through a 16:9 frame.
  • The set repeating. The second take of shot 2 ran past the busker twice by 4 seconds. Shorter shots, or cut where it repeats.
  • A prop that migrates or arrives. The first take of shot 2 gave Vale the saxophone by 2 seconds; the first take of the insert had the case fade in from nothing. Say who holds what in the shot prompt, and for a static prop put it in the frame you animate from.
  • Faces. Wardrobe, hair and props hold across the five shots; whether the face is one person is not measured here.

What it costs

Frames are priced per image, shots by model, length and quality, the score and each effect per run. Studio, the cut and the export are free. Live rates for the models in this film:

  • GPT Image 2.5 FlareEvery plan

    Per image

    Native 8 credits · 2K 21 credits · 4K 58 credits

  • Seedream 5.0 ProEvery plan

    Per image

    1K 5 credits · 2K 9 credits

  • Kling v3Pro and up

    Per 5s clip (audio on by default)

    720p 35 credits · 1080p 47 credits · 4K 117 credits

  • Seedance 1.5 ProPro and up

    Per 5s clip (audio on by default)

    480p 7 credits · 720p 15 credits · 1080p 33 credits

  • Lyria 3.5Every plan

    Per generation

    Standard 5 credits

  • ElevenLabs SFXEvery plan

    Per generation

    Standard 4 credits

Read live from the catalog when this page loads. Every price is shown before you generate. See every rate.

Checklist

  • Character, location and prop saved before the first frame; one sentence each, reused verbatim.
  • One 16:9 frame per shot with the cast and the location attached; wardrobe and props checked at full size.
  • Each frame animated at 4 to 6 seconds with one action and one camera move.
  • Score by tempo and mood; effects one event each.
  • Trimmed where the model drifts; one grade on the timeline; exported once.

Try it yourself

Save your own character and location, then use the frame 1 prompt with the cast and location attached on GPT Image 2.5 Flare at 16:9, and the shot 1 prompt on Kling v3 in Animate an image with the Locked off chip. We ran the frame 1 recipe again on the live composer after the guide was written: it came back on the same platform with the same coat, bob, boots, umbrella and case, in a different stride and with the coat belted closed, although the prompt says it flies open. Expect the pose and the small things to roll; the character and the place are what the references hold. The stills below are our cast, set and prop; the frames and shots are above.

The ingredients

  • Vale, the portraitWebP image · 1200x1600 · 144 KBDownload
  • Ozzy, the portraitWebP image · 1200x1600 · 168 KBDownload
  • The station stillWebP image · 1536x864 · 107 KBDownload
  • The chrome case stillWebP image · 1200x1200 · 68 KBDownload
  • Frame 1, the runWebP image · 1600x900 · 161 KBDownload
Upload the portrait as a character's reference and the station as a location's reference to rebuild the film's cast on your own account.

Questions and answers

Is Vale one person in every shot?
Her coat, hair, boots, umbrella and the case hold across the five shots, and the consistency guide measures that. Whether it is one person from shot to shot is not something we have measured, so we do not claim it.
Why make a frame before every shot?
A frame costs a fraction of a video take and is where the wardrobe, the props and the framing get fixed. It also avoids two problems: a video model refusing a character reference, and a model using a portrait as the first frame.
Which video model should I use?
For the running shot, Kling v3: it keeps a sprint and the props in her hands believable. We planned Seedance 2.0 Fast and it refused every character reference, so the quieter shots are Seedance 1.5 Pro from frames. For anything with weight, use Kling v3 or Seedance 2.5.
How long can the film be?
As long as your shot list. Twenty five seconds was five shots of 4 to 6 seconds. Keep shots short; the models drift and repeat past about 5 seconds.
Can the characters speak?
This film has none. For lines, record them in the audio composer and lay them under the cut, or use a model with native speech for a single shot.
Can I finish it in Premiere or Resolve?
Yes. Studio's Export menu downloads one ZIP with the sequence, the media and the grade as a LUT. Unzip it first, then import the sequence.

Start with the cast

Open Characters

Filmmaking

Keep characters and locations consistent across shots

A saved character, a saved location and one repeated sentence per prop: what actually held across five shots of one film, measured object by object and by colour, and the two gaps that did not.

· 10 min read

Filmmaking

Make a video from a script, step by step

Split a 60 to 90 word script into shots, generate each one in the video composer with a camera move, record the narration, cut it in Studio and export. Every step run for real.

· 12 min read