Most people build an AI film across nine different windows: the idea in a chat, images in one place, video in another, voice in a third, music in a fourth, and the edit at the end. There is another way: connect the chat to the studio over MCP and work against one conversation that remembers everything. Here is an 88-second trailer built that way, and a full working guide: how to connect, how to ask, how to approve a price, and what to keep inside that one conversation.
Most people use AI like a pile of separate tools. Write an idea in a chat. Move to an image tool. Then a video tool. Then a voice tool. Then music. Then an editor. Then go back, because the character in the fourth shot is no longer the person she was in the first, and none of the tools knows what happened in the other three.
This project was built differently. A whole film out of one conversation, where the chat does not only give instructions but actually runs the studio: builds the concept, writes the story, keeps the characters, generates video, attaches voices, generates music, inspects shots, fixes continuity and edits. The thing that makes that possible is called MCP, and the second half of this article is a full working guide to it. The result first.
Trailer · 88 seconds · built in one conversation
The trailer opens on eight seconds that look like the old film: saturated colour, a painted brick road, a green city on the horizon. Then it cuts to the present and never returns to that language. That is not indecision, it is the decision. The opening is the story as the grandmother told it, a childhood memory in the language of a fairy tale. Everything after it is what it looks like when it happens to you, today, on cracked Kansas asphalt.
That separation is also what holds up the rule the rest of the film was built on. Once a prologue takes all the fantasy onto itself, the modern world is completely free of it, and you can forbid it every glow and every particle without giving up the magic. The magic already happened, in the first eight seconds.
In the modern world: 95% reality. 5% impossible.
Which is to say, from the moment the film lands in the present: no fantasy world, no glittering effects, no AI look. A real woman, a real Kansas, a gas station, a cornfield, a garage, a warehouse, a dog. And exactly one impossible thing: the yellow road starts coming back.
The biggest mistake in AI video is opening with "cinematic woman walking". That did not happen here. A concept came first: a cultural icon everyone knows, without imitating the original. Dorothy was a real person. Everything she told the family was treated as a story for years. Then her granddaughter finds out the story never ended.
At this stage the chat was not a generator but a director: who the character is, what she wants, what she discovers, what changes, where the surprise sits, what the payoff is, and what stays with the viewer at the end. Only once the story worked did anything move to shots. That is a full shift in thinking. The AI does not have to invent a film while rendering. The film has to exist before the render.
Worth knowing: Planning the story in the chat is free. A video shot is not. Every minute spent turning an abstract idea into real beats is a minute you are not paying for in credits, and without it you pay twice for the same shot.
If a shot is beautiful and breaks the logic, it goes.
A bible like that is what stops every shot looking like it came from a different film. It gets written once at the start of the conversation, and the conversation holds it all the way through, so no shot is ever sent without it.
This is probably the single most important point in the whole project. If you write "Maya" in the third prompt, the engine does not know who that is. It invents a new woman who fits the description. So an asset library got built: Maya, young Dorothy, the dog, the house, the field, the garage, the warehouse, the shoes, the bricks, the 1955 photograph. Every shot was handed its anchors again, as pictures, not as names.
Instead of "Maya walks by a cornfield", every shot got a shot size, a lens, an aperture, one camera move, action split by the second, lighting, sound, and a list of prohibitions. All of it in English, because that is the language the engine reads best. This is what the camera instruction for the opening shot looked like:
MEDIUM CLOSE-UP, EYE LEVEL · 35mm f/4 · very slow handheld push-in · 5s · 0.0-1.5 Maya finishes opening an old keepsake box. 1.5-3.2 inside are a faded photograph of young Dorothy and a pair of worn red leather shoes. Her expression changes only when she sees the shoes. 3.2-5.0 she lifts one shoe toward the lens until it fills the frame. The final frame is built to cut straight into the same shoe on asphalt. · no glow · no particles · no dialogue · no music · do not open on black
Look at the last line of every shot. The forbidden list is not decoration. It is what stops the engine adding glow, particles, burned-in captions or music of its own, which are exactly the things that turn a good shot into a shot that looks generated.
And on the Kansas shots that prohibition includes the old film itself. The prologue was generated separately, deliberately, in its own language, and after it every shot was told explicitly not to drift back there. Without that the engine would pull the modern world toward the colourful one every time the word Dorothy appeared in a prompt, and that is exactly the kind of bleed that ruins a film halfway through.
The yellow road does not appear. It breaks asphalt. Concrete cracks, dust jumps, gravel flies, a wrench on the bench rattles, a fluorescent fixture swings, paper slides toward the door, the dog startles. Reality reacts to the event, and that is what sells the illusion. When an effect has no consequence in the world it looks like an effect. When it produces a physical result, it feels real.
EXTREME LOW CLOSE-UP · 24mm f/8 · slow push in · 4s · One old mustard-yellow brick punches upward through the concrete floor with grit and broken concrete. CLACK. A second brick rises ahead of it. Then a third. Each impact heavy and plausible, never magical. The line travels in one direction only. · never glowing · no magic particles · no random destruction
Instead of assembling blind, every shot got inspected: length, frames, what actually happens in it. Not what we asked for, what actually came out. That is how an illogical transition got caught, a character who suddenly turns up somewhere else, repetition, and a few beautiful shots that simply did not serve the story. Some never made the film, and that is fine. Not every shot you generate has to go in.
The biggest problem with AI-made trailers is that every clip arrives with music of its own, and then every cut feels like a different film. The rule here was blunt: no shot was generated with music. Only room tone, footsteps, gravel, metal, dog, wind, breath, thunder and speech. The music got built afterwards, as one score.
Worth knowing: A good shot with bad music is not a lost shot. You can separate the voice from the ambience, keep the speech and the effects, and replace only the music. It is one of the most useful tools in AI editing, and without it whole shots get thrown away over one layer.
Maya got one voice. The old man got one voice. Instead of every line sounding like a different person, the same voices came back through the whole film, and a separate trailer narration was written alongside them. The narration does not describe what you see. It gives meaning to what you see.
MAYA: My grandmother used to tell me stories about a road that could find you. I thought they were only stories... until it came back for me.
THE WIZARD: Dorothy thought she ended it. She only delayed it.
THE WIZARD: The road doesn't lead to Oz anymore. It leads back to you.
One of the secrets of a big trailer is that the music does not run as one layer under everything. It changes with the structure. Three cues got built here: the first mysterious, minimal and spacious. The second more rhythmic, with a sense of pursuit and discovery. The third a large emotional orchestral anthem. And most importantly: under dialogue the music drops almost to nothing. Not a little quieter. Genuinely out of the way, and then back.
When the photograph turns over and MAY 1955 appears, no music was added. It was taken away. A moment of near silence, which lets the brain register that something important just happened. A good trailer does not always get stronger by adding. Sometimes it gets stronger by removing.
So the edit does not feel like a pile of clips, speech sometimes starts before the shot it belongs to, or carries on after the picture has already moved. You hear "It's starting again" while the bricks are already on screen. You hear Maya before you see her. That is what makes a trailer feel like one continuous piece instead of a chain.
And transitions were added only when there was a reason. Most of the film runs on hard cuts, impact cuts, sound bridges, a flash frame, a black frame, and a music stop. A good transition does not say "look how I edited this". It makes the brain feel the cut was inevitable.
After the title you need one last moment, the one that makes a viewer think "wait, what?". Here it is the tornado. Title, silence, storm, black. That is the button shot, and it is the difference between a trailer and a summary of a film.
MCP is a standard that lets a chat operate an outside system directly. Think of it like a universal power socket: you connect once, and from that moment the chat does not merely advise you what to write, it actually generates. It knows your balance, your gallery, the characters you saved and the films you already made, so no action starts from a blank page.
That is exactly what made this project possible. One conversation held the production bible, the references and everything already generated, so shot number nine knew who Maya was without anyone describing her again.
IDEA → STORY → RULES → REFERENCES → SHOTS → PHYSICS → PERFORMANCE → SOUND → MUSIC → EDIT → REVIEW
PROMPT → GENERATE → HOPE
Worth knowing: No paid action happens without approval. Every request that costs money comes back with the price in credits first, and only runs once you approve it. If something fails and no result comes back, the credits return. The full guide, with the price table for every engine, lives at the cadabra-mcp page under the site guides.
This is the part people worry about most before they try it, and it is simple. You write what you want, in your own language. A price in credits comes back. You say yes or no. That is the whole thing, and it is the same loop whether you asked for one image or for a film.
Worth knowing: One approval covers one action. If you asked for a film of five shots, ask to see the running total and not only the price of each shot on its own. It is the most useful question to ask before you start: "what will this cost altogether".
The power is not that the chat can generate. It is that it remembers. One conversation that holds the production bible, the references and everything already made is the difference between a film and a pile of clips. In practice, these are the things worth putting into it at the start:
Worth knowing: Many files at once? On the site upload page, at mcp/upload, you upload up to 30 files and get one link that stands for all of them. Paste one link in the chat instead of thirty.
This is the most expensive beginner mistake: running straight to generation. Half the work on this trailer cost nothing, because it was writing, planning and checking. These are things you can ask for without paying:
The two requests that changed this film most did not generate anything. The first is an inspection: "look at the shots and tell me what is actually in them, how long they are, and where a transition does not work". That is what caught the repetition and the shot that turned up in the wrong place, before either reached the edit.
The second is editing a film that already exists: "take this cut, trim the two dead seconds at the top, drop the music under the dialogue, and add captions". No shot is generated again, the original stays whole, and what comes back is a new version of the same film, with a project you can open and edit by hand if you want to.
Worth knowing: Describe what is wrong, not how to fix it. "The old man's line is buried under the music" works better than "drop 6 dB at 0:47", because the second assumes you already measured, and the first asks for the measurement to be done for you.
The film has to exist before the render. Everything else is execution.
עמוד הבית · בלוג · וידאו AI · תמונות AI · קול AI · כל המנועים · קרדיטים ולא מנוי · AI בעברית · קורס מתחילים · מדריכים