Most AI videos come out silent, then you paste a song on in an editor. There is another way: an engine that makes the picture and the sound together, up to 20 seconds, and you can also start from an audio file instead of a still.
The most common frustration after a first AI shot is the silence. The picture moves, the character smiles, the camera pushes in, then you put headphones on and hear nothing. So you open an editor, hunt for a free-to-use track, cut it to the shot length, and discover that the music tempo does not land on motion that was born with no relationship to it.
Some engines generate sound together with the picture. Not an effect you paste on later, but the same pass: wheels on the street, breath, the room, a song someone is singing in frame. When it works, the shot feels like a clip instead of a mute demo.
Not every shot needs this. If you are building a video with narration you write yourself, a clean shot then voice in the edit is better, because you control the words. If a character has to speak an exact Hebrew line, lip sync onto an audio file you made separately is more precise than sound the engine invented.
The simple way is text. Write what happens, how long, and what sound should be heard. "A kitchen in the morning, the coffee machine starts, rain on the window, 10 seconds" already gives the engine both a picture and a soundtrack. Without that last part it will invent sound on its own, and sometimes that is right and sometimes it is generic noise.
The second way is a still. Upload a frame, and the engine moves it with sound. This is the right path when you already have an identity that looks correct: a presenter, a product, a photographed place. The still locks the look, the prompt locks the motion and the sound.
The third way is a first frame and a last frame. Two stills, and the engine invents the transition between them, including what is heard in the middle. Useful when you know how the scene opens and how it closes, and you do not want the engine to wander in between.
The fourth way is audio. A song, narration, or room tone becomes a video, and the length is the length of the file, not the even-second ladder the engine likes for text. Two seconds minimum, up to twenty on the fast path. If there is also a still, it is used as the opening frame. Without a still the engine invents the whole scene from the sound.
Worth knowing: Audio-to-video is billed by the length of the file you uploaded, not by the number you rounded in the picker. A 7-second file costs 7 seconds, not 6 or 8. Check the length before you click.
Describe the sound as an event, not as a genre. "Cinematic music" tells the engine nothing. "A kettle whistles, then the door slams" does. Same rule as acting: not "she is happy", but "she laughs short, draws a breath, and sets the cup down".
An engine with built-in sound holds a face only from the still you handed it on that shot. It is not a trained-character engine. If you need the same person across ten shots, keep a character sheet on an engine built for that, and do not ask the sound engine to "remember" who this is. Hand it the frame, every time.
Built-in sound saves the paste-on step when the sound belongs to the scene itself. It does not replace narration you wrote, and it does not replace an engine that holds a face across a film. On Cadabra that engine is called LTX-2.5: up to 20 seconds on the fast 1080p path, from text or a still or audio, with the price on screen before you click. If the generation fails, the credits come back.
A shot whose sound was born with it stops looking like a demo. It starts looking like something someone would leave the headphones on for until the end.
עמוד הבית · בלוג · וידאו AI · תמונות AI · קול AI · כל המנועים · קרדיטים ולא מנוי · AI בעברית · קורס מתחילים · מדריכים