A picture of a place and a video of a place are not the same thing. There is a way to build a place you can walk around in, pick an angle the way a cinematographer would, and only then turn that frame into a film. Two charges, not one, and that is exactly the point.
Most people ask an AI for "a room in an abandoned hotel at night" and get a pretty picture. Then they ask for a video from the same sentence, and get a completely different room, because the engine invented the place again. There is no angle you chose, no wall you came back to, and no feeling of being inside.
The fix is not a longer prompt. The fix is to build the place first, walk into it, save a frame from your camera, and only then ask a video engine to move that frame. Two steps. Two charges. That is not a pricing bug, it is the split between "building a set" and "shooting a shot".
You get a 3D world you can walk around in the browser: turn, get closer to a wall, look up, choose where to stand. That is not a video. That is a set. The gallery also shows a strange panoramic image, a 360 view that looks like a map rather than a frame. Do not animate it. It is not an opening frame.
Once the world is ready, the system also saves a camera still: a centre crop from the panorama, a flat frame a video engine knows how to read. If you went inside and saved an angle yourself, that is the picture that is worth more, because you chose it. If you did not, the default still is still better than the 360.
Worth knowing: 16 credits build the place. The video on top of it, on whichever engine you pick, is a second charge. If something fails, that step returns to the balance. Not both steps together.
Describe a space, not a montage. "An alley in Jaffa in the evening, wet stone walls, laundry overhead, a cat on a sill, one warm bulb in the distance" gives the model walls to stand between. "A cinematic mood of an old city" gives it nothing to stand on.
You can also upload a photo of a real place, and the world will sit on it. That helps when you already have a set: a shop, an apartment, a street. The interior does not have to match pixel for pixel, but the geography stays readable.
Stand at human height, not from above like a map. Get close to something: a handle, a window, a table. Save. That frame is what the video engine will receive. Now write a short move: "The camera eases toward the window, the curtain shifts in the wind." Do not ask the engine to invent the room again.
LTX-2.5 fits here because it moves a still with sound in the same pass, and the room is already locked. Seedance fits if you need face identity on that set. Kling fits if the presenter must stay the same person. The world does not replace the video engine. It hands the engine a frame that was not invented at the moment you clicked.
On Cadabra this is the studio called Worlds. Write a place in English, go inside, save a picture, and continue to a video engine. You can also do it from the Genie, from Cadabras, or from the marketing agency: the same camera still, the same separate charge for the film.
A set you can pick an angle from is worth more than a pretty picture the engine invented once and that you cannot return to.
עמוד הבית · בלוג · וידאו AI · תמונות AI · קול AI · כל המנועים · קרדיטים ולא מנוי · AI בעברית · קורס מתחילים · מדריכים