Text to video
Describe a scene in words and the AI generates a video clip. Examples:- “A drone shot flying over a mountain range at sunrise”
- “A cat walking across a piano keyboard, close-up”
- “Abstract liquid metal flowing and morphing shapes, slow motion”
First and last frame
Define the start and end of a shot by providing two images. The AI generates the transition between them — creating smooth camera movement, character action, or environmental change. To use this:- Place two images on the canvas — your start frame and your end frame
- Tell the AI: “Create a video transitioning from [first image] to [last image]”
- Scene transitions (day to night, empty room to furnished)
- Controlled camera movement between two compositions
- Character action (standing to sitting, calm to surprised)
- Environmental changes (dry landscape to rainy, calm sea to storm)
Reference video
Provide 1-3 reference images and a text prompt — the AI generates video that maintains visual consistency with your references. This is ideal when you need characters, environments, or props to look the same across multiple shots. Examples:- “Generate a shot of this character walking through a city at night” (with character reference image selected)
- “Create a scene in this environment with camera slowly pushing in” (with environment reference)
- “Animate this character turning to face the camera” (with character reference)
Output format
Generated videos are delivered as MP4 files.Downloading videos
Click any video on the canvas and use the Download button in the toolbar to save it as an MP4 file.Tips for better results
- Be specific about camera work — specify dolly, pan, tilt, rack focus, or establishing shot
- Describe lighting and mood — “warm golden hour” or “harsh overhead fluorescent”
- Use reference video when you need visual consistency across shots
- First and last frame gives you the most control over camera movement
- Start with images — generate a still image you like, then use it as a reference or keyframe for video
Credit costs
Video generation uses more credits than image generation because of the computational complexity involved.

