Create and edit realistic videos from text, images, or references with synchronized audio.
Choose whether to generate a new video or edit an existing clip.
Describe the scene, motion, sound, and action to create, or the edit instruction to apply. Maximum 20,000 characters.
Choose output length for generated videos. Token cost updates automatically.
Sign in to spend tokens
Prompt is required
Your generated video will appear here
Gemini Omni Flash creates and edits realistic videos from text, images, or reference images — with synchronized audio generated alongside the visuals. Prizmad ships it as a built-in tool: pick one of four modes, write a prompt, generate.
Gemini Omni Flash is a video model that generates realistic clips with synchronized audio — ambient sound and scene audio are produced together with the picture, not bolted on afterwards. For ad work that means a UGC-style bathroom scene can come with water sounds, and a studio product reveal can come with a clean audio ambience, all from one prompt.
It runs in four modes on Prizmad. Text to Video generates a clip from a prompt alone. Image to Video animates a starting image — a product photo becomes a moving lifestyle shot. Reference to Video accepts 1–10 reference images you can address directly in the prompt as <IMAGE_REF_0>, <IMAGE_REF_1>, and so on, which is how you lock product identity, characters, scenes, materials, or brand details into the output. Edit Video takes an uploaded 3–10 second clip (MP4, MOV, or M4V up to 90 MB) and applies your edit instruction — though voice editing is not supported. Generated clips run 3 to 10 seconds in 16:9 or 9:16, and prompts can be long: up to 20,000 characters.
On Prizmad, Gemini Omni Flash uses the as , avatars, voiceover, captions, and music. Generate a product image, animate it here with synchronized sound, then — all from one workspace and token balance, with full commercial rights to the output.
Vertical, natural-feeling ads — handheld camera, bathroom or kitchen scenes, product close-ups — with matching ambient sound generated in the same pass.
Turn a static supplier or studio photo into a realistic lifestyle clip: slow push-in, subtle hand interaction, warm daylight, premium audio ambience.
Reference to Video uses your actual product and packaging shots so the launch video shows your real product, not a lookalike.
Edit Video applies a described change to an existing 3–10 second clip instead of reshooting or regenerating from scratch.
Generate the same concept in 16:9 for in-feed and YouTube placements and 9:16 for Stories, Reels, and TikTok-style placements.
Videos come with matching audio — ambient sound and scene atmosphere generated together with the visuals, described right in your prompt.
Text to Video, Image to Video, Reference to Video, and Edit Video cover the full range from blank-page generation to editing an existing clip.
Upload 1–10 references and address them in the prompt as <IMAGE_REF_0>, <IMAGE_REF_1>, etc. — for product identity, characters, scenes, style, materials, or brand details.
Upload a 3–10 second MP4, MOV, or M4V clip (up to 90 MB) and describe the change you want. Voice editing is not supported.
Prompts up to 20,000 characters let you specify scene, motion, sound, and action in detail instead of compressing everything into one sentence.
Generated clips run 3 to 10 seconds in 16:9 or 9:16 — matching in-feed, Stories, and Reels placements.
Choose a mode: Text to Video for pure prompt generation, Image to Video to animate a photo, Reference to Video to lock in products or characters from references, or Edit Video to modify an existing clip.
Write your prompt describing scene, motion, camera work, and sound — for example, "soft synchronized water sounds" or "clean premium audio ambience". You have up to 20,000 characters.
For Image to Video, upload the starting image (JPG, PNG, WebP, GIF, or AVIF up to 10 MB). For Reference to Video, upload 1–10 references and refer to them in the prompt as <IMAGE_REF_0>, <IMAGE_REF_1>, and so on.
For Edit Video, upload a 3–10 second MP4, MOV, or M4V clip up to 90 MB and describe the edit. Keep in mind voice editing is not supported, and token cost follows the verified clip duration.
For generation modes, pick 16:9 or 9:16 and a duration from 3 to 10 seconds — the token cost updates automatically as you change the length.
Click Generate. Results typically arrive in about 3–8 minutes and land in your Prizmad asset library, ready to download or send into the captions tools.
Token cost scales with the length of the clip — the default 8-second generation costs 50 tokens, and the price updates automatically as you change the duration (3–10 seconds). In Edit Video mode, cost is based on the verified duration of your uploaded clip.
Gemini Omni Flash is unlocked on the same token balance you use for AI images, avatars, voiceover, and captions — no separate Gemini account, no API key setup, no extra billing.
If your tokens run dry mid-campaign, buy a one-off top-up directly from the top-up page. Generated videos stay yours with full commercial rights.
It's available to any account with enough tokens rather than offered as a standalone free tool. Generations cost tokens based on clip duration — the default 8-second clip is 50 tokens — drawn from your token balance, with one-off top-ups available.
Cost depends on duration: the default 8-second generation costs 50 tokens, and the exact price is shown in the tool before you generate, updating automatically as you change the length. Edit Video mode bills by the verified duration of the clip you upload.
Yes — synchronized audio is generated together with the visuals. Describe the sound you want in the prompt (ambient water sounds, studio ambience, environmental audio) and it's produced as part of the same generation.
Yes. Videos generated with Gemini Omni Flash on Prizmad are yours to use commercially — in paid ads, on product pages, and across social — with no extra licensing. Make sure you hold the rights to any images or clips you upload as inputs.
In Reference to Video mode you upload 1–10 images (JPG, PNG, WebP, GIF, or AVIF) and refer to them inside the prompt as <IMAGE_REF_0>, <IMAGE_REF_1>, and so on. That tells the model exactly which reference carries the product, character, scene, style, or brand details you want in the video.
Yes. Edit Video mode accepts an uploaded 3–10 second MP4, MOV, or M4V clip up to 90 MB and applies your written edit instruction. Voice editing is not supported, so it won't change spoken audio in the clip.
Generated clips run from 3 to 10 seconds in 16:9 or 9:16. Output is video, delivered to your Prizmad asset library, typically within about 3–8 minutes.
Grok Imagine Video 1.5 is image-to-video only, runs up to 15 seconds, and does not generate synchronized audio as a stated capability — it's the lighter, cheaper option for animating a single reference image. Gemini Omni Flash adds text-to-video, multi-reference generation, video editing, and synchronized audio. Both are on the same token balance.
Published 2026-07-16 · Last updated 2026-07-16