Listen to this lesson
Unlock audio and more
Audio streaming, downloadable PDFs and certificates come with Plus and Pro.
What it means
Text-to-video extends generative imaging across time, and the added dimension is genuinely difficult. Every frame must be plausible individually and consistent with its neighbors: objects keep their identity, lighting stays coherent, motion obeys something like physics.
Progress has been rapid, with clip length, resolution and temporal stability all improving, and synchronized audio generation now appearing alongside. Output remains short-form, and the reliable workflow is generating many candidates and selecting, rather than directing a specific shot.
Compute cost is the practical constraint — video generation is dramatically more expensive per second of output than image generation, which shapes both pricing and how much iteration is affordable.
Why it matters
Video is the most expensive media format to produce conventionally, so the cost collapse has the largest proportional effect in video. It also raises the deepfake stakes: convincing synthetic video of real people is now achievable without specialist skill.
In practice
Budget for iteration — you will generate many clips per usable one. Treat it as a source of B-roll, concepts and short segments rather than as a way to direct a specific scene.
Where this shows up
Tools and models in our catalog.
Veo 3Google DeepMind's video generation models. Veo 3 generates synchronized audio with video. Veo 3.1 (March 2026) upgrades to 4K resolution and 60-second clips. Available via Google AI Studio and Vertex AI.
Runway MLProfessional AI video creation platform used by major studios. Flagship Gen-4.5 (December 2025) tops independent text-to-video rankings, alongside Gen-4 Turbo for fast iteration, Aleph for in-context editing of real footage, and Act-Two for performance capture.
SoraDiscontinued March 2026. Was OpenAI's text-to-video model with cinematic quality output, storyboarding, remix, and synchronized audio. iOS app, API, and sora.com all shutting down.
Kling AIHigh-quality video generation from Kuaishou (China). Strong physics simulation and character motion. Competitive quality at lower cost than US alternatives.
Dream MachineLuma AI's popular text-to-video and image-to-video generation tool. High-quality cinematic output at accessible pricing. Strong community and rapid iteration.