Frontier video direction

Text or image.21 ways to move.

Start from a prompt, animate one image, direct first and last frames, or guide a shot with references. Change the model—not your production process.

Every model animates images Setup stays ready until you generate Results saved to your library
Video Lab / Model stage
21 online
Loading Veo 3.1 sample
Official sampleGoogle logo

Sourced model output / Lens 01

Veo 3.1

Text promptAnimate imageFirst → lastUp to 3 reference imagesAnimate, first/last, and referencesUp to 4K outputOptional generated audio

Director note

Describe the subject, action, camera move, lighting, and sound as separate beats in one concise shot direction.

Sample credited to official sample. Results vary with prompt, settings, and provider updates.

Duration4, 6, or 8 secondsResolution720p, 1080p, or 4KFrame16:9 or 9:16AudioOptional generated audio

From the community

Published Video Lab takes.

Open a real generation, then Generate this in Video Lab.

A better control surface

Model choice without workflow chaos.

Every frontier model has a different motion language. Video Lab keeps prompts, source images, frame roles, format, duration, resolution, audio, and cost in one consistent directing room.

01

Choose the source

Begin with text, animate a frame, direct a transition, or use references when the selected model supports them.

02

Lock the setup

Set the model, duration, format, and controls before any generation is submitted.

03

Keep the take

Completed shots and their settings stay together in your history and Clip Studio Library.

The model slate

21 distinct creative lenses.

Compare every model in detail
Veo 3.1 sample outputOfficial sample01Google logoGoogleVeo 3.1A premium text-and-image video model for cinematic shots, guided frames, references, optional audio, and up to 4K output.Text promptAnimate imageFirst → lastUp to 3 reference imagesSee model detailsVeo 3.1 Fast sample outputOfficial sample02Google logoGoogleVeo 3.1 FastA faster, lower-cost Veo 3.1 SKU for cinematic shots, guided frames, references, optional audio, and up to 4K.Text promptAnimate imageFirst → lastUp to 3 reference imagesSee model detailsSeedance 2.5 sample outputOfficial sample03ByteDance logoByteDanceSeedance 2.5A next-generation text-and-image video model with native audio, 30-second shots, first/last frames, and deep reference control.Text promptAnimate imageFirst → lastUp to 30 reference imagesSee model detailsSeedance 2.0 sample outputIndependent test04ByteDance logoByteDanceSeedance 2.0A flexible text-and-image model with first/last-frame direction, up to nine references, broad formats, and optional audio.Text promptAnimate imageFirst → lastUp to 9 reference imagesSee model detailsKling 3 Pro sample outputOfficial sample05Kuaishou logoKuaishouKling 3 ProA controllable text-and-image model for expressive movement, first/last frames, subject references, and optional audio.Text promptAnimate imageFirst → lastUp to 3 subject elementsSee model detailsKling O3 sample outputOfficial sample06Kuaishou logoKuaishouKling O3Kling’s newer O3 Standard generation for realistic motion, optional audio, intelligent multi-shot, first-to-last frames, and references.Text promptAnimate imageFirst → lastUp to 4 reference imagesSee model detailsHappy Horse 1.1 sample outputOfficial sample07Alibaba logoAlibabaHappy Horse 1.1A native-audio text-and-image model with source-image animation, multi-image references, and broad social formats.Text promptAnimate imageUp to 9 reference imagesSee model detailsGemini Omni Flash sample outputIndependent test08Google logoGoogleGemini Omni FlashA fast multimodal 1.1 model for text, image animation, first/last frames, and references with native audio up to 4K.Text promptAnimate imageFirst → lastUp to 9 reference imagesSee model detailsWan 3.0 sample outputOfficial sample09Alibaba logoAlibabaWan 3.0A longer 1080p one-pass text-and-image model with optional native audio, first/last frames, and up to ten references.Text promptAnimate imageFirst → lastUp to 10 reference imagesSee model detailsWan 3.0 Prime sample outputOfficial sample10Alibaba logoAlibabaWan 3.0 PrimeA faster Wan 3.0 SKU for 2–30 second 1080p shots with optional native audio, first/last frames, and up to ten references.Text promptAnimate imageFirst → lastUp to 10 reference imagesSee model detailsWan 2.7 sample outputOfficial sample11Alibaba logoAlibabaWan 2.7A practical text-and-image generator with first/last frames, multi-image references, 1080p, and repeatable controls.Text promptAnimate imageFirst → lastUp to 9 reference imagesSee model detailsMiniMax H3 Max sample outputOfficial sample12MiniMax logoMiniMaxMiniMax H3 Maxfal’s post-trained MiniMax H3 variant for native-audio clips, image animation, first/last frames, and references at 480p, 768p, or 1080p.Text promptAnimate imageFirst → lastUp to 9 reference imagesSee model detailsMiniMax H3 sample outputOfficial sample13MiniMax logoMiniMaxMiniMax H3A high-detail native-audio model with first/last frames, references, and 480p through 4K output.Text promptAnimate imageFirst → lastUp to 9 reference imagesSee model detailsGrok Imagine Video 1.5 sample outputIndependent test14xAI logoxAIGrok Imagine Video 1.5A flexible native-audio model for text, single-image animation, and multi-reference clips from 480p to 1080p.Text promptAnimate imageUp to 7 reference imagesSee model detailsFLUX 3 sample outputOfficial sample15Black Forest Labs logoBlack Forest LabsFLUX 3Black Forest Labs’ frontier video model for text, single-image motion, first-to-last frames, optional native audio, and clips up to 20 seconds.Text promptAnimate imageFirst → lastSee model detailsPixVerse V6 sample outputOfficial sample16PixVerse logoPixVersePixVerse V6A broad text-and-image toolkit with image animation, first/last transitions, four resolutions, and style controls.Text promptAnimate imageFirst → lastSee model detailsLTX 2.3 Pro sample outputOfficial sample17Lightricks logoLightricksLTX 2.3 ProA production-oriented text-and-image model with first/last frames, 1080p to 4K output, and frame-rate controls.Text promptAnimate imageFirst → lastSee model detailsKling O3 Pro sample outputOfficial sample18Kuaishou logoKuaishouKling O3 ProKling O3 Pro for realistic motion, optional audio, intelligent multi-shot, first-to-last frames, and references.Text promptAnimate imageFirst → lastUp to 4 reference imagesSee model detailsLTX 2.5 Pro sample outputOfficial sample19Lightricks logoLightricksLTX 2.5 ProLightricks’ newer Pro SKU for 720p and 1080p clips with first/last frames, optional audio, and 24, 25, or 50 fps.Text promptAnimate imageFirst → lastSee model detailsSora 2 sample outputOfficial sample20OpenAI logoOpenAISora 2OpenAI Sora 2 for 720p clips with always-on native audio from text or a starting image.Text promptAnimate imageSee model detailsSora 2 Pro sample outputOfficial sample21OpenAI logoOpenAISora 2 ProOpenAI Sora 2 Pro for 720p or true 1080p clips with always-on native audio from text or a starting image.Text promptAnimate imageSee model details

Editorial selection · Last reviewed 11 September 2026

Before the first take

Questions from the directing room.

What is Clip Studio Video Lab?

Video Lab is one workspace for generating videos with text prompts or source images across 21 curated frontier models. Choose a model, add a prompt or image, generate, and keep the completed clip in your Clip Studio Library.

Are these the 21 most popular AI video models?

No objective cross-provider popularity ranking exists. This is an editorial selection of 21 frontier video models chosen for useful differences in quality, speed, control, format support, audio, and resolution.

How is Video Lab priced?

Each plan includes a monthly generation limit. Model, duration, resolution, and other settings use that allowance. Upgrade anytime for a higher monthly limit.

Can I compare the same prompt across models?

Yes. Generate with one model at a time, then use Try another model to preserve your prompt, compatible settings, and available source images while you select a different model.

Which image-to-video workflows are supported?

Every Video Lab model can animate a starting image. Veo, Seedance, Kling, Wan, MiniMax, Gemini, PixVerse, LTX, and FLUX 3 also support first-to-last-frame direction, while Veo, Seedance, Happy Horse, Gemini, Wan, MiniMax H3, MiniMax H3 Max, Kling O3, Kling O3 Pro, and Grok offer multi-image reference workflows. Kling 3 Pro uses subject reference elements inside its image workflow. Sora 2 and Sora 2 Pro animate a starting image; they do not take first-to-last frames or references.

Bring the brief.Choose the motion.

Bring a prompt or source image, choose the right model, and generate.

Enter Video Lab