Google Veo: AI Video Generation Family Guide
Google DeepMind Veo family from Veo 1 to Veo 3.1: how each version raised the bar, which model to use, and where to access it.
Before reading, try Veo 3.1 Fast, the current balanced tier.
Developer: Google DeepMind · Family launched: May 2024 · Current flagship: Veo 3.1 · Official page: Google DeepMind Veo
What Is Google Veo?
Announced at Google I/O in May 2024, Google Veo is a family of generative AI video models developed by DeepMind. The family has evolved through four major generations, with each iteration solving a specific production bottleneck. This progression moved the technology from a silent high-definition experiment (Veo 1) to production-grade 4K resolution (Veo 2), then to native audiovisual generation (Veo 3), and finally to a tiered production platform (Veo 3.1). This evolution reflects a systematic approach to overcoming the core limitations of generative video for professional workflows.
The Veo Timeline: Four Problems, Four Solutions
| Version | Released | Max Resolution | Native Audio | Core Achievement |
|---|---|---|---|---|
| Veo 1 | May 2024 | 1080p | No | Proved HD AI video generation was possible; experimental VideoFX access |
| Veo 2 | December 2024 | 4K | No | Production-grade resolution and physics engine; first commercial pricing |
| Veo 3 | May 2025 | 1080p | Yes | First model to generate video and audio in a single pass |
| Veo 3.1 | October 2025 | Up to 4K | Yes | Three-tier production platform with scene extension and frame control |
This timeline represents more than iterative improvements. Each Veo generation specifically addressed a task category that its predecessor could not perform, expanding the commercial viability of AI video at every step.
How Veo Works: Architecture Overview
Veo models are built on a latent diffusion transformer architecture. Rather than generating video frames directly in pixel space, the model works in a compressed latent representation, encoding visual information into a lower-dimensional space, performing the generation process there, and then decoding back to full-resolution video. This approach makes high-resolution generation computationally tractable and allows the model to reason about temporal coherence across frames more effectively than pixel-space methods.
The architecture processes text prompts, optional reference images, and temporal conditioning signals together as a unified input. This is why Veo models respond well to cinematographic language: the training data and architecture are designed to map concepts like "shallow depth of field," "slow pan," or "golden hour lighting" onto consistent visual patterns. The model doesn't interpret these as decoration; it treats them as spatial and temporal constraints on the generation.
Veo 3 and Veo 3.1 extend this foundation to joint audio-visual generation. The audio and visual streams are generated together using a shared latent representation, which is what makes the synchronization between dialogue and lip movement coherent rather than post-hoc.
How Veo 3.x Generates Audio and Video Together
Beginning with Veo 3, the Veo family shifted from producing silent video to generating audio and visuals in a single pass. This is architecturally distinct from post-hoc audio layering: the audio and visual streams share the same latent representation during generation, so synchronization between lip movement, dialogue, and ambient sound is structural rather than assembled in post-production.
Veo 1 and Veo 2 produced silent video only. The joint audio-visual architecture was introduced in Veo 3 and carried forward into all Veo 3.1 tiers. For production workflows, capability details, and tier-specific audio quality differences, see Veo 3 and Veo 3.1.
SynthID: AI Watermarking Across the Veo Family
Every video generated by any Veo model includes a SynthID watermark. SynthID is Google DeepMind's system for embedding an invisible, cryptographically robust provenance marker into AI-generated content.
The watermark is embedded during generation, not added in post-processing. It survives common editing operations including compression, color grading, and format conversion, and is not perceptible during normal playback. It cannot be removed without degrading the video in detectable ways.
For creators working in commercial or professional contexts, the practical implication is straightforward: any video generated through Veo, including on third-party platforms that access the Veo API, carries SynthID metadata that identifies it as AI-generated. This applies to all four Veo generations, not only Veo 3.x.
Veo 1: The First High-Definition Proof (May 2024)
Introduced at Google I/O 2024, Veo 1 was the foundational release for the model family. Built on an advanced latent diffusion architecture, it was the first Google model to generate high-quality 1080p video capable of exceeding 60 seconds in duration. Access was initially restricted to the experimental VideoFX platform behind a waitlist, limiting its early commercial adoption but providing essential proof of concept for the technology.
Veo 1 mattered because it proved high-definition AI video generation was technically feasible. The model demonstrated a deep understanding of cinematic terminology, reliably interpreting prompts for time-lapse sequences, sweeping pan movements, and aerial shots. However, its purely visual nature defined the constraints that future generations needed to overcome: no audio, no commercial API, and no production-grade resolution.
Veo 2: Resolution and Physics (December 2024)
Released on December 16, 2024, Veo 2 moved the family from experimental demonstration to professional production. The model upgraded the maximum resolution to 4K and extended generation duration past two minutes. This release introduced a significantly upgraded physics engine, which addressed early weaknesses in fluid dynamics, lighting consistency, and shadow tracking. Veo 2 also integrated SynthID watermarking from launch and established the first commercial pricing baseline for the family.
The introduction of 4K output was the inflection point that made AI video usable in professional display contexts. Despite these advances, the model produced completely silent video. This constraint meant Veo 2 still could not natively serve content workflows requiring integrated sound design, which defined what Veo 3 had to solve.
Veo 3: The Audiovisual Leap (May 2025)
Google announced Veo 3 at Google I/O 2025, introducing native audio-visual joint generation. Operating at 1080p and 24 frames per second, Veo 3 was the first model to generate dialogue, ambient sound, and sound effects as part of the same generation pass as the visuals, with no separate synchronization step required.
Google integrated Veo 3 directly into the Gemini API and Google AI Studio, making it the first Veo version with broad developer access at launch. The Veo 3 standalone API was retired in June 2026, fully replaced by Veo 3.1.
Veo 3.1: The Production Platform (October 2025–Present)
Launched in October 2025, Veo 3.1 transformed the single-model concept into a comprehensive production platform with a three-tier architecture: Lite, Fast, and Quality. A January 2026 update restored 4K resolution support. The platform added Scene Extension for building longer sequences, Ingredients to Video for reference-image-guided consistency, and First and Last Frame Control for precise sequence editing. The Lite tier launched on March 31, 2026, completing the rollout.
The tiered structure is the defining architectural shift of Veo 3.1. Rather than offering one capability level, it allows teams to use Lite for rapid high-volume iteration, Fast for most production work, and Quality for final broadcast deliverables, each at a different cost point.
Cross-Version Comparison
| Feature | Veo 1 | Veo 2 | Veo 3 | Veo 3.1 |
|---|---|---|---|---|
| Max Resolution | 1080p | 4K | 1080p | Up to 4K |
| Native Audio | No | No | Yes | Yes |
| Max Duration | 60s+ | 2 min+ | 8 sec base | 8 sec + Scene Extension |
| SynthID Watermark | Yes | Yes | Yes | Yes |
| Pricing tiers | No | No | No | Lite / Fast / Quality |
| Current status | Retired | Retired | Retired | Current |
Source: Google DeepMind official documentation · as of July 2026
Which Veo Model Should You Use?
The decision is simple given current availability:
- New projects: Use Veo 3.1. It is the only active version.
- Within Veo 3.1: Select Lite for high-volume generation and rapid iteration, Fast for most standard production work, and Quality for final broadcast deliverables and large-format output.
- Historical reference: If you need to understand an older version for migration or archival purposes, consult the Veo 2 model guide or the Veo 3 and Veo 3.1 tier guide.
Where to Access Veo
- Gemini Advanced: Consumer-friendly generation interface with a usage quota.
- Google AI Studio: Developer prototyping and prompt engineering; limited free-tier access.
- Vertex AI: Enterprise-scale integration with pay-per-second billing.
- Google Flow: Specialized filmmaking tool built on Veo.
- VioEvo: Direct access to all three Veo 3.1 tiers (Lite, Fast, and Quality) without requiring a separate Google Cloud account.