GoogleGoogle video model

Veo 3.1. Cinema-grade video with native sound.

Google's flagship video model turns a sentence into a finished shot, with dialogue, ambience, and effects generated in sync. Run it in Treza chat or drop it into a pipeline and publish it as an API.

  • 4 to 8 second clips
  • 16:9 and 9:16
  • Native audio
  • First and last frame control
  • 1080p output
The model

Veo 3.1 is the model to reach for when the shot has to look expensive. It renders film-like lighting, believable physics, and camera moves you can direct in plain language: dolly in, rack focus, handheld, aerial. Audio is generated with the picture, so a cafe scene comes back with the espresso machine hissing and the room murmuring behind it.

On Treza, Veo 3.1 is one node on a canvas. Draft the prompt with a language model, generate the shot, then pass it to the editor or straight to your product through a published endpoint. You pay per successful generation from prepaid credits, with no subscription.

Native audio, in sync

Dialogue, ambience, and sound effects are generated with the frames, not layered on after. Characters speak on camera and the room sounds like the room looks.

Direct the camera in words

Veo understands lens and movement language. Ask for a slow push-in, a whip pan, or a top-down drone orbit and the shot follows the direction.

First and last frame control

Pin the opening and closing frames with images and Veo fills the motion between them. Ideal for brand-locked openers and product reveals.

Physics that hold up

Liquids pour, fabric drapes, and reflections track the camera. Fewer uncanny artifacts means fewer retries on hero shots.

Prompt to result

What to make with Veo 3.1

Real prompts, and what comes back. Every one runs in Treza chat or as a node on the canvas.

Slow dolly-in on a barista pouring latte art, golden morning light, quiet cafe ambience

A cinematic 8 second shot with synced pour and room tone.

A founder says 'we ship on Fridays' to camera in a bright loft office

On-camera dialogue with matching lip sync and office ambience.

Aerial drone shot rising over a fjord at sunrise, mist in the valley

A sweeping establishing shot with wind and water audio.

Macro shot of rain hitting a watch face, product stays sharp, shallow depth of field

A commercial-grade product close-up with rainfall sound.

Vertical 9:16 clip of sneakers splashing through neon-lit puddles at night

A social-ready portrait cut with footsteps and city hum.

Start on this product photo, end on this lifestyle photo, smooth transition

A keyframed shot that lands exactly on your closing frame.

Beyond one prompt

Veo 3.1 is one node on a canvas

Chain it with language, image, and video models, then publish the whole pipeline as an API.

  1. Step 01

    Draft the brief

    A language model node expands a one-line idea into a detailed, on-brand prompt before it ever reaches the model.

  2. Step 02

    Generate with Veo 3.1

    Run variants in parallel, compare against other models on the same prompt, and keep every result in run history.

  3. Step 03

    Publish it as an API

    The pipeline becomes a versioned endpoint. Your product calls it with the typed /invoke API or any OpenAI SDK.

FAQ

Veo 3.1, answered

Does Veo 3.1 generate audio?

Yes. Veo 3.1 generates dialogue, ambience, and sound effects in sync with the picture. You can also turn audio off per generation if you plan to score the clip yourself.

What is the difference between Veo 3.1 and Veo 3.1 Fast?

Same family, different trade-off. Veo 3.1 maximizes image quality for final shots. Veo 3.1 Fast returns clips quicker at a lower credit cost, which makes it the better loop for drafts, explorations, and batch runs. A common pipeline drafts on Fast and re-renders the winning prompt on Veo 3.1.

Can I control how a Veo 3.1 clip starts and ends?

Yes. Veo 3.1 accepts a first frame and a last frame image, so you can bookend a shot with exact product frames or brand art and let the model generate the motion between them.

How much does Veo 3.1 cost on Treza?

Treza uses prepaid credits with no subscription. Each clip is charged at Veo 3.1's rate from your balance, and only successful generations are charged. Credit packs start at $5 and never expire.

Can I call Veo 3.1 through an API?

Yes. Build a pipeline that uses Veo 3.1 and publish it as a versioned HTTP endpoint. Call the typed /invoke API from your product, or point any OpenAI SDK at the OpenAI-compatible endpoint. If a better model ships later, swap it on the node without changing the API your product calls.