HeyGen Video API and Pricing, Explained Through a Real Pipeline

HeyGen Video is HeyGen's general-purpose video model, not its avatar product, and it carries the lowest listed per-second rate in our video catalog. Here is the rate card, the one input that doubles it, what the endpoint takes and leaves out, and how to reach it as a versioned API endpoint.

Alex Daro
Alex Daro
HeyGen Video API and Pricing, Explained Through a Real Pipeline

HeyGen Video is HeyGen's general-purpose video model, and it is not the avatar product most people know the name for. It renders scenes, products, and motion from a text prompt, an exact first frame, or a set of reference images, in clips of 5 to 15 seconds. It reached OpenRouter at the end of September 2026, and at 480p it carries the lowest listed per-second rate in our video catalog.

This post covers what it costs at each resolution, the one input that doubles the rate, what the endpoint accepts and what it leaves out, and how to call HeyGen Video as a published API endpoint. Every parameter and price below is read off Treza's live video model catalog as of October 2026, and every behavior described was checked with real renders before we put the model on the canvas.

What you get before you talk about price

CapabilityHeyGen Video
Clip length5 to 15 seconds, every whole second
Resolution480p or 768p, 768p when none is set
Aspect ratios21:9, 16:9, 4:3, 1:1, 3:4, 9:16
AudioNo audio setting; our renders carried a faint ambient track
KeyframesFirst frame only
Reference imagesYes

Two rows stand out. Six aspect ratios, including 21:9 and 4:3, means the shape of a placement is a dropdown rather than a crop. And reference images are supported, which is still uncommon at this price.

The other notable property is speed. In our launch-week test renders, five second clips came back in under 20 seconds, so a prompt can be iterated on in the time a slower model renders once. Capability detail lives on the HeyGen Video model page.

How HeyGen Video bills

Per second of output, at a rate set by the resolution and by whether reference images are attached:

Input480p, per second768p, per second
Prompt, or prompt plus a first frame$0.02$0.03
With reference images$0.04$0.06

A 5 second draft at 480p is 10 cents of provider cost. A 10 second clip at 768p is 30 cents. The most expensive clip the model can render, 15 seconds at 768p with references, is 90 cents.

For scale, Wan 3.0 and Grok Imagine both start at 5 cents a second at 480p, and Veo 3.1 Lite starts at 3 cents a second for a silent clip at 720p. HeyGen Video at 480p is 2 cents.

Reference images double the rate

The rate doubles the moment reference images ride along with the request, at both resolutions. A first frame does not trigger it: a clip that opens on your still bills at the plain rate.

That makes the choice between a first frame and references a pricing decision as well as a creative one. A first frame pins the opening image exactly, and the clip animates out of it. References carry a subject into a scene it was never photographed in. In our test, one product photo of a folded linen shirt on an oak table came back as the same shirt on white marble under different light. If the clip should open on a shot you already have, use the first frame and pay the plain rate. Use references when the subject has to appear somewhere new.

On the Video Generation node the two are separate inputs, and the model takes one or the other per clip. Wire both and the node uses the first frame and notes that the references were set aside, so you never pay the reference rate for images the model did not use.

Resolution: 480p for drafts, 768p for delivery

768p is both the ceiling and the default. Leave Resolution on Model default and a 16:9 clip comes back at 1344 by 768. 480p costs two thirds as much and suits drafts, variants, and anything watched small.

Neither is 1080p. For phone-first vertical work that rarely matters. For a clip shown large, run it through the Video Upscale node, or render that shot on a model that goes higher, such as Veo 3.1 Lite at up to 1080p.

What the endpoint will and will not take

On the endpoint we run, you send a prompt, a duration from 5 to 15 seconds, an aspect ratio, a resolution, and optionally a first frame or reference images. A clip with a first frame animates out of that exact still, the technique the first frame post walks through and the image to video generator covers from the still's side.

It does not take a last frame, so a shot cannot be pinned at both ends. For opening on brand art and landing on a product shot, Kling 3.0 Pro and the Veo 3.1 family take both keyframes. It does not render below 5 seconds, so a 2 second logo sting belongs on a model with a shorter floor, such as Grok Imagine at 1 second.

About the audio

The catalog's description of the model mentions synthesized dialogue, ambience, and sound effects, but it exposes no audio setting, and the clips we rendered carried only a faint ambient track, far below normal listening level. We have not tested prompts that ask for dialogue. Until a render shows otherwise, treat HeyGen Video as picture.

If a clip needs sound, add it in the graph. A Text to Speech node handles narration, an Audio Generation node handles music or effects, and the Sequence node mixes them over the picture and can duck the clip's own audio under a voice.

Not the avatar model

HeyGen is best known for talking-head avatars, and that is a different product. HeyGen's photo-to-talking-head model, Avatar IV, is listed on OpenRouter separately, and the Video Generation node does not run it, because it takes a photo and a voice track rather than a prompt. On Treza, a talking head is built from a clip plus the Lipsync node, which the lip sync and motion transfer post covers.

What you actually pay on Treza

HeyGen Video through a Treza pipeline is metered from a prepaid credit balance. One credit equals $0.01, credit packs start at $5, credits never expire, and a failed generation charges nothing.

Before a run starts, the platform estimates the media cost of the graph and refuses to start one your balance cannot cover. For HeyGen Video the estimator holds the reference rate for the resolution you pick: 6 cents of provider cost per second at 768p or with no resolution set, and 4 cents at 480p. A clip made from a prompt or a first frame settles below that hold, and the difference returns to the balance when the run settles. Reserves are holds, not prices.

For a finished video costed end to end, see the AI video cost breakdown.

Calling HeyGen Video as an API without building the integration

Direct API access to a video model means credentials, a submit-and-poll job pattern, retries, moderation rejections, and somewhere to store the output.

On a Treza AI video pipeline, HeyGen Video is a node on a canvas. Set the duration, aspect ratio, and resolution, wire it next to a script step, a voiceover step, or a captioning step, and publish the whole chain as one versioned endpoint. Your application calls a typed /invoke route with JSON in and JSON out, or points an existing OpenAI SDK at the OpenAI-compatible route.

The swap is what pays off later. When the next model wins on price or quality, you change the dropdown and republish, and the endpoint your product calls does not change. Our hosted MCP server exposes the same pipelines as tools, so an agent in Claude or ChatGPT can render a brief and return the link.

Where HeyGen Video earns its place in a graph

Three jobs where price and speed matter more than resolution. Draft passes, where framings render at 480p for 2 cents a second and the keeper goes to a higher-resolution model with the same prompt. Variant batches, where the node renders up to eight takes of one prompt, so eight 5 second versions of an ad hook at 480p come to 80 cents of provider cost. And short vertical cuts for Shorts and Reels, which are watched on a phone.

The fastest way to learn what it costs for your own work is to run your real prompt at your real duration and read the charge. Open the AI video generator, pick HeyGen Video, and generate once.

Frequently Asked Questions

How much does HeyGen Video cost per second?

$0.02 per second at 480p and $0.03 at 768p, and twice that when reference images are attached: $0.04 and $0.06. Those are the listed rates for the endpoint we run as of October 2026. A first frame bills at the plain rate. On Treza the charge comes from a prepaid credit balance, and a failed generation charges nothing.

Is HeyGen Video the same as HeyGen's avatars?

No. HeyGen Video is HeyGen's general-purpose generator for scenes, products, and motion. Avatar IV, the photo-to-talking-head model, is a separate model that the Video Generation node does not run. For a talking head on Treza, pair a clip with the Lipsync node.

What resolution does HeyGen Video render?

480p or 768p. With no resolution set it renders 768p, which is 1344 by 768 at 16:9. It does not render 1080p, so for a clip shown large, upscale it with the Video Upscale node or render that shot on a model that goes higher.

How long can a HeyGen Video clip be?

5 to 15 seconds in one generation, at any whole second. For longer pieces, stitch generations with the Sequence node, which handles up to 12 shots, or assemble them in the timeline editor.

Does HeyGen Video generate audio?

There is no audio setting, and the clips we rendered carried only a faint ambient track. Plan sound in the graph with Text to Speech or Audio Generation, mixed in the Sequence node.

Can I call HeyGen Video through an API without a HeyGen account?

Yes. On a Treza pipeline, HeyGen Video is a node you configure once and publish as a versioned HTTP endpoint, with a typed /invoke route and an OpenAI-compatible route. A hosted MCP server exposes the same pipelines as tools for agents. Generations are metered from prepaid credits that start at $5 and never expire.