AI video 2026 — what a clip really costs after the prompt
I make commercials that come entirely out of a computer. And the question I hear most often is the wrong one.
It goes: “What does a second cost?” As if that were the bill. It isn’t. The per-second price that fills every comparison table is the smallest line item. The real cost sits elsewhere — and nobody quotes it, because it doesn’t fit neatly into a column.
Let’s start with the per-second price anyway. Just to show how little it explains.
What a second nominally costs
| Model | Price (net list, USD) | Clip length | Audio | Commercial? |
|---|---|---|---|---|
| Kling 2.6 Pro Kling AI | $0.07/s (audio off), $0.14/s (audio on) — via the reseller fal.ai | 5 or 10 s | audio doubles the price | Free users may NOT use output commercially (ToS 4.6); only paying members |
| Ray 3.2 Luma AI | per clip: 1080p/5 s = $1.20 (≈ $0.24/s), 720p/5 s = $0.30 | 5 or 10 s | — | per plan |
| Sora 2 Pro OpenAI | $0.30/s (Sora 2: $0.10/s) | 720×1280 / 1280×720 only | yes | per OpenAI terms |
| Veo 3.1 Google | Standard $0.40/s (720p/1080p), $0.60/s (4K); Fast from $0.10/s; Lite from $0.05/s | 4, 6 or 8 s | always on, included | Google claims no ownership; in the EU only via the paid tier |
| Gen-4.5 Runway | 12 credits/s, 1 credit = $0.01 → $0.12/s | per plan | — | unrestricted; Runway reserves the right to train on outputs |
| Wan 2.7 Alibaba | via Alibaba Cloud Model Studio; open weights only up to Wan 2.2 (Apache-2.0) | 2 to 15 s | own file or auto-generated | per Model Studio terms |
As of 24 July 2026. All prices are net list prices in US dollars taken directly from the provider's page; EU buyers may face VAT or reverse charge (not tax advice). Prices, model IDs and availability change quickly — check with the provider before deciding anything.
Two things stand out.
First, the spread. Between the cheapest and the most expensive quoted per-second price sits roughly a factor of six — Kling at $0.07, Veo Standard at $0.40. Look only at that number and you’d pick Kling and be done. But the number says nothing about how many attempts you need before a clip works. And that’s exactly where the real bill is decided.
Second, the conditions the pricing pages leave off the front. With Kling, free users may not use the output commercially at all — that’s in clause 4.6 of the terms, not on the pricing page. With Veo, audio is always on and included; with Kling it doubles the price. Sora 2 supports a single aspect ratio. And Sora is being shut down in a few weeks — more on that below.
The calculation nobody writes down: discards
Here’s the one cleanly documented practitioner figure I could find, and it’s sobering.
For an AI commercial that aired during the NBA Finals (for the prediction market Kalshi), its maker PJ Accetturo documented the effort openly: around 300 to 400 generations for 15 usable clips. One person, two days. Tools: Veo for the video, Gemini and ChatGPT for the prompts, CapCut and Premiere for the edit, Topaz for upscaling.
Now run that against the per-second price. If you throw away 20 to 25 attempts before one lands, then the single usable clip doesn’t cost you the per-second price — it costs twenty times that. And suddenly the difference between $0.07 and $0.40 is secondary: what matters is which model makes you discard less often, not which has the cheapest single second.
That’s why I’m not crowning a winner here. The discard rate depends on your subject — a still product shot fails differently than a character walking through frame. Anyone telling you “model X is cheapest” has left out the most expensive variable.
And then comes post
The clip out of the model isn’t the finished clip. It’s raw material.
Take the upscaling the Kalshi spot explicitly needed. Runway’s own 4K upscaler bills per frame according to a documented formula. For a ten-second video at 30 frames per second that’s 360 credits — $3.60, just for the upscale. On a model whose generation costs $0.12 per second, upscaling ten seconds therefore costs more than generating them. Topaz, the tool used in that spot, is a separate subscription on top.
And that’s just the upscale. Then come editing, colour, sound mixing, and assembling 15 clips into something that doesn’t feel like 15 prompts glued together. That’s craft, and craft takes time — time that appears in no per-second price.
The honest sentence: AI video is an accelerator, not a replacement. The prompt replaces the shooting. It doesn’t replace the filmmaking.
Which model for what — carefully, because tomorrow it’s different
I deliberately won’t say “the best model”. I’ll say what the evidence shows today, 24 July 2026.
In the blind-comparison arena run by Artificial Analysis, where users rate two clips against each other without brand names, Gemini Omni Flash leads as of the retrieval date — with audio, at an Elo of 1,245 from just under 4,000 votes. That’s a snapshot, not a law of nature: Elo and vote counts shift daily, and behind it press Dreamina Seedance and a row of models nobody had heard of a year ago.
Physics is the hardest breaking point — and it doesn’t track visual realism. Google DeepMind’s Physics-IQ benchmark measures how well a model understands physical processes (does the object fall correctly, does the water flow plausibly). The authors themselves stress that visual realism and physical understanding do not correlate. A clip can look photoreal and still show physics that doesn’t exist. No current model comes close to real video.
The marketing promises about character consistency — the same person across several shots — are exactly that: promises. Google describes Veo 3.1 on its own blog as delivering “superior audio and visual quality” and lets you supply up to three reference images of a character. Google names no independent measurement for it. From my own work: character consistency is the unsolved core problem. Reference images help but guarantee nothing — the more shots, the more drift.
The question almost nobody asks: may I sell this?
For a commercial, “I think so” is not an answer. What the terms of use say differs by provider — and rarely appears on the pricing page.
Kling explicitly forbids free users any commercial use (ToS 4.6: “without our written permission, you may not use … for any commercial purposes”). Only paying members may. And Kling requires the output to be visibly marked as “Kling AI” (ToS 4.5).
Runway permits commercial use without restriction and claims no ownership of your outputs — but reserves the right to train on them. And applications using the API must visibly display “Powered by Runway” and link to runway.com (ToS).
Google claims no ownership of generated videos, but points out explicitly that it may generate the same or similar material for others. Every Veo output carries an invisible SynthID watermark.
In short: “I generated it, so it’s mine” holds automatically with none of these providers. Read the clause before the clip goes on air.
Tool mortality — the catch nobody plans for
On 24 March 2026, OpenAI announced the shutdown of its Videos API and all Sora 2 models, effective 24 September 2026. No successor model is listed on the deprecations page (as of 24 July 2026). What that implies about OpenAI’s strategy I don’t know — it doesn’t say.
But the lesson stands: build a production on a model and you’re building on something the provider can switch off with a few months’ notice. Same pattern at Google: Veo 3 and Veo 2 were shut down on 30 June 2026, and the current Veo 3.1 still carries “preview” in its name. And genuinely open, self-hostable weights exist for Alibaba’s Wan only up to version 2.2 — everything newer is API-only, and therefore provider-dependent.
If you want independence, plan for it: either commit to open weights (and accept the quality gap) or expect to rebuild your workflow every few months.
The German legal position — briefly, and explicitly not legal advice
What follows is a rendering of statutory text and official sources as of 24 July 2026 — not legal advice. Whether and how it applies to your project is a question for a lawyer.
Labelling duty. From 2 August 2026, Article 50 of the EU AI Act (per the European Commission’s FAQ) requires synthetic content to be marked in a machine-readable format and detectable as AI-generated; anyone deploying an AI video must clearly disclose a “deepfake” at first exposure at the latest. The transition period to 2 December 2026 covers only providers’ marking duty for pre-existing systems. In Germany, the Federal Network Agency supervises this following the AI Implementation Act (Bundestag, 11 June 2026). Fines for transparency breaches reach up to €15 million or 3% of worldwide annual turnover (Art. 99(4) AI Act).
The surprising copyright detail: a purely AI-generated video belongs to nobody in Germany. Section 2(2) of the Copyright Act protects only “personal intellectual creations” — a result produced fully autonomously by an AI doesn’t meet that bar. It isn’t copyright-protected; in theory anyone may copy it. Protection arises only through your own creative contribution — editing, selection, composition.
And the flip side: even a completely synthetic face can infringe rights. Section 22 of the Art Copyright Act (right to one’s own image) applies as soon as a real person is recognisable in the generated image — without a camera ever rolling. An AI face that happens, or is made, to look like a real human is not a legal free zone.
My take
If you take one number away, take this one: the per-second price is not the price. Multiply it by your discard rate, add upscaling and editing, and you land at a multiple. A model at half the per-second price that makes you discard three times as often is more expensive.
So the right question isn’t “which is cheapest” but “which discards least on my kind of subject” — and only you can answer that, by pushing your own typical shot through two or three models and counting the hit rate. Not by reading the price list.
And before the clip goes on air: read the terms, add the label, and when in doubt ask someone who knows the law.
Which tools I use for everything around it — image, edit, sound, upscaling — is in my continuously maintained index of 137 AI tools.
Sources
All checked on 24 July 2026 directly with the providers or official bodies. Prices and model states change quickly.
- Sora 2 model documentation and API deprecations — OpenAI (shutdown effective 24 Sep 2026)
- Gemini API pricing and Veo docs — Google (Veo 3.1, Gemini Omni Flash)
- API pricing and Terms of Use — Runway
- Terms of Service — Kling AI
- Kling 2.6 Pro (reseller pricing) — fal.ai
- Luma pricing — Luma AI · Wan Video (open weights) — Alibaba
- Physics-IQ benchmark — Google DeepMind (physical understanding)
- Text-to-video arena — Artificial Analysis (blind comparison)
- Kalshi spot, making-of — PJ Accetturo (300–400 generations for 15 clips)
- Transparency obligations under Article 50 AI Act — European Commission
- Copyright / right to one’s own image: § 2(2) UrhG, § 22 KunstUrhG (Germany)
If something here looks outdated or wrong to you: get in touch.
Comments
Sign in to comment, like and save. Sign in →
No comments yet. Write the first one.