Best Ai Video Generators – Fall 2026

Let’s review the best AI video generators available as of Fall 2026, on the AI service Poe.


One Prompt for 20 AI Video Models

We’re going to test 20 different AI video generators on Poe.

Every Ai video generator will be given this text prompt:

Pullback shot depicting a guitar being strummed by a grunge-rock musician, standing in front of the reflecting pond at Balboa Park in San Diego, California.

Each ai model will takes that text, and produce a video that represents it.

We’re also visiting the actual location (Balboa Park in San Diego, CA) in real life, so we can compare the real Balboa Park against the ‘virtual’ Balboa Park that is generated by these Ai video models.

The real reflecting pond and Botanical Building at Balboa Park, San Diego
Here we are at the reflecting pond of Balboa Park, in real life.

Testing the Best AI Video Generators

There are a LOT of Ai video generators on poe.com.

This article will only be testing Ai video models that can generate their own video AND sound.

There are plenty of other, silent Ai video generators on Poe too.

A full table of every AI video generator on Poe is at the bottom of this article.

Below you’ll find the various Ai video generation “model-family” names, plus their “individually-named” Ai video generation models, for example, “Seedance” (model-family name) > “Seedance-2-Fast” (model name)

Seedance (ByteDance)

1. Seedance-2-Fast, $1.50

6.04s · 1280×720 · 24fps · 135s render · 49,505 pts · $1.50

Seedance-2-Fast clip playing, with its delivered specs and cost on screen
Seedance-2-Fast. Click to play.

It does a good job of making him look like he’s really playing.

It is not realistic to the ‘Balboa Park’ part of the request.

▶ Watch this model in the video (1:41) Raw clip & specs ↗ Try Seedance-2-Fast on Poe ↗

2. Seedance-2.0-Fast-EL, $1.58

6.04s · 1280×720 · 24fps · 164s render · 52,000 pts · $1.58

Seedance-2.0-Fast-EL clip playing, with its delivered specs and cost on screen
Seedance-2.0-Fast-EL. Click to play.

It’s a pullback shot, like our prompt requested.

The music sounds good, and somehow this solo guitarist even has drums & bass.

▶ Watch this model in the video (2:06) Raw clip & specs ↗ Try Seedance-2.0-Fast-EL on Poe ↗

3. Seedance-2.0-Pro-EL, $1.68

6.04s · 1280×720 · 24fps · 128s render · 55,500 pts · $1.68

Seedance-2.0-Pro-EL clip showing a guitarist against a golden domed building at sunset
Seedance-2.0-Pro-EL. Click to play.

This guy looks like he’s at the Taj Mahal, not Balboa Park.

And it cost even more than the earlier one, at $1.68 instead of $1.58.

▶ Watch this model in the video (2:21) Raw clip & specs ↗ Try Seedance-2.0-Pro-EL on Poe ↗

4. Seedance-2.0, $1.88

6.04s · 1280×720 · 24fps · 133s render · 61,880 pts · $1.88

Seedance-2.0 clip playing, with its delivered specs and cost on screen
Seedance-2.0. Click to play.

The audio isn’t in sync with the way he’s playing.

The pullback shot is kind of nonsense.

And there’s water on the cement.

It’s the priciest Seedance of the four, and it doesn’t look like Balboa Park.

▶ Watch this model in the video (2:44) Raw clip & specs ↗ Try Seedance-2.0 on Poe ↗

Kling (Kuaishou)

5. Kling-2.6-Pro, $0.71

5.04s · 1920×1080 · 24fps · 86s render · 23,334 pts · $0.71

Kling-2.6-Pro clip playing, with its delivered specs and cost on screen
Kling-2.6-Pro. Click to play.

It’s fine, at 71 cents, in 1080p.

Totally useful ai video model.

▶ Watch this model in the video (3:30) Raw clip & specs ↗ Try Kling-2.6-Pro on Poe ↗

6. Kling-O3, $1.70

6.04s · 1920×1080 · 24fps · 141s render · 56,000 pts · $1.70

Kling-O3 clip playing, with its delivered specs and cost on screen
Kling-O3. Click to play.

It looks nice, and it’s doing fine.

▶ Watch this model in the video (3:49) Raw clip & specs ↗ Try Kling-O3 on Poe ↗

7. Kling-v3-Pro, $2.04

6.04s · 1920×1080 · 24fps · 124s render · 67,200 pts · $2.04

Kling-v3-Pro clip playing, with its delivered specs and cost on screen
Kling-v3-Pro. Click to play.

It isn’t any more realistic than others, at least not for Balboa Park.

It’s also nearly three times the price of Kling-2.6-Pro, at $2.04 against $0.71.

▶ Watch this model in the video (4:21) Raw clip & specs ↗ Try Kling-v3-Pro on Poe ↗

Veo (Google)

8. Veo-3.1-Lite, $0.36 💰

8.00s · 1280×720 · 24fps · 44s render · 12,000 pts · $0.36

Veo-3.1-Lite clip playing, with its delivered specs and cost on screen
Veo-3.1-Lite. Click to play.

This clip was really cheap to produce, at 36 cents.

It’s the Lite model, so it spent less time processing the request (44 seconds, the fastest render of all twenty.)

So, wow, the cheapest Ai video model is also the fastest.

▶ Watch this model in the video (4:38) Raw clip & specs ↗ Try Veo-3.1-Lite on Poe ↗

9. Veo-3.1-Fast, $0.82

6.00s · 1920×1080 · 24fps · 98s render · 27,000 pts · $0.82

Veo-3.1-Fast clip playing, with its delivered specs and cost on screen
Veo-3.1-Fast. Click to play.

It’s fine.

It’s not really ‘the one’.

▶ Watch this model in the video (5:03) Raw clip & specs ↗ Try Veo-3.1-Fast on Poe ↗

10. Veo-v3.1-Fast, $0.91

6.00s · 1280×720 · 24fps · 48s render · 30,000 pts · $0.91

Veo-v3.1-Fast clip showing a campus-like setting rather than Balboa Park
Veo-v3.1-Fast. Click to play.

The clip reads more like a college campus, in somewhere like Ojai, California.

Veo-3.1-Fast is published by @google, and Veo-v3.1-Fast is published by @fal.

They share a core, but the packager and the price are different, and one gives you 1080p while the other gives you 720p.

▶ Watch this model in the video (6:43) Raw clip & specs ↗ Try Veo-v3.1-Fast on Poe ↗

11. Veo-v3.1, $2.42

6.00s · 1280×720 · 24fps · 54s render · 80,000 pts · $2.42

Veo-v3.1 clip playing, with its delivered specs and cost on screen
Veo-v3.1. Click to play.

This one does two cuts, and that’s weird.

The guitar isn’t synced to his hands either.

That’s $2.42, delivered at 720p when 1080p was requested, so not a great value.

▶ Watch this model in the video (7:07) Raw clip & specs ↗ Try Veo-v3.1 on Poe ↗

12. Veo-3.1, $2.42

6.00s · 1920×1080 · 24fps · 115s render · 80,000 pts · $2.42

Veo-3.1 clip playing, with its delivered specs and cost on screen
Veo-3.1. Click to play.

It’s the same exact thing as the last model, but he’s a little more symmetrical.

It spent twice as long rendering for the identical price: 115 seconds against 54.

What we’re learning is that Veo is very good at making video that looks and sounds good.

Veo’s adherence to a given location is not great, though.

▶ Watch this model in the video (7:37) Raw clip & specs ↗ Try Veo-3.1 on Poe ↗

Sora (OpenAI)

13. Sora-2, $0.73

8.00s · 1280×720 · 30fps · 113s render · 24,000 pts · $0.73

Sora-2 clip playing, with its delivered specs and cost on screen
Sora-2. Click to play.

There’s audio… it’s highly compressed, and it sounds very fake, but there is audio.

The video is soft and overly-blocky/ compressed.

This song also has lyrics, but I never asked for the performer to sing in the original prompt.

▶ Watch this model in the video (8:24) Raw clip & specs ↗ Try Sora-2 on Poe ↗

14. Sora-2-Pro, $3.64 💸

8.00s · 1792×1024 · 30fps · 445s render · 120,000 pts · $3.64

Sora-2-Pro clip showing a guitarist in a colonnade at Balboa Park
Sora-2-Pro. Click to play.

This is among the best-looking clips in the test, and it cost 5x more than Sora-2.

That’s $3.64 instead of 73 cents.

On looks alone, the 5x is justified.

Sore-2 Pro missed the ‘pullback shot’ instructions, and instead returned a static wide shot, but the cheaper Sora-2 got the pullback right.

This one also took 445 seconds to render, and it failed outright on the initial attempt.

Take a moment and compare Sore-2 Pro to Veo-3.1-Lite:

Sora-2-Pro is 10x the money and 10x the wait.

Veo 3.1 Lite was 10x cheaper, 10x faster, and followed directions better, but it had less overall ability to depict the real-world location from the prompt.

▶ Watch this model in the video (9:15) Raw clip & specs ↗ Try Sora-2-Pro on Poe ↗

Wan (Alibaba)

15. Wan-2.5, $0.76

5.04s · 1920×1080 · 24fps · 143s render · 25,000 pts · $0.76

Wan-2.5 clip playing, with its delivered specs and cost on screen
Wan-2.5. Click to play.

The sound fidelity is awful; the singer sounds like a robot.

The video itself is fine-ish, and he’s moving his feet, which is realistic.

▶ Watch this model in the video (9:53) Raw clip & specs ↗ Try Wan-2.5 on Poe ↗

16. Wan-2.6, $0.84 🗣️

6.00s · 1920×1080 · 30fps · 70s render · 27,600 pts · $0.84

Wan-2.6 clip playing, with its delivered specs and cost on screen
Wan-2.6. Click to play.

This one created a unique-looking, smooth-ish guy who monotones some lyrics.

“Yeah, this is where it all starts. Still feels like home, even when I’m lost. This place, it doesn’t care if you’re broke…”

Wan-2.6 has a setting where you can tell it “I want multiple shots within an output.” and somehow that got enabled for this video.

▶ Watch this model in the video (10:32) Raw clip & specs ↗ Try Wan-2.6 on Poe ↗

17. Wan-2.7, $0.91

6.00s · 1920×1080 · 30fps · 126s render · 30,000 pts · $0.91

Wan-2.7 clip playing, with its delivered specs and cost on screen
Wan-2.7. Click to play.

It looks fine.

It checked all the technical boxes for our test, but it’s not very close to ‘Balboa Park’ in appearance.

▶ Watch this model in the video (11:04) Raw clip & specs ↗ Try Wan-2.7 on Poe ↗

18. Wan-3.0, $1.70

6.00s · 1920×1080 · 30fps · 287s render · 56,000 pts · $1.70

Wan-3.0 clip playing, with its delivered specs and cost on screen
Wan-3.0. Click to play.

This one does look decently like where we are.

It also took 287 seconds, second-slowest of the test.

▶ Watch this model in the video (11:13) Raw clip & specs ↗ Try Wan-3.0 on Poe ↗

PixVerse

19. Pixverse-v5.6, $1.52

5.04s · 1920×1080 · 24fps · 84s render · 50,000 pts · $1.52

Pixverse-v5.6 clip showing Plaza de Espana style architecture
Pixverse-v5.6. Click to play.

This one’s actually a zoom out, instead of a pullback shot.

And it looks much more like Spain than San Diego.

▶ Watch this model in the video (11:31) Raw clip & specs ↗ Try Pixverse-v5.6 on Poe ↗

Gemini Omni (Google)

20. Gemini-Omni-1.1-Flash, $1.54

10.00s · 1920×1080 · 24fps · 69s render · 50,921 pts · $1.54

Gemini-Omni-1.1-Flash clip showing the guitarist at the Balboa Park reflecting pond
Gemini-Omni-1.1-Flash. Click to play.

This is the most realistic one yet.

It looks a lot like where we’re supposed to be, and there’s good adherence to all aspects of our text-prompt.

If you can afford it, use Gemini Omni 1.1 Flash

You can get a closer comparison between Gemini Omni Flash’s ‘virtual San Diego’ against the real San Diego, below.

(Hover over the image for a magnified look at any section.)

Side-by-side comparison of the Gemini-Omni-1.1-Flash video frame and a real photograph of the reflecting pond at Balboa Park, San Diego
Notice what matches and what doesn’t. On a phone, press and drag across the image.

▶ Watch this model in the video (11:46) Raw clip & specs ↗ Try Gemini-Omni-1.1-Flash on Poe ↗

What the Best AI Video Generators Cost

Every price below came off Poe’s actual costs on September 4th, 2026.

Model Cost Render Delivered
Veo-3.1-Lite $0.36 44s 720p · 8.00s
Kling-2.6-Pro $0.71 86s 1080p · 5.04s
Sora-2 $0.73 113s 720p · 8.00s
Wan-2.5 $0.76 143s 1080p · 5.04s
Veo-3.1-Fast $0.82 98s 1080p · 6.00s
Wan-2.6 $0.84 70s 1080p · 6.00s
Veo-v3.1-Fast $0.91 48s 720p · 6.00s
Wan-2.7 $0.91 126s 1080p · 6.00s
Seedance-2-Fast $1.50 135s 720p · 6.04s
Pixverse-v5.6 $1.52 84s 1080p · 5.04s
Gemini-Omni-1.1-Flash 🏆 $1.54 69s 1080p · 10.00s
Seedance-2.0-Fast-EL $1.58 164s 720p · 6.04s
Seedance-2.0-Pro-EL $1.68 128s 720p · 6.04s
Kling-O3 $1.70 141s 1080p · 6.04s
Wan-3.0 $1.70 287s 1080p · 6.00s
Seedance-2.0 $1.88 133s 720p · 6.04s
Kling-v3-Pro $2.04 124s 1080p · 6.04s
Veo-v3.1 $2.42 54s 720p · 6.00s
Veo-3.1 $2.42 115s 1080p · 6.00s
Sora-2-Pro $3.64 445s 1792×1024 · 8.00s

Here’s the total cost of this AI video generator experiment.

The twenty clips came to $29.66.

They took 43 minutes of rendering.

The Best AI Video Generator in Fall 2026

Gemini-Omni-1.1-Flash is the pick.

Its video faithfully depicted Balboa Park’s reflecting-pond area, and it followed the actual camera instruction.

It rendered in 69 seconds.

The clip was 10 seconds long and cost $1.54, on September 4th, 2026.

Carleton at Balboa Park delivering the verdict on Gemini-Omni-1.1-Flash
If you can afford it, use this one.

Here are three other candidates for best AI video model, based on your needs.

  • Cheap and fast, for quick experiments: Veo-3.1-Lite. 36 cents, 44 seconds, with audio.
  • Prettiest video: Sora-2-Pro, if you can handle $3.64 a clip and a seven-and-a-half-minute wait.
  • Works simply: Wan-2.7 did exactly what it was told, and it invented nothing.

If you read any of the prior articles, you might remember that I rated a model called Ray-2 as the best AI video generator of its time.

Just a year later, many of those earlier Ai-video generator models no longer exist:

Ray 2, Kling-2.1-Master, all three Pika models, Dream Machine, Veo 2, and Runway Gen-4 Turbo are all gone.

Any video model mentioned today might not exist tomorrow, so enjoy them now. 🙂

Run Your Own Best AI Video Generator Test

I recommend running this test yourself, if you have Poe.

It’s the best way to get an answer with real-world context for your use case.

A bot on Poe called Script-Bot-Creator can run the whole test for you.

There’s a prompt template for it below.

First, put a location in the prompt that you personally know.

Balboa Park is an arbitrary example from the original experiment, not a recommendation.

Ideally you should pick somewhere you can visit yourself.

It could be your street, your local park, or the front of your own building, if it’s a famous landmark.

Open Script-Bot-Creator and paste this brief into a new chat.

The brief now interviews you first.

If you paste it unedited, it stops and asks which location, subject, camera move, length and bots you want, then rewrites the prompt from your answers.

Ask it for the gallery at the end and it builds you a results page with every clip and a download button.

You are running a controlled AI-video benchmark. Precision matters more
than speed or creativity. Call Poe video bots directly and report
structured results.

STEP ZERO: ASK ME WHETHER I CHANGED ANYTHING
Compare the CONTROL PROMPT below against this exact default text:
"Pullback shot depicting a guitar being strummed by a grunge-rock
musician, standing in front of the reflecting pond at Balboa Park in
San Diego, California."

If it still matches word for word, I pasted this brief unedited. Say
so, then ask me all five of these in ONE message, each with your own
recommendation, and WAIT for my answer before doing anything else:

  1. The location. This test only means something if you know the
     place well enough to judge whether a model got it right. Balboa
     Park is an arbitrary example from the original experiment. Think
     about a location that you, personally, know. Name a street, a
     park, a shopfront, a building you can picture with your eyes
     shut.
  2. The subject and what they are doing. A grunge-rock musician
     strumming a guitar is the original example, not a requirement.
  3. The camera move. A pullback is the original example. Push-in,
     crane down, orbit and handheld follow are the usual
     alternatives.
  4. Clip length and resolution. My defaults are the longest each bot
     offers up to 10 seconds, 1080p where it is available, and 16:9
     landscape always.
  5. Which bots to run. My default is every video bot on Poe that
     makes its own audio. Say "all with audio", or "everything", or
     name a list.

Then rewrite the control prompt from my answers, show it back to me
in full, and get my yes before the first generation. Once I say yes,
that exact string is frozen for the whole run and never edited again.

If the control prompt does NOT match the default above, I have
already edited it myself. Do not ask. Read it back to me once and
start.

FIRST, BEFORE YOU PROPOSE ANYTHING, READ THE BOT LIST
Open this page and read it in full.
   URL: https://carletontorpin.com/ai/best-ai-video-generators-fall-2026/

The complete roster of 47 video bots is under the heading "Every AI
Video Generator on Poe". Jump straight to it.
   URL: https://carletontorpin.com/ai/best-ai-video-generators-fall-2026/#full-table

Two traps are waiting there. The first table you meet has only TWO
example rows, which is a preview and not the list. The real 47-row
table sits inside the collapsed grey bar beneath it, labeled "Click to
expand the full table". Fetch the page's raw HTML and all 47 rows are
present whether that block is open or shut. If you are driving a real
browser instead, click that grey bar first, and leave the filter box
above the table empty, because anything typed in it hides rows.

That table is the authority on spelling. Poe bot names are exact
strings and a near miss silently fails, so read the row rather than
guessing. "Veo-3.1-Fast" and "Veo-v3.1-Fast" are different bots, and
so are "Seedance-2-Fast" and "Seedance-2.0-Fast-EL". Every name in the
table links to that bot's own poe.com page. Follow the link instead of
building a URL by hand.

The columns are: Bot with its @publisher underneath, Family, cheapest
points per second, dollars per 5 seconds, max resolution, max
duration, whether it makes its own audio, and what it is for. Poe's
conversion at capture time was 33,000 compute points to $1.00.

Higher up that same page, under "What the Best AI Video Generators
Cost", a second table lists the 20 bots that were actually run, with
what each one really charged, how long it really took, and the specs
it really delivered. Use the 47-row table to pick and spell bots. Use
this one for real costs and real render times.

If you cannot read either table, say so and stop. Do not fall back on
bot names you remember.

THE CONTROL PROMPT: use this EXACT text for every generation. Never
rewrite, expand, shorten, "improve," or add camera tags to it:

"Pullback shot depicting a guitar being strummed by a grunge-rock
musician, standing in front of the reflecting pond at Balboa Park in
San Diego, California."

STANDING RULES
1. Text-to-video only. Never attach an image.
2. Audio ON wherever the bot exposes a toggle or generates it natively.
3. Aspect ratio 16:9 LANDSCAPE, set explicitly. Never accept a vertical
   default.
4. Use exactly the duration and resolution given per bot. If a value
   is unavailable, stop and report rather than guessing.
5. One generation per bot. No re-rolls, no cherry-picking. The first
   output is the result, including when it's bad.
6. If a bot refuses or errors, capture the EXACT text it returned. A
   refusal is a finding, not a retry.
7. Never substitute a different bot for one that fails.
8. Before running a batch, check whether those bots already have
   results in this session. If any do, stop and ask.

REPORT FOR EVERY GENERATION, as a row:
   bot | duration requested | duration delivered | resolution delivered
   | aspect delivered | audio present Y/N | what the audio contains
   | compute points charged | USD | generation time in seconds
   | direct video URL | deviations from what was asked

DEVIATIONS ARE THE POINT. Flag any of: aspect ratio changed, duration
different from requested, unrequested cuts or scene changes, added
dialogue or lyrics, watermark, vertical output.

ON AUDIO, also report: is it music, singing, dialogue, or ambient?
Is it synced to the picture? Did the model invent lyrics?

After each batch, output one markdown table of all rows, then a total
for points and USD. Wait for me to send the next batch.

THE GALLERY, IF I ASK FOR IT
When the run is done, if I ask for a gallery, a results page, or just
"the page", build ONE self-contained HTML file and hand it to me as a
download. Do not publish it as an artifact. Artifact security blocks
video from other domains, so not one clip would play.

WHAT GOES ON THE PAGE
Black background, a single column about 1000px wide, base text 17 to
19px. One card per model, in the order they were run. Each card has:
  a header with the model name, linked to its own page on poe.com,
    with its @publisher beside it
  one spec line of plain facts: seconds, resolution, fps, audio sample
    rate, render time, compute points, dollars
  the clip itself
Facts only on this page. No verdicts, no "stunning", no "impressive".
I supply the opinions.

THE CLIPS
Use a video tag with loop and playsinline and NO controls. One click
plays, another pauses, and starting one pauses all the others. Give
each video an onerror that tells me whether the link expired or was
never captured, with a paste box so I can drop in a replacement.

THE DOWNLOAD BUTTON, WHICH IS THE WHOLE POINT
Poe's video links expire, often within weeks, and at least one model
deletes its output after 48 hours. So put a large button at the top of
the page reading "Copy download command for all clips". Clicking it
copies a ready-to-run block of curl commands to my clipboard, one line
per clip, each saving to a sensible filename such as
01-veo-3-1-lite.mp4. Put a small download link on each card too, for
that one clip.

Say plainly on the page that the links expire and that this button is
how I keep the files.

Do not lean on the download attribute alone for the all-clips button.
Browsers ignore it for files on another domain, so it opens a tab
instead of saving. The copied curl block is the part that always
works.

NAVIGATION
A sticky bar across the bottom with one equal-width chip per model.
Clicking a chip jumps to that card's header rather than the page, so
the title lands in the same spot near the top every time.

Build it as one file with the CSS and the JavaScript inline. Then give
me the file, and tell me how many clips are in it and how many failed.

The control prompt is built from four parts.

Swap any of them out to customize your results.

  • Camera move: “Pullback shot.” Try push-in, crane down, orbit, or handheld follow.
  • Subject: “a grunge-rock musician.” Anything works: a baker, a dog, a kid on a scooter.
  • Subject action: “a guitar being strummed.” Whatever your subject is doing.
  • Environment: “the reflecting pond at Balboa Park.” Name a real, checkable landmark, so you can compare a ground truth against the AI video output.

Then send the bots in small batches of five models per prompt.

Include the duration and resolution you want for each, or ask the bot to “normalize things across the selection” of ai-video generator models you’ve made.

When the clips come back, download every clip on the same day because those Poe content-delivery-network links expire.

Also, take comfort knowing a failed Poe-video generation generally costs nothing.

Every AI Video Generator on Poe

Here’s a compiled list of 47 ai video bots on Poe, across 15 families, captured September 3rd and 4th, 2026.

“Cheapest points/sec” is each bot’s lowest published rate, so every model sits on one scale.

Poe’s own conversion as of September 2026 is 33,000 compute points = $1.00.

Bot Family Cheapest pts/sec ≈$ / 5s Max res Max sec Audio What it’s for
Seedance-1.0-Pro-Fast@Bytedance Seedance 720 $0.11 1080p 12 No Cheapest Seedance 1.0 tier – best value per token in the whole Seedance line
Seedance-1.0-Lite@Bytedance Seedance 1,296 $0.20 1080p 12 No Lightweight Seedance 1.0 for simple T2V/I2V
Click to expand the full table (all 47 AI video generators, 15 families)

No bot matches that.

Bot Family Cheapest pts/sec ≈$ / 5s Max res Max sec Audio What it’s for
Seedance-1.0-Pro-Fast@Bytedance Seedance 720 $0.11 1080p 12 No Cheapest Seedance 1.0 tier – best value per token in the whole Seedance line
Seedance-1.0-Lite@Bytedance Seedance 1,296 $0.20 1080p 12 No Lightweight Seedance 1.0 for simple T2V/I2V
Seedance-1.0-Pro@Bytedance Seedance 1,800 $0.27 1080p 12 No Previous-generation Seedance flagship; strong semantic understanding and prompt following
Seedance-2-Fast@Bytedance Seedance 3,833 $0.58 720p 15 Yes Speed-optimised Seedance 2.0 for rapid iteration and high-volume workflows
Seedance-2.0-Fast-EL@empiriolabsai Seedance 4,067 $0.62 720p 15 Yes Fast multimodal Seedance with the full mode set at roughly 5% less than Pro-EL
Seedance-2.0-Pro-EL@empiriolabsai Seedance 4,286 $0.65 4K (3840×2160, 10-bit H.265) 15 Yes Full multimodal Seedance 2.0 Pro – the only bot here with 4K and video edit/extend
Seedance-2.0@Bytedance Seedance 4,807 $0.73 720p 15 Yes Cinematic T2V/I2V with consistent characters and detailed motion control
Veo-3.1-Lite@google Veo 1,500 $0.23 1080p 8 Yes Cheapest native-audio video on Poe – $0.05/sec at 720p
Veo-v3.1-Fast@fal Veo 3,334 $0.51 1080p 8 Yes Cheapest fal Veo endpoint with an explicit silent discount
Veo-3.1-Fast@google Veo 4,500 $0.68 1080p 8 Yes Veo 3.1 quality at a third of the price – the practical Veo workhorse
Veo-v3.1@fal Veo 6,667 $1.01 1080p 8 Yes fal's Veo 3.1 endpoint – explicit audio/silent price split and first-to-last-frame support
Veo-3-vFast@fal Veo 8,334 $1.26 Not exposed 7 Yes Legacy Veo 3 fast endpoint – text-to-video only, fixed 7 seconds
Veo-3.1@google Veo 13,333 $2.02 1080p 8 Yes Google flagship – native audio, reference mode, strong prompt adherence
Kling-2.1-Std@fal Kling 1,667 $0.25 Not exposed 10 No Cheapest Kling endpoint – cost-efficient image-to-video
Kling-2.6-Pro@fal Kling 2,334 $0.35 Not exposed 10 Yes Best value Kling with native audio – $0.071/sec silent, $0.14/sec with sound
Kling-2.5-Turbo-Pro@fal Kling 2,334 $0.35 Not exposed 10 No Fast, cheap silent Kling for image-driven motion
Kling-2.1-Pro@fal Kling 2,834 $0.43 Not exposed 10 No Cinematic image-to-video with precise camera movement and motion control
Kling-Pro-Effects@fal Kling 3,334 $0.51 Not exposed 10 No Canned photo effects – squish, expand, hug, kiss, heart gesture
Kling-Omni@fal Kling 3,734 $0.57 Not exposed 10 No Image-to-video and first-to-last-frame at a flat mid-range price
Kling-v3-Motion-Ctrl@empiriolabsai Kling 4,667 $0.71 1080p 30 No Motion transfer – drive a character from a still image with a reference video's movement
Kling-O3@empiriolabsai Kling 5,600 $0.85 4K 15 (10 when any video input is used) Yes Multi-scene storytelling in a single prompt – up to 6 scenes with per-scene timing
Kling-2.0-Master@fal Kling 6,000 $0.91 Not exposed 10 No Legacy Kling 2.0 Master; supports CLI-style flags in the prompt
Kling-v3-Pro@fal Kling 7,467 $1.13 Not exposed 15 Yes Newest Kling flagship – T2V/I2V with start+end frames, native audio, 15s max
Wan-2.6@empiriolabsai Wan 750 $0.11 1080p 10 for R2V Yes Multi-shot storytelling with lip-sync; Flash mode is the cheapest video generation on Poe
Wan-2.5@fal Wan 1,667 $0.25 1080p 10 No Simple Wan endpoint with audio-guided generation from an uploaded mp3
Wan-3.0@empiriolabsai Wan 2,333 $0.35 1080p 30 Yes Newest Wan – the longest duration range here (2-30s) plus a speed tier
Wan-Animate@fal Wan 2,500 $0.38 Source-dependent Source-dependent No Character replacement in an existing video, or motion transfer onto a still
Wan-2.7@empiriolabsai Wan 3,333 $0.51 1080p 10 for video-input modes Yes Four-mode Wan with true video editing and voice-timbre reference
Sora-2@openai Sora 3,000 $0.45 720p 20 (panel) Yes Realistic physics, synchronized dialogue and SFX, multi-shot prompt adherence
Sora-2-Pro@openai Sora 9,000 $1.36 1920×1080 / 1792×1024 20 Yes Highest-fidelity Sora – world-state persistence, complex multi-shot, up to 20s
Gemini-Omni-Flash@google Gemini Omni token-metered n/a Not exposed Not exposed Yes Conversational video editing – generate then refine in natural language
Gemini-Omni-1.1-Flash@google Gemini Omni token-metered n/a 4K Not exposed Yes Scene extension and first/last-frame interpolation with 4K output
Pixverse-v4.5@fal PixVerse 2,000 $0.30 1080p 8 No Meme-style canned effects (Hulk, Venom, Kiss Me AI) plus style presets
Pixverse-v5.6@empiriolabsai PixVerse 2,333 $0.35 1080p 10 (8 at 1080p) Yes Style-preset video with optional audio; the deepest control surface of the PixVerse bots
Pixverse-v5@empiriolabsai PixVerse 3,000 $0.45 1080p 8 (5 at 1080p) No Three clean modes decided purely by how many images you attach
Hailuo-02-Standard@fal Hailuo (MiniMax) 1,500 $0.23 768p 10 No Cheap MiniMax image-to-video at 768p
Hailuo-02-Pro@fal Hailuo (MiniMax) 2,667 $0.40 1080p 5 No 1080p MiniMax image-to-video, fixed 5-second clips
Hailuo-AI@fal Hailuo (MiniMax) 2,833 $0.43 Not exposed Not exposed No MiniMax's general text/image-to-video, billed as a flat fee per message
Hailuo-Director-01@fal Hailuo (MiniMax) 3,333 $0.51 Not exposed 5 No THE camera-control bot – explicit shot moves in square brackets
LTX-2-Fast@fal LTX 1,334 $0.20 2160p (4K) 10 Yes Cheapest route to 4K (2160p) and the only 50 FPS option here
LTX-2-Pro@fal LTX 2,000 $0.30 2160p (4K) 10 Yes Professional-grade LTX-2 up to 2K/4K with 50 FPS
Grok-Imagine-Video@xAI Grok Imagine 1,667 $0.25 Not exposed Slider max not labelled No Artistic/creative video plus mp4 video editing; widest aspect-ratio menu here
Runway-Gen-4.5@runwayml Runway Gen 10,000 $1.52 Not exposed 10 No High-end cinematic fidelity with fine-grained motion and camera behaviour
Amazon-Nova-Reel-1.1@empiriolabsai Nova Reel 4,800 $0.73 720p 120 (2 minutes) No Up to TWO-MINUTE multi-shot videos with per-shot prompting
Vidu@fal Vidu 1,333 $0.20 Not exposed 5 No ~170 one-click meme/effect templates – by far the biggest preset library here
SVI-2.0-Pro@empiriolabsai Stable Video Infinity 1,910 $0.29 720p 121.5 (480p) No LONG-FORM video – up to 121.5 seconds in a single generation, the longest here by far
OmniHuman@Bytedance OmniHuman 4,667 $0.71 Not exposed 30 No Audio-driven talking/singing avatar from one photo

Poe changes rates often, so check the Rates dialog before budgeting anything.

See Also: The Best AI Image Generators

Here’s the same experiment, but for AI image generators instead of video.

What Is The Best Ai Image Generator? tests 29 image models on photorealism, spatial reasoning, historical accuracy, and text rendering.

There’s more of it around the site: the AI section, the Photo section, etc.

One of the AI image-generator comparisons from the image article
One comparison from that article. There are plenty more inside, across 29 image models.

Fun to Notice

This is the fourth time in two years that I’ve tested the best AI video generators.

The time it takes me to make each article keeps dropping, too.

The first article, in Fall 2024, took about 12 hours, from start to finished article, companion video included.

The second article, in March 2025, took 6 hours, once AI could handle enough of the video generation on its own.

The third article, in January 2026, took about 4 hours, and so did this one.

There will certainly be a point in the next two years where this will be a test I can run in seconds, complete with an article I can author and post in under a minute.

Twenty models, one prompt, Every clip is clickable.

Have fun making stuff, and share your results!


Previously: Text to Video AI with Poe (Oct 2024) · Best Ai Video Generators in Poe Ai (Mar 2025) · Ai Video Generation Models Compared (Jan 2026)

Related: What Is The Best Ai Image Generator? · Ai Audio Generators on Poe · Script Bot Creator Poe Ai

Costs from Poe’s billing ledger, September 4th, 2026. Specs measured with ffprobe on the delivered files, not read off a product page. One output per model on one prompt, which is hardly representative of any model in general.

CategoriesAi

Leave a Reply

Your email address will not be published. Required fields are marked *