Let’s review the best AI video generators available as of Fall 2026, on the AI service Poe.
One Prompt for 20 AI Video Models
We’re going to test 20 different AI video generators on Poe.
Every Ai video generator will be given this text prompt:
Pullback shot depicting a guitar being strummed by a grunge-rock musician, standing in front of the reflecting pond at Balboa Park in San Diego, California.
Each ai model will takes that text, and produce a video that represents it.
We’re also visiting the actual location (Balboa Park in San Diego, CA) in real life, so we can compare the real Balboa Park against the ‘virtual’ Balboa Park that is generated by these Ai video models.
Testing the Best AI Video Generators
There are a LOT of Ai video generators on poe.com.
This article will only be testing Ai video models that can generate their own video AND sound.
There are plenty of other, silent Ai video generators on Poe too.
A full table of every AI video generator on Poe is at the bottom of this article.
Below you’ll find the various Ai video generation “model-family” names, plus their “individually-named” Ai video generation models, for example, “Seedance” (model-family name) > “Seedance-2-Fast” (model name)
Seedance (ByteDance)
1. Seedance-2-Fast, $1.50
6.04s · 1280×720 · 24fps · 135s render · 49,505 pts · $1.50
It does a good job of making him look like he’s really playing.
It is not realistic to the ‘Balboa Park’ part of the request.
▶ Watch this model in the video (1:41) Raw clip & specs ↗ Try Seedance-2-Fast on Poe ↗
2. Seedance-2.0-Fast-EL, $1.58
6.04s · 1280×720 · 24fps · 164s render · 52,000 pts · $1.58
It’s a pullback shot, like our prompt requested.
The music sounds good, and somehow this solo guitarist even has drums & bass.
▶ Watch this model in the video (2:06) Raw clip & specs ↗ Try Seedance-2.0-Fast-EL on Poe ↗
3. Seedance-2.0-Pro-EL, $1.68
6.04s · 1280×720 · 24fps · 128s render · 55,500 pts · $1.68
This guy looks like he’s at the Taj Mahal, not Balboa Park.
And it cost even more than the earlier one, at $1.68 instead of $1.58.
▶ Watch this model in the video (2:21) Raw clip & specs ↗ Try Seedance-2.0-Pro-EL on Poe ↗
4. Seedance-2.0, $1.88
6.04s · 1280×720 · 24fps · 133s render · 61,880 pts · $1.88
The audio isn’t in sync with the way he’s playing.
The pullback shot is kind of nonsense.
And there’s water on the cement.
It’s the priciest Seedance of the four, and it doesn’t look like Balboa Park.
▶ Watch this model in the video (2:44) Raw clip & specs ↗ Try Seedance-2.0 on Poe ↗
Kling (Kuaishou)
5. Kling-2.6-Pro, $0.71
5.04s · 1920×1080 · 24fps · 86s render · 23,334 pts · $0.71
It’s fine, at 71 cents, in 1080p.
Totally useful ai video model.
▶ Watch this model in the video (3:30) Raw clip & specs ↗ Try Kling-2.6-Pro on Poe ↗
6. Kling-O3, $1.70
6.04s · 1920×1080 · 24fps · 141s render · 56,000 pts · $1.70
It looks nice, and it’s doing fine.
▶ Watch this model in the video (3:49) Raw clip & specs ↗ Try Kling-O3 on Poe ↗
7. Kling-v3-Pro, $2.04
6.04s · 1920×1080 · 24fps · 124s render · 67,200 pts · $2.04
It isn’t any more realistic than others, at least not for Balboa Park.
It’s also nearly three times the price of Kling-2.6-Pro, at $2.04 against $0.71.
▶ Watch this model in the video (4:21) Raw clip & specs ↗ Try Kling-v3-Pro on Poe ↗
Veo (Google)
8. Veo-3.1-Lite, $0.36 💰
8.00s · 1280×720 · 24fps · 44s render · 12,000 pts · $0.36
This clip was really cheap to produce, at 36 cents.
It’s the Lite model, so it spent less time processing the request (44 seconds, the fastest render of all twenty.)
So, wow, the cheapest Ai video model is also the fastest.
▶ Watch this model in the video (4:38) Raw clip & specs ↗ Try Veo-3.1-Lite on Poe ↗
9. Veo-3.1-Fast, $0.82
6.00s · 1920×1080 · 24fps · 98s render · 27,000 pts · $0.82
It’s fine.
It’s not really ‘the one’.
▶ Watch this model in the video (5:03) Raw clip & specs ↗ Try Veo-3.1-Fast on Poe ↗
10. Veo-v3.1-Fast, $0.91
6.00s · 1280×720 · 24fps · 48s render · 30,000 pts · $0.91
The clip reads more like a college campus, in somewhere like Ojai, California.
Veo-3.1-Fast is published by @google, and Veo-v3.1-Fast is published by @fal.
They share a core, but the packager and the price are different, and one gives you 1080p while the other gives you 720p.
▶ Watch this model in the video (6:43) Raw clip & specs ↗ Try Veo-v3.1-Fast on Poe ↗
11. Veo-v3.1, $2.42
6.00s · 1280×720 · 24fps · 54s render · 80,000 pts · $2.42
This one does two cuts, and that’s weird.
The guitar isn’t synced to his hands either.
That’s $2.42, delivered at 720p when 1080p was requested, so not a great value.
▶ Watch this model in the video (7:07) Raw clip & specs ↗ Try Veo-v3.1 on Poe ↗
12. Veo-3.1, $2.42
6.00s · 1920×1080 · 24fps · 115s render · 80,000 pts · $2.42
It’s the same exact thing as the last model, but he’s a little more symmetrical.
It spent twice as long rendering for the identical price: 115 seconds against 54.
What we’re learning is that Veo is very good at making video that looks and sounds good.
Veo’s adherence to a given location is not great, though.
▶ Watch this model in the video (7:37) Raw clip & specs ↗ Try Veo-3.1 on Poe ↗
Sora (OpenAI)
13. Sora-2, $0.73
8.00s · 1280×720 · 30fps · 113s render · 24,000 pts · $0.73
There’s audio… it’s highly compressed, and it sounds very fake, but there is audio.
The video is soft and overly-blocky/ compressed.
This song also has lyrics, but I never asked for the performer to sing in the original prompt.
▶ Watch this model in the video (8:24) Raw clip & specs ↗ Try Sora-2 on Poe ↗
14. Sora-2-Pro, $3.64 💸
8.00s · 1792×1024 · 30fps · 445s render · 120,000 pts · $3.64
This is among the best-looking clips in the test, and it cost 5x more than Sora-2.
That’s $3.64 instead of 73 cents.
On looks alone, the 5x is justified.
Sore-2 Pro missed the ‘pullback shot’ instructions, and instead returned a static wide shot, but the cheaper Sora-2 got the pullback right.
This one also took 445 seconds to render, and it failed outright on the initial attempt.
Take a moment and compare Sore-2 Pro to Veo-3.1-Lite:
Sora-2-Pro is 10x the money and 10x the wait.
Veo 3.1 Lite was 10x cheaper, 10x faster, and followed directions better, but it had less overall ability to depict the real-world location from the prompt.
▶ Watch this model in the video (9:15) Raw clip & specs ↗ Try Sora-2-Pro on Poe ↗
Wan (Alibaba)
15. Wan-2.5, $0.76
5.04s · 1920×1080 · 24fps · 143s render · 25,000 pts · $0.76
The sound fidelity is awful; the singer sounds like a robot.
The video itself is fine-ish, and he’s moving his feet, which is realistic.
▶ Watch this model in the video (9:53) Raw clip & specs ↗ Try Wan-2.5 on Poe ↗
16. Wan-2.6, $0.84 🗣️
6.00s · 1920×1080 · 30fps · 70s render · 27,600 pts · $0.84
This one created a unique-looking, smooth-ish guy who monotones some lyrics.
“Yeah, this is where it all starts. Still feels like home, even when I’m lost. This place, it doesn’t care if you’re broke…”
Wan-2.6 has a setting where you can tell it “I want multiple shots within an output.” and somehow that got enabled for this video.
▶ Watch this model in the video (10:32) Raw clip & specs ↗ Try Wan-2.6 on Poe ↗
17. Wan-2.7, $0.91
6.00s · 1920×1080 · 30fps · 126s render · 30,000 pts · $0.91
It looks fine.
It checked all the technical boxes for our test, but it’s not very close to ‘Balboa Park’ in appearance.
▶ Watch this model in the video (11:04) Raw clip & specs ↗ Try Wan-2.7 on Poe ↗
18. Wan-3.0, $1.70
6.00s · 1920×1080 · 30fps · 287s render · 56,000 pts · $1.70
This one does look decently like where we are.
It also took 287 seconds, second-slowest of the test.
▶ Watch this model in the video (11:13) Raw clip & specs ↗ Try Wan-3.0 on Poe ↗
PixVerse
19. Pixverse-v5.6, $1.52
5.04s · 1920×1080 · 24fps · 84s render · 50,000 pts · $1.52
This one’s actually a zoom out, instead of a pullback shot.
And it looks much more like Spain than San Diego.
▶ Watch this model in the video (11:31) Raw clip & specs ↗ Try Pixverse-v5.6 on Poe ↗
Gemini Omni (Google)
20. Gemini-Omni-1.1-Flash, $1.54
10.00s · 1920×1080 · 24fps · 69s render · 50,921 pts · $1.54
This is the most realistic one yet.
It looks a lot like where we’re supposed to be, and there’s good adherence to all aspects of our text-prompt.
If you can afford it, use Gemini Omni 1.1 Flash
You can get a closer comparison between Gemini Omni Flash’s ‘virtual San Diego’ against the real San Diego, below.
(Hover over the image for a magnified look at any section.)
▶ Watch this model in the video (11:46) Raw clip & specs ↗ Try Gemini-Omni-1.1-Flash on Poe ↗
What the Best AI Video Generators Cost
Every price below came off Poe’s actual costs on September 4th, 2026.
| Model | Cost | Render | Delivered |
|---|---|---|---|
| Veo-3.1-Lite | $0.36 | 44s | 720p · 8.00s |
| Kling-2.6-Pro | $0.71 | 86s | 1080p · 5.04s |
| Sora-2 | $0.73 | 113s | 720p · 8.00s |
| Wan-2.5 | $0.76 | 143s | 1080p · 5.04s |
| Veo-3.1-Fast | $0.82 | 98s | 1080p · 6.00s |
| Wan-2.6 | $0.84 | 70s | 1080p · 6.00s |
| Veo-v3.1-Fast | $0.91 | 48s | 720p · 6.00s |
| Wan-2.7 | $0.91 | 126s | 1080p · 6.00s |
| Seedance-2-Fast | $1.50 | 135s | 720p · 6.04s |
| Pixverse-v5.6 | $1.52 | 84s | 1080p · 5.04s |
| Gemini-Omni-1.1-Flash 🏆 | $1.54 | 69s | 1080p · 10.00s |
| Seedance-2.0-Fast-EL | $1.58 | 164s | 720p · 6.04s |
| Seedance-2.0-Pro-EL | $1.68 | 128s | 720p · 6.04s |
| Kling-O3 | $1.70 | 141s | 1080p · 6.04s |
| Wan-3.0 | $1.70 | 287s | 1080p · 6.00s |
| Seedance-2.0 | $1.88 | 133s | 720p · 6.04s |
| Kling-v3-Pro | $2.04 | 124s | 1080p · 6.04s |
| Veo-v3.1 | $2.42 | 54s | 720p · 6.00s |
| Veo-3.1 | $2.42 | 115s | 1080p · 6.00s |
| Sora-2-Pro | $3.64 | 445s | 1792×1024 · 8.00s |
Here’s the total cost of this AI video generator experiment.
The twenty clips came to $29.66.
They took 43 minutes of rendering.
The Best AI Video Generator in Fall 2026
Gemini-Omni-1.1-Flash is the pick.
Its video faithfully depicted Balboa Park’s reflecting-pond area, and it followed the actual camera instruction.
It rendered in 69 seconds.
The clip was 10 seconds long and cost $1.54, on September 4th, 2026.
Here are three other candidates for best AI video model, based on your needs.
- Cheap and fast, for quick experiments: Veo-3.1-Lite. 36 cents, 44 seconds, with audio.
- Prettiest video: Sora-2-Pro, if you can handle $3.64 a clip and a seven-and-a-half-minute wait.
- Works simply: Wan-2.7 did exactly what it was told, and it invented nothing.
If you read any of the prior articles, you might remember that I rated a model called Ray-2 as the best AI video generator of its time.
Just a year later, many of those earlier Ai-video generator models no longer exist:
Ray 2, Kling-2.1-Master, all three Pika models, Dream Machine, Veo 2, and Runway Gen-4 Turbo are all gone.
Any video model mentioned today might not exist tomorrow, so enjoy them now. 🙂
Run Your Own Best AI Video Generator Test
I recommend running this test yourself, if you have Poe.
It’s the best way to get an answer with real-world context for your use case.
A bot on Poe called Script-Bot-Creator can run the whole test for you.
There’s a prompt template for it below.
First, put a location in the prompt that you personally know.
Balboa Park is an arbitrary example from the original experiment, not a recommendation.
Ideally you should pick somewhere you can visit yourself.
It could be your street, your local park, or the front of your own building, if it’s a famous landmark.
Open Script-Bot-Creator and paste this brief into a new chat.
The brief now interviews you first.
If you paste it unedited, it stops and asks which location, subject, camera move, length and bots you want, then rewrites the prompt from your answers.
Ask it for the gallery at the end and it builds you a results page with every clip and a download button.
You are running a controlled AI-video benchmark. Precision matters more
than speed or creativity. Call Poe video bots directly and report
structured results.
STEP ZERO: ASK ME WHETHER I CHANGED ANYTHING
Compare the CONTROL PROMPT below against this exact default text:
"Pullback shot depicting a guitar being strummed by a grunge-rock
musician, standing in front of the reflecting pond at Balboa Park in
San Diego, California."
If it still matches word for word, I pasted this brief unedited. Say
so, then ask me all five of these in ONE message, each with your own
recommendation, and WAIT for my answer before doing anything else:
1. The location. This test only means something if you know the
place well enough to judge whether a model got it right. Balboa
Park is an arbitrary example from the original experiment. Think
about a location that you, personally, know. Name a street, a
park, a shopfront, a building you can picture with your eyes
shut.
2. The subject and what they are doing. A grunge-rock musician
strumming a guitar is the original example, not a requirement.
3. The camera move. A pullback is the original example. Push-in,
crane down, orbit and handheld follow are the usual
alternatives.
4. Clip length and resolution. My defaults are the longest each bot
offers up to 10 seconds, 1080p where it is available, and 16:9
landscape always.
5. Which bots to run. My default is every video bot on Poe that
makes its own audio. Say "all with audio", or "everything", or
name a list.
Then rewrite the control prompt from my answers, show it back to me
in full, and get my yes before the first generation. Once I say yes,
that exact string is frozen for the whole run and never edited again.
If the control prompt does NOT match the default above, I have
already edited it myself. Do not ask. Read it back to me once and
start.
FIRST, BEFORE YOU PROPOSE ANYTHING, READ THE BOT LIST
Open this page and read it in full.
URL: https://carletontorpin.com/ai/best-ai-video-generators-fall-2026/
The complete roster of 47 video bots is under the heading "Every AI
Video Generator on Poe". Jump straight to it.
URL: https://carletontorpin.com/ai/best-ai-video-generators-fall-2026/#full-table
Two traps are waiting there. The first table you meet has only TWO
example rows, which is a preview and not the list. The real 47-row
table sits inside the collapsed grey bar beneath it, labeled "Click to
expand the full table". Fetch the page's raw HTML and all 47 rows are
present whether that block is open or shut. If you are driving a real
browser instead, click that grey bar first, and leave the filter box
above the table empty, because anything typed in it hides rows.
That table is the authority on spelling. Poe bot names are exact
strings and a near miss silently fails, so read the row rather than
guessing. "Veo-3.1-Fast" and "Veo-v3.1-Fast" are different bots, and
so are "Seedance-2-Fast" and "Seedance-2.0-Fast-EL". Every name in the
table links to that bot's own poe.com page. Follow the link instead of
building a URL by hand.
The columns are: Bot with its @publisher underneath, Family, cheapest
points per second, dollars per 5 seconds, max resolution, max
duration, whether it makes its own audio, and what it is for. Poe's
conversion at capture time was 33,000 compute points to $1.00.
Higher up that same page, under "What the Best AI Video Generators
Cost", a second table lists the 20 bots that were actually run, with
what each one really charged, how long it really took, and the specs
it really delivered. Use the 47-row table to pick and spell bots. Use
this one for real costs and real render times.
If you cannot read either table, say so and stop. Do not fall back on
bot names you remember.
THE CONTROL PROMPT: use this EXACT text for every generation. Never
rewrite, expand, shorten, "improve," or add camera tags to it:
"Pullback shot depicting a guitar being strummed by a grunge-rock
musician, standing in front of the reflecting pond at Balboa Park in
San Diego, California."
STANDING RULES
1. Text-to-video only. Never attach an image.
2. Audio ON wherever the bot exposes a toggle or generates it natively.
3. Aspect ratio 16:9 LANDSCAPE, set explicitly. Never accept a vertical
default.
4. Use exactly the duration and resolution given per bot. If a value
is unavailable, stop and report rather than guessing.
5. One generation per bot. No re-rolls, no cherry-picking. The first
output is the result, including when it's bad.
6. If a bot refuses or errors, capture the EXACT text it returned. A
refusal is a finding, not a retry.
7. Never substitute a different bot for one that fails.
8. Before running a batch, check whether those bots already have
results in this session. If any do, stop and ask.
REPORT FOR EVERY GENERATION, as a row:
bot | duration requested | duration delivered | resolution delivered
| aspect delivered | audio present Y/N | what the audio contains
| compute points charged | USD | generation time in seconds
| direct video URL | deviations from what was asked
DEVIATIONS ARE THE POINT. Flag any of: aspect ratio changed, duration
different from requested, unrequested cuts or scene changes, added
dialogue or lyrics, watermark, vertical output.
ON AUDIO, also report: is it music, singing, dialogue, or ambient?
Is it synced to the picture? Did the model invent lyrics?
After each batch, output one markdown table of all rows, then a total
for points and USD. Wait for me to send the next batch.
THE GALLERY, IF I ASK FOR IT
When the run is done, if I ask for a gallery, a results page, or just
"the page", build ONE self-contained HTML file and hand it to me as a
download. Do not publish it as an artifact. Artifact security blocks
video from other domains, so not one clip would play.
WHAT GOES ON THE PAGE
Black background, a single column about 1000px wide, base text 17 to
19px. One card per model, in the order they were run. Each card has:
a header with the model name, linked to its own page on poe.com,
with its @publisher beside it
one spec line of plain facts: seconds, resolution, fps, audio sample
rate, render time, compute points, dollars
the clip itself
Facts only on this page. No verdicts, no "stunning", no "impressive".
I supply the opinions.
THE CLIPS
Use a video tag with loop and playsinline and NO controls. One click
plays, another pauses, and starting one pauses all the others. Give
each video an onerror that tells me whether the link expired or was
never captured, with a paste box so I can drop in a replacement.
THE DOWNLOAD BUTTON, WHICH IS THE WHOLE POINT
Poe's video links expire, often within weeks, and at least one model
deletes its output after 48 hours. So put a large button at the top of
the page reading "Copy download command for all clips". Clicking it
copies a ready-to-run block of curl commands to my clipboard, one line
per clip, each saving to a sensible filename such as
01-veo-3-1-lite.mp4. Put a small download link on each card too, for
that one clip.
Say plainly on the page that the links expire and that this button is
how I keep the files.
Do not lean on the download attribute alone for the all-clips button.
Browsers ignore it for files on another domain, so it opens a tab
instead of saving. The copied curl block is the part that always
works.
NAVIGATION
A sticky bar across the bottom with one equal-width chip per model.
Clicking a chip jumps to that card's header rather than the page, so
the title lands in the same spot near the top every time.
Build it as one file with the CSS and the JavaScript inline. Then give
me the file, and tell me how many clips are in it and how many failed. The control prompt is built from four parts.
Swap any of them out to customize your results.
- Camera move: “Pullback shot.” Try push-in, crane down, orbit, or handheld follow.
- Subject: “a grunge-rock musician.” Anything works: a baker, a dog, a kid on a scooter.
- Subject action: “a guitar being strummed.” Whatever your subject is doing.
- Environment: “the reflecting pond at Balboa Park.” Name a real, checkable landmark, so you can compare a ground truth against the AI video output.
Then send the bots in small batches of five models per prompt.
Include the duration and resolution you want for each, or ask the bot to “normalize things across the selection” of ai-video generator models you’ve made.
When the clips come back, download every clip on the same day because those Poe content-delivery-network links expire.
Also, take comfort knowing a failed Poe-video generation generally costs nothing.
Every AI Video Generator on Poe
Here’s a compiled list of 47 ai video bots on Poe, across 15 families, captured September 3rd and 4th, 2026.
“Cheapest points/sec” is each bot’s lowest published rate, so every model sits on one scale.
Poe’s own conversion as of September 2026 is 33,000 compute points = $1.00.
| Bot | Family | Cheapest pts/sec | ≈$ / 5s | Max res | Max sec | Audio | What it’s for |
|---|---|---|---|---|---|---|---|
| Seedance-1.0-Pro-Fast@Bytedance | Seedance | 720 | $0.11 | 1080p | 12 | No | Cheapest Seedance 1.0 tier – best value per token in the whole Seedance line |
| Seedance-1.0-Lite@Bytedance | Seedance | 1,296 | $0.20 | 1080p | 12 | No | Lightweight Seedance 1.0 for simple T2V/I2V |
Click to expand the full table (all 47 AI video generators, 15 families)
No bot matches that.
| Bot | Family | Cheapest pts/sec | ≈$ / 5s | Max res | Max sec | Audio | What it’s for |
|---|---|---|---|---|---|---|---|
| Seedance-1.0-Pro-Fast@Bytedance | Seedance | 720 | $0.11 | 1080p | 12 | No | Cheapest Seedance 1.0 tier – best value per token in the whole Seedance line |
| Seedance-1.0-Lite@Bytedance | Seedance | 1,296 | $0.20 | 1080p | 12 | No | Lightweight Seedance 1.0 for simple T2V/I2V |
| Seedance-1.0-Pro@Bytedance | Seedance | 1,800 | $0.27 | 1080p | 12 | No | Previous-generation Seedance flagship; strong semantic understanding and prompt following |
| Seedance-2-Fast@Bytedance | Seedance | 3,833 | $0.58 | 720p | 15 | Yes | Speed-optimised Seedance 2.0 for rapid iteration and high-volume workflows |
| Seedance-2.0-Fast-EL@empiriolabsai | Seedance | 4,067 | $0.62 | 720p | 15 | Yes | Fast multimodal Seedance with the full mode set at roughly 5% less than Pro-EL |
| Seedance-2.0-Pro-EL@empiriolabsai | Seedance | 4,286 | $0.65 | 4K (3840×2160, 10-bit H.265) | 15 | Yes | Full multimodal Seedance 2.0 Pro – the only bot here with 4K and video edit/extend |
| Seedance-2.0@Bytedance | Seedance | 4,807 | $0.73 | 720p | 15 | Yes | Cinematic T2V/I2V with consistent characters and detailed motion control |
| Veo-3.1-Lite@google | Veo | 1,500 | $0.23 | 1080p | 8 | Yes | Cheapest native-audio video on Poe – $0.05/sec at 720p |
| Veo-v3.1-Fast@fal | Veo | 3,334 | $0.51 | 1080p | 8 | Yes | Cheapest fal Veo endpoint with an explicit silent discount |
| Veo-3.1-Fast@google | Veo | 4,500 | $0.68 | 1080p | 8 | Yes | Veo 3.1 quality at a third of the price – the practical Veo workhorse |
| Veo-v3.1@fal | Veo | 6,667 | $1.01 | 1080p | 8 | Yes | fal's Veo 3.1 endpoint – explicit audio/silent price split and first-to-last-frame support |
| Veo-3-vFast@fal | Veo | 8,334 | $1.26 | Not exposed | 7 | Yes | Legacy Veo 3 fast endpoint – text-to-video only, fixed 7 seconds |
| Veo-3.1@google | Veo | 13,333 | $2.02 | 1080p | 8 | Yes | Google flagship – native audio, reference mode, strong prompt adherence |
| Kling-2.1-Std@fal | Kling | 1,667 | $0.25 | Not exposed | 10 | No | Cheapest Kling endpoint – cost-efficient image-to-video |
| Kling-2.6-Pro@fal | Kling | 2,334 | $0.35 | Not exposed | 10 | Yes | Best value Kling with native audio – $0.071/sec silent, $0.14/sec with sound |
| Kling-2.5-Turbo-Pro@fal | Kling | 2,334 | $0.35 | Not exposed | 10 | No | Fast, cheap silent Kling for image-driven motion |
| Kling-2.1-Pro@fal | Kling | 2,834 | $0.43 | Not exposed | 10 | No | Cinematic image-to-video with precise camera movement and motion control |
| Kling-Pro-Effects@fal | Kling | 3,334 | $0.51 | Not exposed | 10 | No | Canned photo effects – squish, expand, hug, kiss, heart gesture |
| Kling-Omni@fal | Kling | 3,734 | $0.57 | Not exposed | 10 | No | Image-to-video and first-to-last-frame at a flat mid-range price |
| Kling-v3-Motion-Ctrl@empiriolabsai | Kling | 4,667 | $0.71 | 1080p | 30 | No | Motion transfer – drive a character from a still image with a reference video's movement |
| Kling-O3@empiriolabsai | Kling | 5,600 | $0.85 | 4K | 15 (10 when any video input is used) | Yes | Multi-scene storytelling in a single prompt – up to 6 scenes with per-scene timing |
| Kling-2.0-Master@fal | Kling | 6,000 | $0.91 | Not exposed | 10 | No | Legacy Kling 2.0 Master; supports CLI-style flags in the prompt |
| Kling-v3-Pro@fal | Kling | 7,467 | $1.13 | Not exposed | 15 | Yes | Newest Kling flagship – T2V/I2V with start+end frames, native audio, 15s max |
| Wan-2.6@empiriolabsai | Wan | 750 | $0.11 | 1080p | 10 for R2V | Yes | Multi-shot storytelling with lip-sync; Flash mode is the cheapest video generation on Poe |
| Wan-2.5@fal | Wan | 1,667 | $0.25 | 1080p | 10 | No | Simple Wan endpoint with audio-guided generation from an uploaded mp3 |
| Wan-3.0@empiriolabsai | Wan | 2,333 | $0.35 | 1080p | 30 | Yes | Newest Wan – the longest duration range here (2-30s) plus a speed tier |
| Wan-Animate@fal | Wan | 2,500 | $0.38 | Source-dependent | Source-dependent | No | Character replacement in an existing video, or motion transfer onto a still |
| Wan-2.7@empiriolabsai | Wan | 3,333 | $0.51 | 1080p | 10 for video-input modes | Yes | Four-mode Wan with true video editing and voice-timbre reference |
| Sora-2@openai | Sora | 3,000 | $0.45 | 720p | 20 (panel) | Yes | Realistic physics, synchronized dialogue and SFX, multi-shot prompt adherence |
| Sora-2-Pro@openai | Sora | 9,000 | $1.36 | 1920×1080 / 1792×1024 | 20 | Yes | Highest-fidelity Sora – world-state persistence, complex multi-shot, up to 20s |
| Gemini-Omni-Flash@google | Gemini Omni | token-metered | n/a | Not exposed | Not exposed | Yes | Conversational video editing – generate then refine in natural language |
| Gemini-Omni-1.1-Flash@google | Gemini Omni | token-metered | n/a | 4K | Not exposed | Yes | Scene extension and first/last-frame interpolation with 4K output |
| Pixverse-v4.5@fal | PixVerse | 2,000 | $0.30 | 1080p | 8 | No | Meme-style canned effects (Hulk, Venom, Kiss Me AI) plus style presets |
| Pixverse-v5.6@empiriolabsai | PixVerse | 2,333 | $0.35 | 1080p | 10 (8 at 1080p) | Yes | Style-preset video with optional audio; the deepest control surface of the PixVerse bots |
| Pixverse-v5@empiriolabsai | PixVerse | 3,000 | $0.45 | 1080p | 8 (5 at 1080p) | No | Three clean modes decided purely by how many images you attach |
| Hailuo-02-Standard@fal | Hailuo (MiniMax) | 1,500 | $0.23 | 768p | 10 | No | Cheap MiniMax image-to-video at 768p |
| Hailuo-02-Pro@fal | Hailuo (MiniMax) | 2,667 | $0.40 | 1080p | 5 | No | 1080p MiniMax image-to-video, fixed 5-second clips |
| Hailuo-AI@fal | Hailuo (MiniMax) | 2,833 | $0.43 | Not exposed | Not exposed | No | MiniMax's general text/image-to-video, billed as a flat fee per message |
| Hailuo-Director-01@fal | Hailuo (MiniMax) | 3,333 | $0.51 | Not exposed | 5 | No | THE camera-control bot – explicit shot moves in square brackets |
| LTX-2-Fast@fal | LTX | 1,334 | $0.20 | 2160p (4K) | 10 | Yes | Cheapest route to 4K (2160p) and the only 50 FPS option here |
| LTX-2-Pro@fal | LTX | 2,000 | $0.30 | 2160p (4K) | 10 | Yes | Professional-grade LTX-2 up to 2K/4K with 50 FPS |
| Grok-Imagine-Video@xAI | Grok Imagine | 1,667 | $0.25 | Not exposed | Slider max not labelled | No | Artistic/creative video plus mp4 video editing; widest aspect-ratio menu here |
| Runway-Gen-4.5@runwayml | Runway Gen | 10,000 | $1.52 | Not exposed | 10 | No | High-end cinematic fidelity with fine-grained motion and camera behaviour |
| Amazon-Nova-Reel-1.1@empiriolabsai | Nova Reel | 4,800 | $0.73 | 720p | 120 (2 minutes) | No | Up to TWO-MINUTE multi-shot videos with per-shot prompting |
| Vidu@fal | Vidu | 1,333 | $0.20 | Not exposed | 5 | No | ~170 one-click meme/effect templates – by far the biggest preset library here |
| SVI-2.0-Pro@empiriolabsai | Stable Video Infinity | 1,910 | $0.29 | 720p | 121.5 (480p) | No | LONG-FORM video – up to 121.5 seconds in a single generation, the longest here by far |
| OmniHuman@Bytedance | OmniHuman | 4,667 | $0.71 | Not exposed | 30 | No | Audio-driven talking/singing avatar from one photo |
Poe changes rates often, so check the Rates dialog before budgeting anything.
See Also: The Best AI Image Generators
Here’s the same experiment, but for AI image generators instead of video.
What Is The Best Ai Image Generator? tests 29 image models on photorealism, spatial reasoning, historical accuracy, and text rendering.
There’s more of it around the site: the AI section, the Photo section, etc.
Fun to Notice
This is the fourth time in two years that I’ve tested the best AI video generators.
The time it takes me to make each article keeps dropping, too.
The first article, in Fall 2024, took about 12 hours, from start to finished article, companion video included.
The second article, in March 2025, took 6 hours, once AI could handle enough of the video generation on its own.
The third article, in January 2026, took about 4 hours, and so did this one.
There will certainly be a point in the next two years where this will be a test I can run in seconds, complete with an article I can author and post in under a minute.
Have Fun With AI Video
Twenty models, one prompt, Every clip is clickable.
Seedance-2-Fast
Seedance-2.0-Fast-EL
Seedance-2.0-Pro-EL
Seedance-2.0
Kling-2.6-Pro
Kling-O3
Kling-v3-Pro
Veo-3.1-Lite
Veo-3.1-Fast
Veo-v3.1-Fast
Veo-v3.1
Veo-3.1
Sora-2
Sora-2-Pro
Wan-2.5
Wan-2.6
Wan-2.7
Wan-3.0
Pixverse-v5.6
Gemini-Omni-1.1-Flash
Have fun making stuff, and share your results!
Previously: Text to Video AI with Poe (Oct 2024) · Best Ai Video Generators in Poe Ai (Mar 2025) · Ai Video Generation Models Compared (Jan 2026)
Related: What Is The Best Ai Image Generator? · Ai Audio Generators on Poe · Script Bot Creator Poe Ai
Costs from Poe’s billing ledger, September 4th, 2026. Specs measured with ffprobe on the delivered files, not read off a product page. One output per model on one prompt, which is hardly representative of any model in general.