AI video generation/4K output
Describe the shot.Get the shot.
Text, stills and reference footage go in. Finished cuts with sound come out.
- 22
- video models
- 4K
- on 7 of 22
- 60s
- longest clip
- 7/22
- ship with audio
[01] Ways in
Six doors, one workspace.
Every workflow below has its own page, its own model defaults, and its own set of controls. Start wherever your material already is.
Text to video
Turn a written scene direction into a short AI video.
- Start with
- A written scene and motion brief
- Create
- A short generated video
Image to video
Start from a still image, then direct the motion you want to see.
- Start with
- A supported reference image
- Create
- A short motion-guided video
Source-video transformation
Use a source video and a written direction to create a new video version.
- Start with
- A supported source video
- Direct with
- A written transformation direction
Multi-shot advertising video
Turn a campaign beat sheet into a multi-shot AI ad video.
- Plan
- Hook, product proof, and closing shot
- Model anchor
- Kling 3.0 multi-shot and sound
Vertical 9:16 short video
Create a 9:16 AI video composed for the phone screen.
- Frame
- Native 9:16 vertical composition
- Plan for
- Reels, Shorts, and story placements
First-and-last-frame animation
Animate the transition between a planned first frame and last frame.
- Start with
- An approved opening frame and ending frame
- Keep clear
- The transition, subject continuity, and camera path
[02] The pipeline
How to create a video from text
A short, focused direction is easier to inspect than a prompt that asks for several unrelated scenes at once.
Write the subject and action first
State what is visible, what happens, and where it happens before adding atmosphere or camera direction.
Add camera and timing cues
Include the viewing angle, movement, speed, or mood that should guide the clip.
Choose a compatible video model
Use the current model controls to select a frame, duration, and quality level that are available for that model.
Generate and inspect the whole clip
Review the opening, motion, timing, visible text, and final frame before using the result publicly.
01 / TEXT TO VIDEO
WRITE IT. WATCH IT MOVE.
A sentence is a shot list. Name the subject, the lens, the light — get back footage with sound already on it.
Open the workflow →02 / IMAGE TO VIDEO
ONE STILL. THIRTY SECONDS.
Drop in a frame you already own and let the model find the movement that was implied in it.
Open the workflow →Source-video transformation
Keep the motion. Change the look.
Choose the source-video workflow when an uploaded clip provides the movement and timing you want to begin from. Add the supported video, describe the visual transformation, then inspect the generated version as a new result before reuse.
A movement-led starting point
Use an approved source clip when its movement or timing gives the next version a useful foundation.
A specific visual transformation
Explain what should change in the new version, such as scene treatment, material feeling, or overall visual direction.
A source-aware review step
Compare the result with the source clip before relying on continuity, product details, or exact visible information.
[03] The board
Twenty-two models, fully spec’d.
| Model | Max res | Duration | Ratios | Audio | Input |
|---|---|---|---|---|---|
| Seedance 2 MiniAudio, web search, and multimodal reference | 720p | 4–15s | 6 | YES | Text or image |
| Seedance 2 FastFast multimodal video with audio and web search | 720p | 4–15s | 6 | YES | Text or image |
| Seedance 2Multimodal video with audio, web search, and up to 4K | 4k | 4–15s | 6 | YES | Text or image |
| Seedance 2.5Long-form multimodal video with first and last frames | 720p | 1–30s | 6 | YES | Text or image |
| HappyHorse 1.1Text or image to video at 720p or 1080p | 1080p | 3–15s | 9 | — | Text or image |
| HappyHorseText or first frame to video at 720p or 1080p | 1080p | 3–15s | 5 | — | Text or image |
| Kling 2.6Text or image to video with sound · 5 or 10 seconds | 720p | 5–10s | 3 | YES | Text or image |
| Kling V3 TurboText or image to video at 720p or 1080p | 1080p | 3–15s | 3 | — | Text or image |
| Kling 3.0 VideoMulti-shot video with sound at up to 4K | 4k | 3–15s | 3 | YES | Text or image |
| Kling O3Text or image to video with native audio | 4k | 3–15s | 3 | YES | Text or image |
| Wan 2.6 VideoPrompt-guided source video transformations | 1080p | 5–10s | — | — | Text or image |
| Kling 2.6 Motion ControlTransfer a driving video motion to a reference character | 1080p | 3–30s | — | — | Image required |
| Kling 3.0 Motion ControlTransfer a driving video motion to a reference character | 1080p | 3–30s | — | — | Image required |
| Wan 2.7 VideoText or frames, clip, and driving audio to video | 1080p | 2–15s | 5 | — | Text or image |
| Grok Imagine VideoText or one image to video · 480p or 720p | 720p | 6–30s | 5 | — | Text or image |
| Grok Imagine Video 1.5 PreviewText or images to video at 480p or 720p | 720p | 1–15s | 5 | — | Text or image |
| Gemini Omni VideoPrompt with optional images at up to 4K | 4k | 4–10s | 2 | — | Text or image |
| MiniMax H3Text or first-frame video at 768P or 2K | 2k | 4–15s | 6 | — | Text or image |
| Veo 3.1 QualityHighest-fidelity text or first-and-last-frame video | 4k | 4–8s | 2 | — | Text or image |
| Veo 3.1 FastFaster text or first-and-last-frame video | 4k | 4–8s | 2 | — | Text or image |
| Veo 3.1 LiteCost-efficient text or first-and-last-frame video | 4k | 4–8s | 2 | — | Text or image |
| OmniHuman 1.5Portrait image and audio driven video | 1080p | 1–60s | — | — | Image required |
Showing 22 of 22 video models
[04] Controls
What every dial does.
Resolution
480p → 4K
Set per model. Only 7 of the 22 reach 4K.
Duration
1s → 60s
Some models expose fixed lengths, others a full range.
Aspect ratio
9 max
Chosen at generation time, not cropped afterwards.
Audio
7/22
Generated on the same track as the picture.
Prompt length
30k
Long enough for a full shot description, not a slogan.
Start/end frames
5/22
Pin where the shot begins and where it lands.
[R] Framing
Framed for where it lands.
Pick the aspect ratio at generation time instead of cropping a 16:9 master afterwards and losing the composition. Support varies by model — the count under each frame is the current one.
9:16
Shorts, Reels, TikTok
18/22 models
1:1
Feed and marketplace
14/22 models
16:9
YouTube and web
18/22 models
21:9
Cinematic and hero
6/22 models
22
video models
60s
longest single clip
4K
top resolution
7
ship with audio
5
take start/end frames
[05] Made here
Twelve, playing at once.
[W] Why here
Twenty-two subscriptions, or one.
01
One account, every model
Seedance, Kling, Veo, Wan, Grok, MiniMax and the rest sit behind the same prompt box and the same credit balance. No twenty subscriptions, no twenty dashboards.
02
Limits published up front
Every model page lists its resolutions, durations, framing options and whether audio comes with it — before you spend a credit finding out.
03
The whole take is kept
Prompt, reference, model and settings stay attached to every result, so a direction that worked can be returned to instead of reconstructed.
[L] Before you publish
What to check before using text-to-video output
01
A prompt does not lock every visible detail
Generated motion, objects, text, and timing can vary, so important details still need a full visual review.
02
Controls vary by video model
The available duration, frame, quality, sound, and reference-media options depend on the model you select.
03
Only publish material you can use
Review the finished clip and use only source material, visual references, and claims you are allowed to publish.
[C] Catalog
Read before you generate.
[$] Credits
One balance, every model.
Every new account starts with 50 free credits. Video models cost more than image models — the current price is shown before you generate.
Free
Get started
$0
No payment method required
50 welcome credits
- 50 welcome credits at sign-up
- Start with available models
- No subscription required
Starter
For occasional projects
$9.90/ month
Billed monthly. Cancel before the next renewal.
1,000 credits every month
Regular rate: approx. $0.99 per 100 credits
- Use credits for image and video generation
- Choose any available model and output settings
- Commercial use
- Unused subscription credits roll over and remain available.
Pro
For frequent work
$19.90/ month
Billed monthly. Cancel before the next renewal.
2,000 credits every month
Regular rate: approx. $0.99 per 100 credits
FREE GIFT
+200 credits · one-time on your first subscription
First month total: 2,200 credits
- Use credits for image and video generation
- Choose any available model and output settings
- Commercial use
- Unused subscription credits roll over and remain available.
Ultimate
For continuous production
$49.90/ month
Billed monthly. Cancel before the next renewal.
5,000 credits every month
Regular rate: approx. $1 per 100 credits
FREE GIFT
+500 credits · one-time on your first subscription
First month total: 5,500 credits
- Use credits for image and video generation
- Choose any available model and output settings
- Commercial use
- Unused subscription credits roll over and remain available.
Need extra credits?
Add a one-time credit pack whenever you need more capacity. These credits can be used alongside a subscription and never expire.
Credits roll into one balance and work across every image and video model.
[F] Frequently asked
Asked before you ask.
Every answer here is the same one on the workflow page it came from, so nothing on this page contradicts the documentation.
A text-to-video generator creates a short video from a written scene direction. In ImagineClip, the prompt and selected video model determine which inputs and output controls are available.
Start with the subject, action, setting, and camera direction. Add timing, light, atmosphere, or style only after the core motion is clear.
You can start from an image when you choose a video model that accepts an image reference. The supported formats, limits, and output controls are shown by the active model.
Focus on the change over time: action, camera movement, timing, environment, and visual feeling. The image already provides the starting appearance, so the prompt should explain how it should develop.
A video-to-video workflow starts from a supported source clip and uses a written direction to create a new generated version. The result should be reviewed as a new video, not as an exact edit of the source.
Yes. ImagineClip includes a source-video workflow for compatible video models. Upload requirements and available output controls are shown in the active workspace.
It turns a campaign beat sheet into a sequence with distinct shots rather than one continuous generic scene. ImagineClip exposes Kling 3.0 multi-shot controls, while final claims still require review.
Kling 3.0 Video supports multi-shot generation and generated sound, with an available route to 4K delivery. Available controls are shown in the workspace.
It creates video in a native vertical frame for phone-first placements. ImagineClip provides compatible 9:16 model controls while final captions and publishing details remain separate.
Kling 3.0, Seedance 2.5, Wan 2.7, and several other current models expose 9:16. Duration, sound, reference inputs, and quality still vary by model.
It uses one image as the opening state and another as the ending state, then generates the motion between them. The two images guide the endpoints but do not guarantee every middle frame.
Seedance 2.5, Veo 3.1 variants, and Wan 2.7 expose first-and-last-frame workflows in ImagineClip. Their duration, quality, sound, and reference options differ.