Pricing
Credits per model
Every generation is priced from the same table the dashboard and API use, so the estimate you see before generating is what gets deducted. Video is billed per second of output; images per image; audio per generation. Minimum charge is 5 credits. Failed generations are refunded automatically.
Video models
| Model | Description | Credits / second | Note |
|---|---|---|---|
| H3 Max | Generate video from a prompt or an image. | 38/s (480P) · 60/s (768P) · 120/s (1080P) | e.g. 5s at default res ≈ 600 |
| H3 Max Turbo | Generate video from a prompt or an image. | 19/s (480P) · 30/s (768P) · 60/s (1080P) | e.g. 5s at default res ≈ 300 |
| Wan 3.0 | Generate video from a prompt or an image. | 38/s (480p) · 75/s (720p) · 150/s (1080p) | e.g. 5s at default res ≈ 750 |
| LTX 2.5 Pro | Generate video from a prompt or an image. | 90/s (720p) · 128/s (1080p) | e.g. 6s at default res ≈ 765 |
| LTX 2.5 Fast | Generate video from a prompt or an image. | 68/s (720p) · 98/s (1080p) | e.g. 6s at default res ≈ 585 |
| Grok Imagine Video 1.5 | Generate video from a prompt or an image. | 68/s (480p) · 113/s (720p) · 195/s (1080p) | e.g. 5s at default res ≈ 975 |
| PixVerse V6 | Generate video from a prompt or an image. | 26/s (360p) · 34/s (540p) · 45/s (720p) · 86/s (1080p) | e.g. 5s at default res ≈ 431 |
| Happy Horse 1.1 | Generate video from a prompt or an image. | 105/s (720p) · 135/s (1080p) | e.g. 5s at default res ≈ 675 |
| Veo 3.1 | Cinematic quality + audio | 300/s | e.g. 4s ≈ 1200 |
| Veo 3.1 Fast | Fast generation | 113/s | e.g. 4s ≈ 450 |
| Gemini Omni Flash | Google - text/image → video, plus conversational video edit | 98/s | e.g. 3s ≈ 293 |
| Hailuo 2.3 | MiniMax - image-to-video, 768p, expressive motion | 42/s | e.g. 6s ≈ 252 |
| Kling 3.0 Pro | Cinematic + native audio - multi-shot | 42/s | e.g. 5s ≈ 210 |
| Kling 3.0 Standard | Cinematic + native audio - budget | 21/s | e.g. 5s ≈ 105 |
| Kling 2.6 Pro | Native audio - high quality | 42/s | e.g. 5s ≈ 210 |
| Kling 2.6 Motion Control Pro | Transfers motion from a reference video - up to 30s | 42/s | e.g. 5s ≈ 210 |
| Kling 2.6 Motion Control Standard | Motion transfer - budget | 21/s | e.g. 5s ≈ 105 |
| Seedance 2.5 | ByteDance - native 30s single-pass, joint audio-video, best lip-sync | 165/s (480p) · 355/s (720p) | e.g. 4s at default res ≈ 1419 |
| Seedance 2.5 Ref2Video | ByteDance - up to 30 image refs (50 files total) → one coherent 30s clip, native audio | 165/s (480p) · 355/s (720p) | e.g. 4s at default res ≈ 1419 |
| Wan 2.7 | Alibaba - first/last frame control, 2-15s, audio-driven, open weights | 75/s (720p) · 113/s (1080p) | e.g. 3s at default res ≈ 225 |
| Seedance 2.0 | ByteDance - unified audio-video, lip-sync, cinematic quality | 101/s (480p) · 227/s (720p) · 511/s (1080p) · 1167/s (4k) | e.g. 4s at default res ≈ 906 |
| Seedance 2.0 Fast | ByteDance - fast generation, lower latency | 182/s | e.g. 4s ≈ 726 |
| Seedance 2.0 Ref2Video | ByteDance - up to 9 ref images → one coherent multi-shot clip, native audio | 101/s (480p) · 227/s (720p) · 511/s (1080p) · 1167/s (4k) | e.g. 4s at default res ≈ 906 |
| Seedance v1 Lite | ByteDance - fast and lightweight | 15/s | e.g. 5s ≈ 75 |
| Kling O3 Pro | Omni - multi-image - up to 15s | 105/s | e.g. 5s ≈ 525 |
| Seedance v1.5 Pro | ByteDance - native audio, cinematic quality | 60/s | e.g. 5s ≈ 300 |
| Seedance v1 Pro Fast | ByteDance - fast, 1080p native | 30/s | e.g. 5s ≈ 150 |
| Seedance Ref2Video | ByteDance - up to 4 ref images to video, character preservation | 27/s | e.g. 5s ≈ 135 |
| WAN Replace | WAN 2.2 14B - inserts character into source video preserving scene | 60/s | e.g. 5s ≈ 300 |
| WAN Move | WAN 2.2 14B - transfers motion and expressions, preserves identity | 60/s | e.g. 5s ≈ 300 |
| Grok Imagine Video | Fast video generation with audio, up to 10s | 38/s | e.g. 5s ≈ 188 |
| HeyGen Avatar V5 | HeyGen - V5 digital twin, TTS or audio_url, up to 4K, MP4 or WebM | 113/s | e.g. 5s ≈ 563 |
| HeyGen Avatar 4 | HeyGen - V4 digital twin, TTS or audio_url | 75/s | e.g. 5s ≈ 375 |
| HeyGen Avatar 3 (budget) | HeyGen - V3 digital twin, fast + cheap, TTS or audio_url | 26/s | e.g. 5s ≈ 128 |
| HeyGen Avatar 4 (face image) | HeyGen - animate any face image with TTS or audio_url | 75/s | e.g. 5s ≈ 375 |
| HeyGen V3 Video Agent | HeyGen - one prompt, agent picks avatar + voice + script, attach up to 20 refs | 26/s | e.g. 5s ≈ 128 |
| HeyGen V2 Video Agent | HeyGen - prompt → talking-avatar video, V2 model | 26/s | e.g. 30s ≈ 765 |
| HeyGen Lipsync Precision | HeyGen - high-accuracy avatar-inference lipsync, replaces audio on an existing video | 75/s | e.g. 5s ≈ 375 |
| HeyGen Lipsync Speed | HeyGen - fast audio-only lipsync, ~½ the price of Precision | 38/s | e.g. 5s ≈ 188 |
| Kling AI Avatar v2 | Animate a portrait image to speak from audio | 42/s | e.g. 5s ≈ 211 |
| Sync LipSync 2 Pro | High-fidelity lip-sync on an existing video + audio, up to 4K | 62/s | e.g. 5s ≈ 311 |
| HeyGen Translate Precision | HeyGen - high-precision video dubbing into another language | 75/s | e.g. 5s ≈ 375 |
| HeyGen Translate Speed | HeyGen - fast video dubbing into another language, ~½ price of Precision | 38/s | e.g. 5s ≈ 188 |
| Topaz Video Upscale | Topaz Labs - AI video upscaling up to 4x | 19/s | e.g. 5s ≈ 94 |
Image models
| Model | Description | Credits / image | Note |
|---|---|---|---|
| GPT Image 2.5 Flare | Fast generation and precise edits with transparent backgrounds. | 23 (low) · 30 (medium) · 60 (high) · 98 (xhigh) · 188 (max) | per image, by quality |
| GPT Image 2.5 Flare Edit | Fast generation and precise edits with transparent backgrounds. | 23 (low) · 30 (medium) · 60 (high) · 98 (xhigh) · 188 (max) | per image, by quality |
| GPT Image 2.5 Sunburst | Detailed images and careful edits for intricate visual work. | 23 (low) · 30 (medium) · 60 (high) · 98 (xhigh) · 188 (max) | per image, by quality |
| GPT Image 2.5 Sunburst Edit | Detailed images and careful edits for intricate visual work. | 23 (low) · 30 (medium) · 60 (high) · 98 (xhigh) · 188 (max) | per image, by quality |
| Grok Imagine Image 2.0 | Detailed imagery and typography from xAI. | 30 (low) · 45 (medium) | per image, by quality |
| Grok Imagine Image 2.0 Edit | Detailed imagery and typography from xAI. | 38 (low) · 53 (medium) | per image, by quality |
| Qwen Image 3 | Precise prompts, multilingual text and image editing. | 56 | per image |
| Qwen Image 3 Edit | Precise prompts, multilingual text and image editing. | 56 | per image |
| Ideogram 4 | Posters, logos and crisp typography. | 23 | per image |
| Krea 2 Large | Expressive aesthetics and creative image generation. | 45 | per image |
| Flux 2 Max | High-quality image generation and reference-guided editing from Black Forest Labs. | 75 | per image |
| Flux 2 Max Edit | High-quality image generation and reference-guided editing from Black Forest Labs. | 98 | per image |
| Flux 2 Flex | High-quality image generation and reference-guided editing from Black Forest Labs. | 75 | per image |
| Flux 2 Flex Edit | High-quality image generation and reference-guided editing from Black Forest Labs. | 113 | per image |
| Flux Pro | Black Forest Labs - max quality, 12B params | 41 | per image |
| Flux Dev | Black Forest Labs - open-weights, flexible | 19 | per image |
| Flux Pro Ultra | Black Forest Labs - ultra HD, up to 4K | 45 | per image |
| Flux 2 | Black Forest Labs - second generation | 23 | per image |
| Flux 2 Pro | Black Forest Labs - Flux 2 Pro, top quality | 41 | per image |
| Flux 2 LoRA | Black Forest Labs - Flux 2 with custom LoRA | 26 | per image |
| Seedream 4.5 | ByteDance - gen+edit unified, up to 4K | 23 | per image |
| Seedream 5.0 Lite | ByteDance - Seedream 5 fast, up to 4K | 23 | per image |
| Seedream 5.0 Pro | ByteDance - flagship Seedream 5, native text in 14 languages, dense layouts | 101 | per image |
| Nano Banana | Gemini 2.5 Flash - text+images to image, fast | 29 | per image |
| Nano Banana Pro | Gemini 3 Pro - reasoning, 4K, up to 14 refs | 113 | per image |
| Nano Banana 2 | Gemini 3.1 Flash - reasoning, text rendering, character consistency | 60 | per image |
| Nano Banana 2 Lite | Gemini - fast, cheap 1K text-to-image (sub-2s) | 38 | per image |
| Nano Banana Lite | Gemini - fast, cheap 1K text-to-image | 38 | per image |
| Nano Banana Lite Edit | Gemini - fast image edit from one or more references | 38 | per image |
| Ideogram v3 | Ideogram - best-in-class typography + graphic design | 68 | per image |
| Qwen Image | Tongyi - strong text rendering + editing, 1MP | 19 | per image |
| Z-Image Turbo | Tongyi-MAI - 6B params, super fast, 1MP | 6 | per image |
| GPT Image 2 | OpenAI - text rendering + photorealism, up to 4K | 5 (low) · 40 (medium) · 158 (high) | per image, by quality |
| GPT Image 1.5 | High fidelity, GPT-5 architecture | 7 (low) · 26 (medium) · 100 (high) | per image, by quality |
| GPT Image 1 | Excellent text rendering | 8 (low) · 32 (medium) · 125 (high) | per image, by quality |
| Grok Imagine | Highly aesthetic generation, up to 2K | 15 | per image |
| GPT Image 2 Edit | OpenAI - photorealistic image-to-image with optional mask | 5 (low) · 40 (medium) · 158 (high) | per image, by quality |
| GPT Image 1.5 Edit | Image-to-image, preserves composition | 7 (low) · 26 (medium) · 100 (high) | per image, by quality |
| Grok Imagine Edit | Precise image editing | 15 | per image |
| Flux Kontext | BFL - edit with text + reference image | 30 | per image |
| Seedream 4.5 Edit | ByteDance - multi-image editing (max 10 refs) | 30 | per image |
| Flux 2 Edit | BFL - Flux 2 image-to-image editing | 30 | per image |
| Flux 2 Pro Edit | BFL - Flux 2 Pro image editing, top quality | 53 | per image |
| Flux 2 LoRA Edit | BFL - Flux 2 image editing with custom LoRA | 34 | per image |
| Seedream 5.0 Lite Edit | ByteDance - multi-image editing Seedream 5 (max 10 refs) | 30 | per image |
| Seedream 5.0 Pro Edit | ByteDance - region-precise editing, sketch/annotation guidance (max 10 refs) | 101 | per image |
| FASHN Try-On | Virtual try-on (ref 1 = person, ref 2 = garment) | 23 | per image |
| Kling Kolors Try-On | Kling - virtual try-on (ref 1 = person, ref 2 = garment) | 15 | per image |
| CAT-VTON | Virtual try-on (ref 1 = person, ref 2 = garment) | 15 | per image |
| ESRGAN 4x | Classic AI upscaling up to 4x | 5 | per image |
| Clarity Skin | Upscaling + skin and texture enhancement | 38 | per image |
| Topaz Upscale | Topaz Labs - professional AI upscaling up to 4x | 9 | per image |
| CCSR Upscaler | Content-Consistent Super Resolution - precise details | 5 | per image |
| Crystal Upscaler | ClarityAI - ultra-sharp crystal clarity upscaling | 38 | per image |
Audio models
| Model | Description | Credits | Note |
|---|---|---|---|
| Lyria 3 Pro | Google — full 3-minute songs with vocals and timed lyrics | 60 | per generation |
| ElevenLabs Music | Long-form, up to 10 min — vocals or instrumental-only | 600 | per generation |
| Lyria 3 | Google — 30-second clip, same engine as Pro | 30 | per generation |
| Stable Audio 3 | Instrumental beds up to 6 min — commercial-safe | 29 | per generation |
| MiniMax Music | Song generation from a style prompt + lyrics | 113 | per generation |
| ACE-Step | Fastest and cheapest — draft quality | 5 | per generation |
| ElevenLabs SFX | Sound effects 0.5–22s, seamless looping | 33 | per generation |
| Mirelo SFX | Up to 60s — ambience mode for loopable room tone | 75 | per generation |
| ElevenLabs v3 | Most expressive multilingual speech | 75 | per generation |
| Seed Speech v2 | ByteDance — a third the cost, steerable delivery | 23 | per generation |