Grok Imagine Video 1.5 — xAI's #1Image-to-Video Model with Native Audio
Generate cinematic 720p/24fps videos up to 15 seconds from text prompts and reference images. Native audio — dialogue, sound effects, and ambience — generated in a single pass with the Aurora autoregressive engine.
See What Grok Imagine Video 1.5 Can Create
From native audio-driven cinematic dialogue to anime motion, commercial transitions, and fantasy worldbuilding — explore the kind of stunning videos Grok Imagine Video 1.5 can generate from text prompts and reference images with 720p/24fps output.
Native Audio & Speech — All in One Pass
Dialogue, sound effects, and ambience are generated together with video — not dubbed in later. Speech lands on the action, clearer and better synced.
1950s hotel elevator. A woman in an emerald gown speaks to the operator in a red uniform as the gold-trimmed doors close. Soft dramatic lighting, rich film colors.
Dynamic Commercial-to-Set Transitions
Create engaging behind-the-scenes and transition-focused marketing content. Grok Imagine Video 1.5 smoothly transitions from pristine commercial product shots to complex studio sets with fully synchronized foley.
A continuous pull-out shot. A hand pours milk into a mason jar of iced coffee next to a stack of cookies. The camera pulls back dynamically, revealing a woman taking a bite of the cookie with a synchronized crunch, on a busy green-screen production studio set with crew.
Stylized Anime & Motion Consistency
Render vibrant anime art and fluid character motion. Grok Imagine Video 1.5 maintains flawless character details, complex fabric physics, and expressive facial acting across dynamic stylized shots.
Stylized anime 3D animation. A cute blue-haired elf girl with red eyes, wearing a black cyberpunk Qipao with a blue dragon print and tactical straps, dances playfully in front of a traditional temple with red lanterns. Smooth fluid movement, expressive winks and smiles, vibrant lighting, highly detailed.
Multi-Agent Physics & Interactions
Simulate hyper-realistic animal locomotion and chaotic city physics. Grok Imagine Video 1.5 seamlessly coordinates natural animal movement, flocking bird dynamics, and volumetric steam in crowded environments.
A rabbit sprinting through NYC, fast-paced, photorealistic.
Narrative Character Growth & Worldbuilding
Deliver continuous character evolution and epic world-scale transitions. Grok Imagine Video 1.5 simulates biological growth — like hatching and aging — while maintaining character identity across vast, physics-rich fantasy environments.
A cinematic fantasy sequence. A cute white baby dragon hatches from a shimmering, iridescent egg surrounded by glowing crystals. The dragon grows and spreads its wings on a cliffside, then takes off to fly smoothly through fluffy clouds. It transitions into soaring majestically toward a breathtaking sunset over a vast landscape of floating islands and giant crystal spires.
Grok Imagine Video 1.5 vs Seedance 2.0: AI Video Model Comparison
Both Grok Imagine Video 1.5 and Seedance 2.0 are top-tier image-to-video models with native audio, but they serve different priorities. Grok Imagine Video 1.5 prioritizes generation speed and single-pass audio-visual coherence. Seedance 2.0 prioritizes reference depth and multi-shot control.
Grok Imagine Video 1.5 — built on Aurora autoregressive (MoE) engine. Generates 720p video with native audio in a single pass at ~25s for a 6s clip.
Seedance 2.0 — Dual Branch Diffusion Transformer. Excels at multi-shot storytelling with broader reference input support and 1080p output.
| Comparison Point | Grok Imagine Video 1.5 | Seedance 2.0 | Key Difference |
|---|---|---|---|
| Developer | xAI | ByteDance | Different research teams and architectures |
| Architecture | Aurora autoregressive (MoE) | Dual Branch Diffusion Transformer | Grok uses autoregressive; Seedance uses diffusion-based generation |
| Generation Speed (6s clip) | ~25s (Fast mode) | ~120s | Grok Imagine Video 1.5 is ~5× faster |
| Max Resolution | 720p | 1080p | Seedance offers higher max resolution |
| Max Duration | 15s | 15s | Both support up to 15-second clips |
| Native Audio Output | Single-pass: dialogue, SFX, ambience | Dialogue, SFX, lip-sync | Both deliver complete audio-visual generation |
| Input Type | Image + Text prompt | Image + Text + Multi-ref support | Seedance accepts more reference images per generation |
| Arena Leaderboard (I2V) | #1 (May 2026) | #2 | Grok Imagine Video 1.5 currently leads |
| Best For | Fast I2V, native audio, rapid iteration | Reference depth, multi-shot, 1080p output | Grok for speed; Seedance for reference variety |
Grok Imagine Video — Model Evolution Timeline
From the launch of xAI's first image-to-video model to the Arena-topping 1.5 — here's how Grok Imagine Video has evolved.
Grok Imagine Video 1.0 Launch
xAI launched its first dedicated image-to-video model — separate from the Grok chatbot. Built on the proprietary Aurora autoregressive engine, it generated up to 10-second 720p clips at 24fps from text and image inputs, quickly gaining traction among creators.
Multi-Image Support & Extension
xAI added multi-image support and video extension capabilities to Grok Imagine Video 1.0, allowing creators to chain reference images and extend generated clips for more complex storytelling workflows.
API Preview & Developer Access
Grok Imagine Video 1.0 became available via the xAI developer platform API, opening the model to third-party integrations and creative tools like Topview for broader production use.
Grok Imagine Video 1.5 Preview
xAI released Grok Imagine Video 1.5 in preview. It immediately claimed the #1 position on the Image-to-Video Arena leaderboard with a 52 Elo point jump over version 1.0, surpassing Seedance 2.0 and other competitors. Key upgrades: faster generation (~25s Fast mode), native audio improvements, and extended 15-second clip duration.
Grok Imagine Video 1.5 Generally Available
Grok Imagine Video 1.5 exits preview and becomes generally available (GA) on the xAI API. Alongside the GA launch, xAI rolled out new creative workflow features — Projects for organizing work, parallel agent execution for running multiple prompts at once, and search for finding past generations quickly.
Ecosystem & Platform Growth
Grok Imagine Video models are available through the xAI API (grok-imagine-video-1.5), grok.com/imagine web app, iOS and Android apps, and third-party creative platforms. Built on Aurora autoregressive architecture trained on 110,000 NVIDIA GB200 GPUs.
Now AvailableModel Parameters
Core Grok Imagine Video 1.5 specifications relevant to creators evaluating output quality, generation speed, and production fit.
Grok Imagine Video 1.5
xAI's latest image-to-video model, launched June 2026
#1 Image-to-Video
1404 Elo, +52 over v1.0 (May 2026)
Native: dialogue, SFX, ambience, music
Generated in single pass with video — no post-dubbing
Aurora Autoregressive (MoE)
Proprietary mixture-of-experts architecture by xAI
110,000 GB200 GPUs
One of the largest GPU clusters for video AI
480p (draft) / 720p (output)
24fps cinema-standard frame rate
6s - 15s per clip
Extendable via chaining
7 (1:1, 16:9, 9:16, 4:3, 3:4, 3:2, 2:3)
Full platform coverage from square to vertical
Image + Text
Upload a reference image with natural language prompt
H.264 MP4
Input accepts JPG, PNG, WEBP, GIF, AVIF
~25s (6s 720p Fast)
~2× faster than v1.0; ~4-5× faster than competitors
grok-imagine-video-1.5
Generally Available via xAI API
What's New in Grok Imagine Video 1.5 — v1.0 vs v1.5 Comparison
Grok Imagine Video 1.5 is xAI's latest image-to-video model, built on the upgraded Aurora autoregressive engine. It delivers faster speeds, better motion physics, improved audio sync, and longer clip durations compared to version 1.0.
| Capability | Grok Imagine Video 1.0 | Grok Imagine Video 1.5 | Improvement |
|---|---|---|---|
| Max Resolution | 720p | 720p | Better detail, less warping |
| Generation Speed (6s 720p) | ~40+ seconds | ~25 seconds (Fast) | Nearly 2x faster |
| Max Duration | 10s | 15s | 50% longer clips |
| Native Audio | Basic sync | Clearer dialogue, better lip-sync, event-aligned SFX | More polished audio-visual coherence |
| Motion Physics | Some warping | Better momentum, fewer warps, believable weight | More realistic movement |
| Aspect Ratios | 5 formats | 7 formats (1:1 to 16:9 and vertical) | Full platform-native support |
| API Status | Preview | Generally Available (GA) | Production-ready API |
| Arena Leaderboard | Strong contender | #1 Image-to-Video (+52 Elo jump) | Top-ranked model |
| Platform Features | Basic generation | Projects, parallel agents, search library | New creative workflow tools |
| Training Compute | Previous cluster | 110,000 GB200 GPUs | Massive infrastructure scale |
Grok Imagine Video 1.5 vs Seedance 2 vs Veo 4 vs Sora 2 - Model Comparison
Choosing the right AI video model in 2026 means comparing output quality, speed, and workflow fit. This comparison focuses on the features that matter most for creators, marketers, and production teams.
| Feature | Grok Imagine Video 1.5 | Seedance 2 | Veo 4 | Sora 2 |
|---|---|---|---|---|
| Developer | xAI | ByteDance | OpenAI | |
| Max Duration | 15s | 15s | 20s+ | 12s |
| Max Resolution | 720p | 1080p | 4K | 1080p |
| Native Audio | Dialogue + SFX + ambience (single-pass) | Dialogue + SFX + lip-sync | Dialogue + ambience mix | Generated audio |
| Input Type | Image + Text | Image + Text + Multi-ref | Image + Text | Image + Text |
| Architecture | Aurora autoregressive (MoE) | Dual Branch Diffusion Transformer | Diffusion Transformer | Diffusion Transformer |
| Generation Speed | ~25s (6s 720p Fast) | ~2 min | ~2.5 min | ~3 min |
| Multi-Shot Sequences | Via chaining | Yes | Yes | Yes |
| Arena Ranking (I2V) | #1 (May 2026) | #2 | Top 5 | Top 5 |
| API Available | GA (grok-imagine-video-1.5) | Full | Full | Limited |
| Best For | Fast I2V with native audio, rapid iteration | Reference depth and multi-shot storytelling | Cinematic polish and 4K output | Physics realism and text-to-video |
Grok Imagine Video 1.5 stands out as the fastest image-to-video model with native audio — generating high-quality 720p clips in about 25 seconds, roughly 4-5× faster than competitors. It ranked #1 on the Image-to-Video Arena leaderboard as of May 2026. For creators prioritizing speed, native audio-visual coherence, and production efficiency, Grok Imagine Video 1.5 is the clear frontrunner.
Who Should Use Grok Imagine Video 1.5 on Topview
Grok Imagine Video 1.5 is built for teams that need fast image-to-video generation with native audio — from cinematic storytellers and product marketing teams to social content creators.
Filmmakers and Story-First Creators
When you need cinematic framing, camera language, and scene composition from a reference image, Grok Imagine Video 1.5's Aurora engine delivers coherent motion and native audio in about 25 seconds — fast enough for creative exploration.
Fashion, Beauty, and Product Teams
Start from a product photo and generate polished product showcase videos. Grok Imagine Video 1.5 excels at maintaining product detail and lighting mood from the reference image with realistic motion and ambiance.
Performance Marketers and Ad Teams
Grok Imagine Video 1.5's ~25-second Fast mode makes it ideal for ad variant testing. Generate multiple hooks, angles, and versions rapidly — compare performance and scale what works without slowing down your creative pipeline.
Music and Dance Creators
Native audio-visual sync means beat-aware motion and rhythm-driven visuals. Generate performance clips that match music energy without external audio alignment work — all in a single generation pass.
Viral Social and Trend Creators
Grok Imagine Video 1.5's speed makes it perfect for social-first creators who need trending hooks, pet videos, and POV concepts at the pace of platform culture. 720p is the sweet spot for social platforms.
Creative Teams That Value Speed
If your bottleneck is generation speed, Grok Imagine Video 1.5's 25-second Fast mode is a significant advantage. More iterations, more variants, more chances to find the creative that performs.
How to Use Grok Imagine Video 1.5

Upload a reference image and write a prompt
Start with your key visual — a product photo, character design, or scene reference. Describe the motion, camera movement, and audio atmosphere you want.

Generate Video
Click generate and watch Grok Imagine Video 1.5 create a 720p/24fps video with native audio in about 25 seconds (Fast mode).

Download the video
Export a clean MP4 with synchronized audio when you're ready to publish to any platform.
Experience Grok Imagine Video 1.5 — The #1 Image-to-Video AI
No expensive GPUs required. Generate cinema-grade 720p video with native audio from text prompts and reference images — all in about 25 seconds with Grok Imagine Video 1.5 on Topview.
Start free · No credit card required · All leading AI video models in one workspace
