Grok Imagine Video 1.5: What It Does and How to Prompt It
Grok Imagine Video 1.5 is xAI's dedicated video generation model, launched in preview on May 30, 2026, with API access opening June 3. It sits inside xAI's broader Grok Imagine suite, which also covers text-to-image and image editing, but the video model is a separate product built for one job specifically: taking a still image and turning it into motion.
The thing that makes it genuinely different from most of the models in this guide series is the image relationship. When you upload a still to Grok Imagine Video 1.5, it becomes the literal first frame of the clip, not a reference the model interprets or reinvents. The original lighting, color, texture, and composition are preserved and animated outward from there. Your prompt only drives what changes from that fixed starting point.
What is Grok Imagine Video 1.5?
Grok Imagine Video 1.5 runs on xAI's Aurora autoregressive engine. It is not the same as the Grok chatbot or the reasoning model. The brand name overlaps, but the video product is a dedicated generation tool that shares no functionality with the conversational side.
Six workflows cover everything the model does: image-to-video (its primary and strongest mode), text-to-video, reference-to-video (holding a character or visual style across clips), video editing (changing something specific in an existing clip while preserving the rest), video extension (continuing from the final frame of an existing clip), and reference-guided generation (style-locking a new generation without fixing the first frame).
Output is 1 to 15 seconds at 480p or 720p, 24 FPS, across seven aspect ratios (16:9, 9:16, 4:3, 3:4, 3:2, 2:3, and 1:1). No 1080p, no 4K. For web, social, and iteration, 720p is fine. For broadcast delivery or large-screen output, this model isn't the right call.
Native audio generates in the same pass as the video: dialogue with lip-sync, sound effects, ambient environment, and music, all from the same prompt. No separate audio step.
Where Grok Imagine Video 1.5 is strongest
Animating a still you already have. This is the model's actual purpose, and it's where it earns its position. If you have a finished photograph, a product render, a painting, or a character concept, the model brings it to life while preserving every detail in the original. The starting image is a hard constraint, not a suggestion. The lighting that was there stays. The color that was there stays. The composition that was there stays. The only thing the prompt describes is what moves, how the camera behaves, and what the scene sounds like.
Speed. Independent testing puts clip delivery at under a minute for most generations. For the drafting and approval stage of a production workflow where you need to show a client twelve options before committing, that turnaround matters more than theoretical quality differences between models.
Prompt adherence on camera work. Named camera moves execute reliably: dolly, push-in, orbit, tracking shot, handheld, pan, crane, rack focus. The model reads cinematography vocabulary and applies it to the shot. This is notably more reliable here than on some models where named camera moves produce something in the general neighborhood of the instruction.
Audio-visual sync. Like Flux 3, audio here generates alongside video rather than through a separate pass. A footstep lands on the frame a foot hits the ground. A door click lands on the frame the hand pulls it shut. For physical event-driven sound design, the results are tight.
Where it falls short
720p is the hard ceiling. The most practically significant constraint. Broadcast, large-screen delivery, or any project where source quality needs to survive post-production upscaling: this model stops at 720p. That's not a draft-mode limitation. That's the model's actual maximum.
5 to 8 seconds is the stability window. Clips longer than 8 to 10 seconds show more drift: subject softening, motion artifacts, and lighting shifts compounding through the later frames. 15 seconds is the technical maximum. It's not the reliable working range.
Audio fires on most clips, not all. Roughly 3 in 5 generations return synced sound effects. The rest return music-only. This isn't a prompting problem specifically, it's inconsistent model behavior. When sound effects matter to the output, plan for a re-roll rather than assuming first-generation audio.
Faces soften under fast motion. High-frequency facial detail is the first thing to go when the body moves quickly. A hero close-up with heavy face movement is where the model is most likely to produce something that reads as wrong. Still or slow-moving faces hold much better.
Text-to-video is the weaker mode. The model is designed around a fixed first frame. Without one, the generation is more variable and less controlled. If the brief is text-only and a specific starting composition matters, generate the first frame separately and then animate it.
How Grok Imagine Video 1.5 reads a prompt
A strong prompt is a short shot brief. Not a caption, not a description of an existing image. A brief describing what moves, how the camera behaves, and what it sounds like.
Five components, in this order: Subject and action first, then camera movement, then atmosphere and lighting, then audio, then any style or finishing notes.
Subject and action is the first sentence. Name who or what is in the frame and what happens. Concrete and observable: something a camera lens could resolve. Skip internal states and mood words unless they translate into something visible on screen.
Camera movement is named explicitly. "Cinematic" tells the model nothing. "Slow dolly push-in from a medium shot to a close-up over 6 seconds" tells it exactly what to do. One camera move per clip. Stacking two moves produces warping.
Atmosphere and lighting describe visible conditions: "warm side light from the left," "cold neon overcast," "golden hour backlight with a rim light on the shoulders." These inform the model's motion decisions as much as the visual output.
Audio goes in the same prompt as motion. Name each layer by what produces it: "the soft scrape of the chair on the floor," "rain on a corrugated metal roof," "she says: I didn't expect to see you here." Generic audio notes ("ambient sound," "realistic audio") give the model nothing specific to generate.
What you leave out, the model fills in. An unspecified camera move produces a drift or a static lock, chosen by the model. Unspecified audio produces music-only on roughly 2 in 5 generations. The prompt budget is better spent specifying than describing what's already fixed in the starting frame.
The image-first workflow
The most reliable production pattern from sources across this research: generate or select the starting frame before writing the video prompt, not after.
If the starting frame is a photograph or existing asset, the work is done. If it needs to be generated, use GPT Image 2 or Ideogram 4 first (both have dedicated guides on this blog), approve the composition, then write the video prompt around what's actually in the frame rather than what you'd like to be in it.
For talking-head content specifically: a front-facing portrait with the mouth clearly in frame, consistent neutral light, and the head occupying at least a third of the frame produces the cleanest lip-sync results. Profile-facing subjects, heavily shadowed mouths, or extreme-angle starting frames produce more variable dialogue sync.
Video extension for longer sequences
Extension continues from the final frame of an existing clip: same subject, same lighting, same motion energy. It reads the last few seconds as context and continues from there. One action per clip, then extend. Repeat for as many beats as the sequence needs.
The thing to watch is compounding drift. Each extension inherits any drift from the previous clip plus introduces its own. A 6-second clip extended twice to 18 seconds may look noticeably different at the end than it did at the start. Review across the full sequence, not just the most recent extension.
For a sequence built from extensions, keep each individual clip to the 5 to 8 second stability window rather than pushing each segment to 15. The total can still reach whatever length the sequence needs.
Five prompts across different workflows
The afternoon portrait, Image-to-video, 6 seconds
Start frame: a woman in a cream linen shirt seated by an open window, late afternoon light on one side of her face, the other in soft shadow. She's looking slightly away from camera.
She turns her head slowly toward camera, a brief pause as her eyes meet the lens, then a small exhale and the beginning of a smile. The light stays exactly as it is in the starting frame. Slow push-in from the medium shot as she turns. Her breath, the creak of the chair as she shifts, soft ambient room tone, no music.
The product pour, Image-to-video, 5 seconds
Start frame: a whiskey glass on a dark wooden surface, ice already in the glass, warm single-source light from the upper right.
A pour of amber whiskey descends in a slow arc into the glass, landing on the ice and blooming outward. Camera holds static. The glass stays exactly as it is in the starting frame. The liquid hit on ice, a low crystalline ring as it settles, soft room silence underneath. No music.
The city commuter, Reference-to-video, 8 seconds
Reference image: a woman in a tan trench coat, defined consistently.
The character from the reference image walks through the turnstile of a busy underground station at rush hour, head down, then looks up as she steps onto the platform and sees the train already open. Handheld tracking shot following from slightly behind and to the right as she navigates the crowd. Turnstile click, platform noise, crowd, train doors chiming, no music.
The wave extension chain, Image-to-video + extension, 6 + 6 seconds
Start frame: a rocky coastline, a wave mid-break, late afternoon sun catching the spray.
Clip 1:
The wave completes its break across the rocks, the spray catching the low sun and scattering into light. Camera holds static from a low angle. The full sound of the break, the water pulling back over the rocks as the foam settles.
Extension:
Continue from the final frame. A second wave rolls in behind the first, beginning to lift as the foreground foam from the previous break is still retreating. Same camera, same light. The low roll of the incoming wave building, the stones clicking as water pulls between them.
The empty stage, Image-to-video, 8 seconds
Start frame: an empty theater shot from the back of the stalls, plush red seats, gilded balconies, a single pool of stage light on the empty boards.
The house lights fade slowly over the full 8 seconds as the stage light brightens slightly in response. No movement in the frame except the light shifting. Static camera, no camera move. The murmur of an arriving audience fading to silence as the house goes dark, a single cough somewhere in the stalls, then nothing.
Common questions
What is Grok Imagine Video 1.5?
xAI's dedicated image-to-video generation model, launched May 30, 2026. It animates a still image into a video clip with native synchronized audio, treating the uploaded image as the literal first frame rather than a loose reference.
Is Grok Imagine Video 1.5 available on Aristotto?
Yes, Grok Imagine Video 1.5 is available on Aristotto.
Does Grok Imagine Video 1.5 support 4K or 1080p?
No. Grok Imagine Video 1.5 output is capped at 720p. It's well-suited to web and social delivery. For broadcast, large-screen, or post-production workflows requiring high-resolution source material, a different model is the better choice.
How does Grok Imagine Video 1.5 use the starting image?
Grok Imagine Video 1.5 uses starting image as a literal first frame of the clip. The original lighting, color, composition, and subject identity are preserved. The prompt only describes what changes
Does Grok Imagine Video 1.5 always generate audio?
Audio generates natively with the video in most cases, but not all. Roughly 3 in 5 generations return synced sound effects. The rest return music-only. Always include at least one audio cue in the prompt and plan for a re-roll when specific sound is critical.
How do I get clean lip-sync in Grok Imagine Video 1.5?
Start from a front-facing portrait with the mouth clearly in frame and the head occupying at least a third of the frame. Keep the spoken line short. Neutral consistent lighting on the starting frame produces the most reliable results.
What is Grok Imagine Video 1.5 video extension?
A workflow that continues a generated clip from its final frame: same subject, same lighting, same motion energy. Repeat to build multi-part sequences. Keep each segment within the 5 to 8 second stability window and review the full sequence for compounding drift.
What is Grok Imagine Video 1.5 reference-to-video?
A workflow where reference images guide character or style across clips without fixing the first frame. Useful for keeping a consistent character or visual look across separate generations that don't share a starting still.
How is Grok Imagine Video 1.5 different from Seedance or Wan?
Grok Imagine Video 1.5 is image-first: the uploaded still becomes the literal first frame and is preserved exactly. Seedance and Wan are generation-first: they produce motion from a prompt and use reference images as guidance rather than a fixed anchor.



