xAI adds image, text, and voice references to Grok Imagine Video 1.5 with native 1080p output
xAI has expanded Grok Imagine Video 1.5 with reference-based video generation, letting users lock a specific face, product, or location into a generated clip, alongside a jump to native 1080p output and full text-to-video generation with no starting image required.
What's new
xAI describes the update simply: "Our best video model, now with text, image, and voice references — generating up to 1080p." The release adds three input modes on top of the model's existing image-to-video capability:
- Text-to-video — users can "describe the shot — no starting image needed," generating a clip directly from a text prompt.
- Reference-to-video — up to seven reference elements can be locked into a scene at once. As xAI puts it, "Each reference image locks one thing in place — a face, a product, a location," letting the rest of the scene vary while that element stays consistent across shots.
- Native 1080p — output resolution for both text-to-video and image-to-video generation moves up from the model's previous cap.
xAI's developer documentation confirms the same update on the API side: grok-imagine-video-1.5 "now supports text-to-video, image-to-video, and reference-to-video (including optional preset voices), with native 1080p for T2V and I2V," available through the Video Generation, Image-to-Video, and Reference-to-Video endpoints.
Context
Grok Imagine Video launched as xAI's image-to-video and text-to-video product earlier in 2026 and currently ranks first on the Image-to-Video Arena leaderboard. This update is a capability expansion on that existing model rather than a new model release — it broadens the input modes (adding text and reference inputs to the original image-to-video flow) and raises the resolution ceiling, rather than replacing the underlying model. It lands the same week xAI also shipped Grok Voice Think Fast 2.0 and an adjustable voice-activity-detection threshold for its Speech to Text API, part of a broader push across xAI's media and audio stack in late July.
Why it matters
Reference-locking is the feature that determines whether an AI video tool is usable for anything beyond one-off clips — ads, series, or any content needing the same character or product to reappear consistently across shots. By adding multi-reference support (up to seven locked elements) alongside a resolution bump to 1080p, xAI is directly targeting the consistency and quality gaps that have limited AI-generated video's use in professional production work, an area where competitors like Runway, Luma, and Google's Veo have also been racing to close the gap between demo-quality and broadcast-quality output.
Corroborating sources
- X
https://x.ai/news/grok-imagine-video-1-5-references
“Our best video model, now with text, image, and voice references — generating up to 1080p.”