Pegasus Journey of the Forgotten
A metaphorical prompt depicting forgotten impoverished people transformed into royalty, riding pegasi to a castle.
1,523 prompts
A metaphorical prompt depicting forgotten impoverished people transformed into royalty, riding pegasi to a castle.
A cinematic visual sequence showing the crescent moon and Mars appearing close together in the sky at dawn.
A conceptual video prompt depicting photons transforming negative emotions like fear and pain into love and happiness.
A highly detailed cinematic video prompt for a 15-second sequence of a woman discovering a flickering warm light in a rainforest.
A peaceful prompt depicting a full moon hanging over rolling hills and a shimmering lake.
A cinematic video prompt depicting a subject walking forward as a Starship rocket launches behind a mountain range in the background.
A cinematic and melancholic video prompt depicting a little girl sitting inside a giant rusty pocket watch in the desert.
A cinematic wide-angle shot of a massive, rusty industrial pipe spanning a valley with a lone figure walking atop it.
A high-resolution sci-fi prompt for generating a disc-shaped mothership with volumetric lighting effects.
A prompt for creating a 10-second, 9:16 vertical cinematic 3D cartoon video.
A highly detailed multi-shot video prompt with specific timestamps for creating a samurai duel scene in the rain.
A realistic portrait video of a woman in a black lace top standing by a window in natural side lighting.
An animation prompt for Grok-imagine designed to animate a puppy photo, showing it raising a paw as if hailing a taxi.
A detailed cinematic video prompt depicting a victorious gladiator in glowing cyborg armor captured in a dramatic orbital shot.
A detailed animation prompt of a tabby cat in an eggshell, featuring realistic meowing and dust particles in sunlight.
A cinematic and atmospheric photograph depicting a woman sitting on a semi-submerged plane wreck in a stormy ocean.
A cinematic and melancholic video prompt featuring an old woman on an abandoned train platform.
A cinematic video prompt for a miniature world rescue mission, featuring detailed environmental elements and realistic action.
A cinematic video prompt for a giant transforming robot, featuring detailed mechanical parts, sparks, and slow-motion effects.
A detailed cinematic prompt for Grok Imagine featuring a soot-covered female miner in a dark tunnel, focusing on the dramatic lighting as her eyes turn into burning amber.
A poetic video prompt capturing a woman taking a deep breath among roses under warm sunlight and falling petals.
A detailed video generation prompt visualizing a future Mars city built by Tesla's Optimus robots under Elon Musk's vision.
A cinematic video prompt describing a man walking through a lush greenhouse filled with glowing plants at night.
An advanced video generation prompt designed to transform static images into 15-second suspenseful cinematic sequences with precise timing and character movement.
Last reviewed August 17, 2026 · Editorial synthesis of current Topview samples and public prompting guidance
Quick answer
A useful Grok Imagine video prompt can be concise: identify the subject, give it one readable action, and say how the camera observes that action. Add the environment, light, visual treatment, dialogue, sound, or constraints only when each detail resolves a production decision. For an image-led workflow, treat the source frame as the visual anchor and describe what should move, what should stay recognizable, and where the action ends. For a more complex clip, arrange two or three short beats in chronological order rather than combining unrelated events in one sentence. Natural language is enough; concrete verbs such as turns, catches, tracks, or pulls back are more directable than labels such as epic or dynamic. If the selected Topview mode provides audio or dialogue generation, name the speaker, exact line, delivery, ambience, and sound cue separately. Finish with a small set of observable realism rules, then review the output for subject drift, motion, framing, physics, speech, and the final frame. Prompt wording can improve clarity, but it cannot guarantee identity, lip sync, sound, or physically perfect motion in every generation.
[output context] + [subject or source-frame anchor] + [primary action or short beats] + [shot and camera] + [environment, light, and style] + [dialogue and sound when available] + [continuity and realism constraints]
Name the intended format, orientation, and clip purpose when they affect composition or pacing. Keep resolution and duration in product settings when the selected workflow exposes those controls.
Define the person, product, animal, vehicle, or environment viewers should follow. In an image-led workflow, say which visible details should remain recognizable without redescribing every pixel.
Use a concrete motion verb and a clear direction, speed, or end state. For a sequence, arrange a small number of actions in the order they should happen.
Choose a shot size, viewpoint, and one motivated camera move. Explain what the move reveals instead of stacking cinematic terms.
Add weather, atmosphere, light direction, palette, and one coherent capture language that supports the action.
When audio is available in the selected Topview mode, separate exact speech from ambience, effects, and music, then connect important sounds to visible events.
Name the few details, screen directions, object states, or physical outcomes that would visibly break the idea if they changed.
Start with the smallest instruction that communicates the idea, then add control only where ambiguity could change the result. Most clips do not need every layer.
| Goal | Prompt pattern | Example wording |
|---|---|---|
| Capture the core idea | [subject] + [one visible action] | A red fox trots across a snow-covered country road. |
| Clarify motion | [subject] + [action direction, pace, and end state] | A red fox trots from frame left to frame right, slows at the tire tracks, and stops to listen. |
| Direct the viewer | [action] + [shot size] + [one camera behavior] | Low side-tracking medium shot follows the fox at its pace, then settles when the fox stops. |
| Set a coherent world | [environment] + [light] + [one visual treatment] + [physical atmosphere] | Quiet rural road at blue hour, cold backlight on the fur, naturalistic documentary texture, loose snow moving in the wind. |
| Protect the result | [important sound if available] + [two or three observable constraints] | For an audio-enabled mode: soft paw steps and winter wind. One fox only, stable leg count, no sudden cut or camera roll. |
The current Topview samples range from one-line actions to detailed shot plans. Use the simplest pattern that still makes the intended event readable.
| Goal | Prompt pattern | Example wording |
|---|---|---|
| Animate a source image | Anchor the existing subject → describe what begins moving → define a restrained end state | Starting from the supplied portrait, she lifts her eyes toward the window, takes one quiet breath, and ends with a slight smile; keep the framing and wardrobe recognizable. |
| Stage a single action | Initial position → one action with direction and weight → clear completion | The cyclist enters from frame right, brakes on the wet pavement with believable momentum, plants one foot, and comes to a complete stop beside the kiosk. |
| Show a transformation | Stable starting form → visible transformation mechanism → finished form and hold | The paper bird unfolds along its creases into a small mechanical swallow; brass joints lock into place, the wings open once, and the finished form holds on the table. |
| Build a short narrative arc | Establish → trigger → reaction or payoff, with one main event per beat | Start wide on the empty platform. A suitcase rolls into view by itself. The waiting conductor notices it, steps back once, and the camera ends on his reaction. |
| Direct a product reveal | Product anchor → controlled material motion → readable hero state | A ribbon of condensation travels down the unchanged bottle while the turntable rotates a quarter turn; the label finishes facing camera in a clean centered hero frame. |
Select camera language according to the information the shot must reveal. One clear move is usually easier to evaluate than several simultaneous moves.
| Goal | Prompt pattern | Example wording |
|---|---|---|
| Reveal emotion | Stable medium shot → slow push-in → stop before an intimate close-up | Medium shot at eye level; slow push-in as the runner hears the announcement, ending on her restrained reaction. |
| Follow movement | Side or rear tracking shot at subject speed + consistent screen direction | Waist-height side-tracking shot follows the skateboarder moving left to right; keep the face near the upper-left third. |
| Show form or scale | Close detail or low angle → controlled orbit or pullback → wider context | Begin on the rover wheel pressing into red dust, then pull back slowly to reveal the vehicle alone beneath the canyon wall. |
| Create grounded realism | Locked or lightly handheld camera + specific imperfection + movement limit | Locked street-level camera with mild autofocus recovery as the bus crosses foreground; no zoom and no camera shake after focus settles. |
| Protect spatial continuity | Declare entrance, travel direction, eyeline, and final screen position | The mug enters from the left, is caught once at center frame, then exits fully to the right; the hand ends empty. |
Use physical details that can be seen or heard. Audio instructions are relevant only when the selected mode supports them, and every output still needs review.
| Goal | Prompt pattern | Example wording |
|---|---|---|
| Make atmosphere visible | [weather or particles] + [how they react to subject and light] | Fine rain streaks through the storefront light, splashes under each step, and beads on the jacket without becoming fog. |
| Keep lighting coherent | [key light source] + [direction] + [surface response] + [continuity rule] | Warm window light remains camera-left, creating one stable rim on the glass and soft reflections across the metal cap. |
| Write short dialogue | [named speaker] says "[exact line]" in [delivery]; [listener or camera behavior] | In an audio-enabled mode, Nia says, "Leave the light on," quietly and without smiling; the listener remains silent in a locked two-shot. |
| Synchronize sound and action | [visible event] lands with [specific effect]; [ambience or music rule] | The station sign flickers out with one electrical pop; distant train ambience continues, with no music. |
| Request believable physics | [weight, contact, momentum, or material behavior] + [observable failure exclusions] | The heavy crate compresses the wet soil on landing, slides only a few centimeters, and stops; no bounce, floating, duplicated crate, or changing dimensions. |
Animate the bicycle courier in a cinematic and realistic way with dramatic camera movement, great sound, and lots of action.
Short 9:16 image-led street scene using the uploaded frame as the visual anchor. Keep the same bicycle courier, yellow rain jacket, black helmet, cargo bag, and red bicycle recognizable. One continuous street-level shot. Start in a locked medium-wide frame as the courier looks over the left shoulder. The bicycle then rolls forward from right to left at a controlled pace while the camera begins a smooth side track. A delivery receipt slips from the cargo bag; the courier brakes once, plants the left foot on the wet pavement, reaches down, and picks it up. End with the bicycle fully stopped and the receipt visible in the courier’s empty right hand. Overcast afternoon, soft storefront reflections, light rain, natural tire spray, believable weight and braking momentum. For an audio-enabled mode: quiet traffic, rain on fabric, one brake squeak, no music or dialogue. Preserve screen direction and bicycle proportions. No cuts, camera roll, extra rider, duplicated wheels, floating receipt, or wardrobe change.
The weak prompt names a mood but does not identify a readable event, camera path, sound cue, or finish state. The stronger version turns the idea into one reviewable action arc, gives the source frame a bounded role, preserves one screen direction, uses a single motivated tracking move, describes contact and momentum, and defines which sounds apply only in an audio-enabled mode. Its constraints target failures that would break this particular shot instead of promising perfect realism or adding a generic negative list.
Add one concrete action, its direction or pace, and a clear end state. In an image-led workflow, focus on what changes after the starting frame.
Choose an observable shot size and move—such as a slow push-in, locked low angle, or side track—and say what it should reveal.
Keep one primary action arc. If the idea needs multiple locations or major actions, divide it into separate generations or a small number of chronological beats.
Use one main camera behavior per beat. Reserve a pan, pullback, or angle change for the moment it adds new information.
State weight, contact, momentum, material response, or environmental interaction, then review the output for the exact failure you care about.
First confirm the selected Topview mode supports the needed audio input or generation. Keep lines short, name the speaker and delivery, and verify synchronization after generation.
List the few visible invariants that matter—face, clothing, product geometry, prop count, light direction, or screen direction—without implying a guarantee.
Describe the intended action positively, then exclude only a few scene-specific failures such as an extra object, an unwanted cut, or a changed product shape.
A good prompt identifies the subject, one readable action, how the camera observes it, and the intended environment or finish. Add sound and constraints only when they matter to the chosen mode and concept. Clear production decisions are more useful than a long string of quality adjectives.
Yes. Many current Topview examples use only a subject and one action. A concise prompt is appropriate when the idea is simple or a source image already establishes the composition. Add detail when you need to control direction, timing, camera, atmosphere, sound, continuity, or the final state.
There is no useful universal word count. Use one or two sentences for a simple movement and a short ordered brief for a more complex action. Remove repeated adjectives, contradictory styles, and instructions that do not change what should move, appear, sound, or remain stable.
Text-to-video needs enough description to establish the subject and scene. Image-to-video can use the starting frame as a visual anchor, so the prompt should concentrate on motion, camera, atmosphere, sound, and the details you want to remain recognizable. Available inputs depend on the Topview mode you select.
Use timestamps or numbered beats when order and pacing would otherwise be ambiguous. For one continuous action, a start, progression, and end state is usually clearer. Multiple cuts increase continuity demands, so split a larger idea into separate clips when each scene needs its own setup.
Begin with familiar instructions such as locked shot, close-up, wide shot, slow push-in, pullback, pan, orbit, side track, handheld follow, POV, or low angle. Choose one main move and connect it to what the audience should notice.
When the selected Topview mode supports audio, write short exact dialogue, identify the speaker and delivery, and separate speech from ambience, effects, and music. Attach important sound cues to visible actions, and review lip sync and timing rather than assuming they will be exact.
Describe observable physics: where weight lands, what touches the ground, how momentum is absorbed, how fabric or particles react, and what state the subject ends in. Keep the action load manageable and evaluate contact, proportions, and object count after generation.
Use a clear source-frame or subject anchor, repeat only the most important visual invariants, limit unnecessary cuts and style changes, and keep product geometry or wardrobe descriptions stable. These steps can improve direction but do not guarantee identical details across every frame.
Start with positive instructions for the intended action, framing, materials, and end state. Then add a few exclusions tied to visible risks in that scene. Whether a separate negative-prompt field is available depends on the selected Topview workflow.
How this guide was built
This guide is an editorial synthesis of a snapshot of 1,188 Grok Imagine records in the Topview prompt library reviewed on August 17, 2026; every record in that snapshot was classified as video generation. About seven in ten non-empty prompts used 40 words or fewer, while longer examples added camera or shot direction, ordered beats, dialogue or sound, preservation language, and scene-specific realism constraints when the concept required more control. The structure was also checked against official xAI material describing image-led motion, camera, pacing, atmosphere, physics, sound design, and current Video 1.5 audio features, plus public Grok Imagine guides covering prompt length, motion verbs, camera vocabulary, use-case examples, mistakes, and FAQs. All formulas and examples above were written from scratch for Topview. The official 1.5 material was used to validate general vocabulary, not to claim that every Topview Grok Imagine mode exposes the same inputs or behavior. These recommendations are starting points, not a guarantee that any prompt has been independently tested or will reproduce the same identity, dialogue, sound, or motion on every run.