Ludo.ai

You've got questions? We've got answers!

Explore our comprehensive documentation for in-depth information about Ludo.ai and its powerful features.
Generating Game AssetsAPI & MCPAudio GeneratorSprite GeneratorVideo Generator3D Asset GeneratorFAQUser Account and SubscriptionProjectImage GeneratorGame IdeatorAsk Ludo

Video Generator


  • Introduction and Getting Started

    The Video Generator is a powerful AI tool within the Ludo.ai platform designed to create short, animated game videos based on images and text descriptions. This feature allows you to quickly visualize motion, animate characters and assets, or create dynamic gameplay mockups.

    With the Video Generator, you can:

    • Generate videos from a detailed text prompt (Text to Video).
    • Generate videos from a starting image and a motion prompt (Image to Video).
    • Set a custom final frame for precise animation control.
    • Generate a video of your characters and objects from up to five reference images (From References).
    • Edit an existing video with a text instruction and optional reference images (Edit Video).
    • Get synchronized sound with every video — the Griffin models render sound effects, dialogue and music as part of the clip, driven by your prompt.
    • Pick clip lengths from 5 to 15 seconds and pack several actions, camera moves and lines of dialogue into one clip.
    • Choose between Griffin and the higher-resolution Griffin HD.
    • Upscale a finished video 2x.
    • Utilize a two-step generation process for maximum control.
    • Apply filters for art style, perspective, aspect ratio, and more.
    • Edit the initial or final frame of a video using the integrated Image Editor.
    • Download, save, and iterate on your video creations.

    To get started:

    1. Navigate to the Video Generator tool from the main menu.
    2. Choose your desired generation mode: Text to Video, Image to Video, From References or Edit Video.
    3. Provide your inputs (text prompt, an image, reference images, or a video to edit).
    4. Use filters and pick a model and duration to refine your request.
    5. Generate the first frame or the final video.

  • Generation Modes Explained

    The Video Generator offers four modes, as tabs at the top of the page.

    Text to Video

    This mode is ideal when you are starting from scratch with just an idea. You provide a single, detailed text prompt that describes both the visual content of the video's first frame and the motion you want to see. The tool then generates an initial static image based on your description, which you can then approve to generate the final animated video.

    Image to Video

    Use this mode when you already have a specific image you want to animate. This could be an image you've previously generated on Ludo.ai, or one you upload. You provide the starting image, an optional final frame, and a separate text prompt to describe the desired motion. The tool will then generate a video that animates your provided image.

    From References

    Use this mode when you have the cast but not the shot: up to five reference images of characters, objects or scenes, plus a prompt describing the video you want them in. Unlike Image to Video, the references are not the first frame — the model composes a new shot with those subjects in it. See From References.

    Edit Video

    Use this mode to change an existing video with a text instruction — swap an item, recolor, add an effect, restyle — while keeping its motion and timing. You can add reference images to show exactly what you mean. See Editing a Video.


  • Text to Video: The Two-Step Process

    The "Text to Video" mode uses a two-step workflow so you can review and refine the starting visuals before committing credits to a full video render.

    Step 1: Generate the First Frame

    1. Select the Text to Video tab.
    2. In the prompt box, write a detailed description. Your prompt should describe both the subject of the video AND the motion you want to happen. For example: "In a pixelated world, a retro spaceship narrowly evades incoming asteroid clusters while its cannons unleash a barrage of retro laser fire."
    3. Pick a Video Type (Art, Gameplay, Icon, Asset, Sprite) and, optionally, set filters (Art Style, Perspective, Aspect Ratio) to further guide the AI.
    4. Click the "Generate First Frame" button.
    5. Ludo.ai will generate one or more static images that represent the starting point of your video.
    6. If you are not happy with the result, generate again with the same inputs or adjust the prompt and filters.

    Step 2: Turn the Frame into a Video

    1. Review the generated first frame(s). You can click on an image to zoom in.
    2. On the frame that best matches your vision, click "Use as First Frame" in the action bar.
    3. Ludo.ai switches you to the Image to Video tab with that frame pre-loaded as the first frame and your original prompt copied over as the motion prompt.
    4. In the Image to Video tab you can:
      • Optionally set a Final Frame for precise control over how the animation ends.
      • Edit the motion prompt to better describe what should happen.
      • Choose a Model and a Duration (see Choosing a Model below).
    5. Click "Generate Video" to render the final animated video. It will appear in the gallery below once ready.

  • Image to Video

    The "Image to Video" mode is a direct way to animate an existing visual. It is also where you land after clicking "Use as First Frame" on a frame generated in the Text to Video flow.

    1. Select the Image to Video tab.
    2. Click "Choose initial image" on the First Frame card. You can drag and drop a file or browse your computer.
    3. (Optional) Set a Final Frame. Click the "Choose final image" card. You can upload an image, copy the first frame, or open the editor to derive a final frame from your starting image. The final frame is guaranteed to be the last frame of the video, giving you precise control over the animation's end point — ideal for perfect loops or specific scene transitions.
    4. In the prompt box, describe what should happen in the video. This prompt should focus on what happens — actions in order, camera, dialogue and sound — since the initial (and optional final) images already define the visuals. You can leave it blank to let the AI decide on the motion. See Writing Effective Prompts.
    5. Pick a Model and a Duration — Griffin is the default; Griffin HD renders the same clip at a higher resolution. Durations run from 5 to 15 seconds in whole seconds, and every Griffin video comes with an audio track. See Choosing a Model below.
    6. Click the "Generate Video" button. The final video will appear in the gallery once it is ready.

    Editing Your Keyframes for Better Control

    Before generating the video, you can precisely control the starting and ending visuals using the integrated Image Editor. After uploading or selecting an image for the First Frame or Final Frame, click the "Open In Editor" button on its card. This allows you to:

    • Add, remove, or change specific elements in the image.
    • Modify colors, styles, or backgrounds.
    • Refine the composition to better suit the animation you have in mind.

    Editing your keyframes ensures the animation begins and ends with the exact visuals you want. Once you save your edits, you will return to the Video Generator to describe the motion and create the video.

    Input Image Guidelines and Limitations

    To ensure the best results and adhere to platform policies, please follow these guidelines for your input images:

    • No Photos of Real People: The generator is not designed to animate photographs of identifiable individuals.
    • No Not-Safe-For-Work (NSFW) Content: All uploaded and generated content must be suitable for a general audience.
    • Content moderation applies to the image and the prompt. A blocked request fails with a moderation message; to get past it you have to change the image, the prompt, or both. See Troubleshooting.
    • Aspect ratio follows the image. The video keeps the aspect ratio of your first frame.

  • Writing Effective Prompts

    The quality of your prompt is key to getting great results. What a good prompt looks like depends on the tab: in Text to Video the prompt paints the first frame, in the other tabs it directs the clip.

    Text to Video: describe the frame and the motion

    Here the prompt is used first to generate a still image, so it needs the visuals as well as the action. A good structure includes:

    • Subject: The main focus (e.g., a hero sprite).
    • Setting: The environment (e.g., in a top-down pixel art perspective).
    • Action/Motion: What the subject and camera are doing (e.g., executes a rapid dash attack through a cluster of blocky enemies, triggering bright particle effects).
    • Details: Specifics about the scene (e.g., a health bar depletes).

    Image to Video, From References and Edit Video: direct the clip

    The image already fixes the look, so don't re-describe it. Write what happens, and use the whole duration. Ludo turns your instruction into a paced, shot-by-shot description for the video model, so a longer clip can carry a real sequence rather than one motion stretched thin:

    • Several actions, in order. List what happens as a sequence: "the knight draws her sword, blocks the falling boulder, then turns and sprints for the gate as it starts to close." On a 10 or 15-second clip, two to four beats — an opening state, a development, a turn, a result — each get real screen time. On a 5-second clip keep it to two or three actions that flow into each other.
    • Camera moves. Name one camera idea per shot, as a plain sentence: "the camera slowly pushes in on his face", "a fast pan follows the arrow", "static shot". Say how big and how fast when it matters.
    • Dialogue. Put spoken lines in double quotes, and say who speaks and how: the shopkeeper leans in and whispers: "Not for sale." Lines are reproduced verbatim and lip-synced; nothing is invented that you didn't write.
    • Sound. Describe the sounds and music you want — footsteps on gravel, a metallic clang on impact, tense low strings — and they are rendered with the picture. Music is only added when you ask for it.
    • What stays the same. If something must not change (a logo, a colour scheme, on-screen text), say so. Text that exists in the image is kept verbatim.
    • Cuts. You can ask for a cut to a new viewpoint on longer clips; otherwise the clip is one continuous shot. With a final frame set, the clip is always one continuous shot, so it can interpolate to that frame.

    What to avoid: mood words and adjectives ("epic", "beautiful", "cinematic") do nothing; concrete nouns and verbs do the work. Don't write the duration or timestamps yourself — the pacing is derived from the duration you selected.

    Sprites and game assets. The asset's type (a sprite, a background, a VFX element, an icon) tells Ludo what a typical animation for it looks like — an idle loop for a background, a static camera for a sprite — and that is used as a default. Your instruction always wins: ask a background to explode and it explodes.

    Simple prompts for exploration

    Simple prompts are excellent for brainstorming and discovering unexpected ideas. Don't be afraid to start with a broad concept and let the AI fill in the details.

    • Simple Prompt Example: "a spaceship shooting lasers"
    • AI Interpretation: The AI might generate a side-scrolling view, a top-down view, a cinematic fly-by, or something else entirely. Each result can be a new source of inspiration.

    Use detailed prompts when you have a specific vision in mind, and use simple prompts when you want the AI to be a creative partner in the ideation process.


  • Filters and Options

    The Video Generator uses a focused set of options to steer the AI. The options available depend on which tab you are in.

    Text to Video options

    These options shape the first frame generated from your prompt:

    • Video Type: Choose the category of the starting image (Art, Gameplay, Icon, Asset, Sprite).
    • Art Style: Choose from styles like Pixel Art, Cartoonish, Photorealistic 3D, and more.
    • Perspective: Define the camera viewpoint, such as Top-Down, First-Person, or Side-Scroll.
    • Aspect Ratio: Control the dimensions of the generated first frame (e.g., Square, Landscape, Portrait). The video inherits the aspect ratio of the frame.

    Image to Video options

    In Image to Video, the first frame is already chosen, so the options focus on the animation itself:

    • First Frame and (optional) Final Frame: The keyframes that bound the animation.
    • Model: Griffin or Griffin HD. See Choosing a Model.
    • Duration: Video length, in whole seconds from 5 to 15. The credit cost depends on the model and the duration.

    From References options

    • Reference Images: one to five images of the subjects the video should contain.
    • Text Description: the video to generate from those references.
    • Model: Griffin or Griffin HD.
    • Duration: 5 to 15 seconds.
    • Aspect Ratio: the output shape — Default (follows the first reference), Landscape 16:9, Portrait 9:16, Square 1:1, Standard 4:3, Portrait 3:4 or Cinematic 21:9.

    Edit Video options

    • Video to edit: any video from your history.
    • Edit Instruction: what should change.
    • Reference Images (optional): up to five images showing the object, character or style to bring in.
    • Model: Griffin HD.
    • Max Duration: an upper limit for the edited video. The result follows the source video's length and never exceeds this value.

  • Interacting with Generated Videos

    Once a video is generated, it appears in the gallery with several options:

    • Play/Zoom: Click on a video thumbnail to open it in a larger player window.
    • Mute / Unmute: Every Griffin video has a soundtrack, so a speaker icon appears on the thumbnail. Clicking it toggles all videos between muted and unmuted, so previews don't talk over each other. The sound is part of the generation — to change it, change the prompt and generate again (see Writing Effective Prompts).
    • Action Bar (on hover):
      • Edit Video: Opens the Edit Video tab with this video as the source. See Editing a Video.
      • Upscale 2x: Renders a new copy of the video at twice the resolution, keeping its frame rate, duration and audio. Not available for videos that are already 960×960 or larger. The result lands in the gallery as a new video.
      • Generate from Image: Takes the first frame of the selected video and brings it into the Image to Video tab. This is perfect for trying a different motion prompt while keeping the same starting visual.
    • Three-dot Menu (top-right of the thumbnail):
      • Download Video: Downloads the video file (as an .mp4) to your computer.
      • Get Help With This: Opens the documentation assistant with this video attached, so you can ask why a result came out the way it did. It also shows the video's ID, which is what our team needs if you report the result on Discord.
      • Delete: Removes the video from your gallery.

    First-frame thumbnails (Text to Video flow)

    Between Step 1 and Step 2 of the Text to Video flow, the gallery shows the generated first frames as image-only thumbnails. These have their own actions:

    • Use as First Frame: Sends the image to the Image to Video tab as the first frame, with the original motion prompt copied over. This is the bridge to Step 2 of the two-step flow.
    • Edit Image: Opens the frame in the integrated Image Editor. Save your edits to use the edited version as the basis for a new video.

  • Use Cases in Game Development

    The Video Generator is a versatile tool that can accelerate various stages of the game creation process, from initial design to marketing.

    Concept and Design

    • Visualize Mechanics: Quickly create short clips to demonstrate how a core gameplay mechanic, like a dash attack or a magic spell, would look in motion.
    • Animate Characters: Bring character sprites or concept art to life to get a feel for their movement and personality.
    • Dynamic Mood Boards: Instead of a static mood board, create a collection of short, atmospheric videos to establish the tone and feel of your game world.

    Development and Prototyping

    • Animated Placeholders: Generate simple animated assets (e.g., a spinning coin, a pulsating power-up, an opening treasure chest) to use as placeholders in early prototypes.
    • VFX Ideation: Brainstorm ideas for visual effects like explosions, smoke, or energy blasts.
    • Sound Design Brainstorming: Quickly generate placeholder sound effects or background music to match the mood of your animated concepts, helping to establish the audio-visual tone early on.
    • UI Animation Mockups: Create animated UI elements, such as a glowing button, a loading bar, or a level-up notification, to test in your UI mockups.

    Marketing and Social Media

    • Eye-Catching Social Media Clips: Generate short, looping videos perfect for sharing on platforms like Twitter, TikTok, or Instagram to build community interest.
    • Ad Creative Mockups: Rapidly produce multiple variations of a short video to test different concepts for marketing campaigns.
    • Conceptual Trailers: Stitch together several generated clips to create a simple "vision" trailer that communicates your game's concept to potential team members or investors.

  • Troubleshooting

    If you encounter issues while using the Video Generator, consider these steps:

    • Irrelevant or Static Videos:
      • Ensure your prompt includes a clear description of what happens. If no action is described, the resulting video may be static.
      • Rephrase your prompt as a sequence of concrete actions — see Writing Effective Prompts.
    • The clip feels rushed or crammed:
      • You asked for more than the duration can hold. Raise the duration (up to 15 seconds) or cut the prompt down to the beats that matter.
    • The clip is longer than expected:
      • Griffin renders whole seconds from 5 to 15. A shorter request comes back as a 5-second video.
    • Initial Frame is Not What You Expected:
      • Refine your initial text prompt with more detail.
      • Use the filters (Art Style, Perspective, etc.) to better guide the AI.
      • Click "Generate First Frame" again to get new variations.
    • The audio doesn't match:
      • The soundtrack is rendered from your prompt together with the picture, so there is no separate audio pass. Name the exact sounds, music and lines you want in the prompt (e.g. "footsteps on gravel," "futuristic synthwave music," "sword clashing sound effects") and generate again.
    • Upscale 2x is unavailable:
      • The source is already 960×960 or larger; there is no 2x step above that.
    • Video Generation is Slow:
      • Video generation is computationally intensive and can take longer than image generation, and longer clips take longer.
      • If generation seems stuck, please wait a few moments. If the problem persists, try refreshing the page.
    • The request was blocked by content moderation:
      • The video models apply their own content moderation to every input image, reference and prompt, and a blocked request fails with a moderation message. Ludo cannot override it, and retrying the same inputs gives the same result.
      • Change what tripped it: reword the prompt (violence, weapons, gore, injuries and suggestive content are the usual causes, even when they are ordinary for a game), and/or swap or edit the image or references. Blood, wounds and realistic weapons in the image are blocked as often as in the prompt.
      • Photos of real, identifiable people are blocked outright — use generated or illustrated characters instead.
    • Video Fails to Generate:
      • This can sometimes happen if the requested motion is too complex or conflicts with the initial image.
      • Try simplifying your prompt.
      • Try generating a video from a different first frame.
    • Undesired Motion:
      • Use the "Generate from Image" option on the video to try again with a refined or different motion prompt.

    If problems persist, please contact Ludo.ai support or ask on our Discord server for assistance.


  • Choosing a Model

    The Image to Video, From References and Edit Video tabs let you pick the video model. Since September 2026 the lineup is two models from the same family:

    • Griffin — the default for Image to Video and From References. It renders at 480p, always with a synchronized audio track, and takes durations from 5 to 15 seconds. It is the right choice for almost everything: it follows multi-action prompts, camera directions and dialogue, and it is the cheaper of the two.
    • Griffin HD — the same model rendered at a higher resolution (768p), at a higher per-second cost. Pick it when the video is the deliverable rather than a mockup — trailers, store pages, social clips — or when you'll be cropping into the frame. It is the model used by Edit Video.

    Both follow the aspect ratio of your first frame (Image to Video) or the Aspect Ratio you choose (From References).

    The Model dropdown

    Open the Model dropdown to compare the models side by side. Each card shows the model's name, the duration range it supports (for example, 5 to 15s), a short description of what it is good at, and the cost in credits per second, plus a minimum charge where applicable.

    Legacy models

    Older models — Blitz, Anvil, Eagle and Eagle with Audio — have been retired to a collapsed Legacy models group at the bottom of the dropdown. They exist only so that projects already built on them can keep generating matching videos, and should not be used for new work. New accounts don't see the legacy group at all — if the dropdown only lists Griffin and Griffin HD, nothing is missing.

    Duration

    After you pick a model, the Duration control shows only the durations that model supports. Griffin and Griffin HD take whole seconds from 5 to 15, so the slider snaps to those.

    The credit cost shown next to the duration updates live based on your model + duration choice, so you always know what a generation will cost before clicking Generate Video. If you switch models and your current duration is not supported by the new model, Ludo.ai automatically selects the closest valid duration.


  • From References

    The From References tab generates a new video featuring the subjects in your reference images, rather than animating one image. Give it your character sheet, an enemy design and a background, describe the shot, and you get a clip with all of them in it.

    When to use it

    • You have a character (or several) and want to see them in a scene that doesn't exist yet as a single image.
    • You want a whole cast to stay on-model across many clips — reuse the same references each time.
    • You want to control the output shape independently of any input image.

    How to Use:

    1. Select the "From References" tab.
    2. Add Reference Images — one to five. Each image should show one subject clearly: a character, an object, a location. Ludo-generated images and uploads both work.
    3. Write the Text Description. Describe the video, and refer to each reference by its position: "Image 1 is the knight, Image 2 is the dragon. The knight walks into the cave from the left, the dragon lifts its head and roars, smoke curls up as the knight raises her shield." Ludo tells the model what each image defines, so a short role line per reference at the start of the prompt keeps identities straight. Everything from Writing Effective Prompts applies: actions in order, camera, dialogue in quotes, sound.
    4. Choose a Model, Duration and Aspect Ratio. Griffin is the default; Griffin HD for a higher-resolution render. Duration is 5 to 15 seconds. Aspect Ratio sets the output shape — by default it follows the first reference; otherwise pick Landscape 16:9, Portrait 9:16, Square 1:1, Standard 4:3, Portrait 3:4 or Cinematic 21:9.
    5. Click "Generate Video." The result appears in the gallery like any other video, with the same actions (edit, upscale, download).

    Tips

    • Fewer, cleaner references beat many busy ones. The subject should fill the reference and sit on a plain background where possible.
    • Say what must not change — armour colour, a logo, a hairstyle — if it matters.
    • The references define who and what; the prompt defines where and what happens. A background reference is a good way to fix the setting.

  • Editing a Video

    The Edit Video tab changes what's in an existing video while keeping its motion, timing and framing. Instead of re-generating and hoping for the same movement, you edit the clip you already have.

    What you can do with a text instruction

    • Replace elements — "Replace the sword with a flaming axe."
    • Add or remove elements — "Add a red cape," "Remove the shield."
    • Change colours and materials — "Make the armour golden."
    • Add effects — "Add a glowing blue aura around the character."
    • Restyle the whole clip — "Make it look hand-drawn with thick outlines."

    Reference images (optional, up to 5): show the AI exactly what to bring in — your game's actual axe, a target art style, an outfit — instead of describing it.

    How to Use:

    1. Click Edit Video on any video card, or select the Edit Video tab and pick a video from your history with Choose Video.
    2. Write the Edit Instruction. Describe what should change; the motion and everything you don't mention are preserved. Keep each edit to one clear change and run a second edit for more.
    3. (Optional) Add Reference Images.
    4. Set the Max Duration. The edited video follows the source's length and never exceeds this limit — lower it to edit only the opening seconds of a long clip.
    5. Click "Edit Video." The result appears as a new video; the original is untouched, so you can iterate freely.

    Edits run on Griffin HD. Every edit re-renders the whole clip, so colours and detail drift a little with each pass — if an edit didn't work, go back to the original rather than stacking another edit on top.

    Upscale 2x

    Upscale 2x on a video card renders the same video at twice the resolution, keeping its frame rate, duration and soundtrack. It is a good final step for a clip you want to publish: edit and iterate at the normal size, then upscale the one you keep. It is unavailable for videos that are already 960×960 or larger.