11 sept. 2026Guide

How to Generate AI Images, Videos & Lip Sync Videos with Fish Creative - Full Tutorial

How to Generate AI Images, Videos & Lip Sync Videos with Fish Creative - Full Tutorial

Generate an image from a prompt, turn it into a video, and make the person on screen speak naturally. You can do all three in Fish Creative.

Fish Creative brings your AI creative tools together in one workspace, including image generation, video generation, and lip sync. Reuse anything you have already created without downloading and uploading files between tools. Whether you are making a virtual presenter, a product demo, a social media video, or a character with a voice of their own, you can start with the same basic workflow.

This tutorial walks you through the process from start to end.

Follow along with the tutorial and start creating your own images and videos. Try Fish Creative.


What Can You Do with Fish Creative?

Sign in to Fish Audio and select Image & Video in the left sidebar to open Fish Creative. This tutorial focuses on three core tools:

  • Image: Generate images from text or upload a reference image to guide a new creation. You can also use Topaz HD to upscale individual images and recover detail.
  • Video: Create videos from text, images, first and last frames, reference images, reference videos, or a mix of inputs that includes audio.
  • Lip Sync: Add audio to an image or video to make the person on screen speak in sync with it. Your creations are saved automatically to your asset library. Once you have made an image, you can send it straight to video generation. When the video is ready, send it to Lip Sync. You can stay in the same workspace throughout.

Sign up for a Fish Audio account to start using Fish Creative. New users receive free credits to try image generation, video creation, and lip sync.


1. From an Idea to a Finished Video: The Full Workflow

Watch this section in the full tutorial above: 00:19–01:06.

You do not need to learn every model and button before you start. It helps to see how a piece of content comes together first, then work through the details in the sections below.

For this tutorial, we will make a short video featuring a virtual presenter. Here is the basic sequence:

Plan the video → Create a character image → Animate the image → Prepare the voice → Add lip sync

Follow these steps:

  1. Sign in to Fish Audio and open Image & Video from the left sidebar.
  2. Decide where the video will be used and choose an aspect ratio. For a vertical talking-head video on social media, for example, start with 9:16.
  3. Open Image and write a prompt describing the character, setting, composition, style, and lighting.
  4. Choose a model and settings, generate your images, and pick one with a clearly defined subject and a composition that will work well in motion.
  5. Click Generate Video to send the selected image directly to the video tool.
  6. Write a prompt describing the character's movements and expressions, the camera movement, and the pacing.
  7. Set the video's aspect ratio, duration, resolution, and number of outputs, then generate it.
  8. Check the face, continuity of movement, and visibility of the mouth. Choose a video that will work well for lip sync.
  9. Open Text to Speech in Fish Audio, enter your script, and choose a library voice or a cloned voice to generate the audio. You can also use your own recording.
  10. Send the character video directly to Lip Sync, then select audio from your asset library or upload an audio file.
  11. Set the aspect ratio, resolution, and number of outputs, then generate the finished video.
  12. Preview the lip sync. From there, you can regenerate, extend, download, or reuse the result as a reference.

Tip: You can adapt this workflow to the project. For a quick video of someone speaking to the camera, skip video generation and use a character image with audio in Lip Sync. For a film or an ecommerce ad, create several reference images and video shots, then edit them into a complete piece.

The next sections cover image generation, video generation, and lip sync in more detail, with tips for getting consistent results that match your idea.


2. Create Your First AI Image

Watch this section in the full tutorial above: 01:07–01:43.

Step 1: Open the Image Tool

Open Fish Creative and select Image at the top of the page. Use the prompt box in the center to describe what you want to create. You can also add reference images to guide the character, product, composition, or visual style.

Tip: If you are starting with an idea, text is all you need. You do not have to prepare an image first.

Step 2: Write Your Image Prompt

An image prompt tells the model what to create. A clear prompt usually covers:

  • Subject: The person, product, or object in the image.
  • Setting: Where the subject is.
  • Composition: A close-up, full-body shot, overhead view, wide-angle shot, or another framing choice.
  • Style: Photography, illustration, 3D, a cinematic look, or another visual style.
  • Lighting: Soft light, natural light, neon, golden hour, or a similar direction. We will use a virtual creator facing the camera so the image is easy to work with when we move on to video and lip sync.

Example prompt:

A young tech content creator standing in a modern creative studio.
Medium close-up, facing the camera with a confident, friendly expression.
A clean, uncluttered background, soft cinematic lighting, and a photorealistic style.
Clear facial features and rich detail throughout the image.

Tip: If you are planning a vertical short video, specify vertical framing in the prompt and choose the matching aspect ratio in the settings.

Step 3: Add Reference Images If You Need Them

If you already have a character, product, or visual style in mind, upload a reference image. References help the model understand which features, composition, or style you want to preserve. They are also useful when you are creating a series featuring the same character.

Choose reference images with:

  • A clear subject and good resolution.
  • An unobstructed face or product.
  • Framing close to what you want in the final result.
  • Consistent lighting and color.

Tip: If you do not need a reference, skip this step.

Step 4: Choose a Model and Settings

Fish Creative currently offers several image models, including:

  • Nano Banana 2 Lite
  • GPT Image 2
  • Nano Banana 2
  • Nano Banana Pro
  • Seedream 5.0
  • Seedream 4.5

Tip: The model lineup is updated over time, so check Fish Creative for the current selection.

You can also adjust:

  • Aspect ratio: Landscape, portrait, or square.
  • Resolution: The dimensions of the output image.
  • Quality: The quality level for the result.
  • Number of outputs: Generate one image or several at once.

Tip: The Generate button shows the estimated credit cost and how many more images your balance can cover at the current settings. Check it before submitting, and adjust the model, quality, or number of outputs as needed.

Step 5: Generate and Choose an Image

Once you are happy with the prompt and settings, click Generate. Preview the results and choose the image that works best for your video.

For each generated image, you can:

  • Use it as a reference.
  • Recreate it.
  • Upscale it.
  • Use it for lip sync.
  • Generate a video from it.
  • Publicly share the generation record.
  • Download it.

Tip: If the subject looks right but you need a larger image or more detail, use Topaz HD to upscale the image before moving on.

Suggestions for Images You Will Use in Video or Lip Sync

When creating a character image for the next steps:

RecommendedAvoid
A front-facing view or a slight angle, with the eyes, nose, and mouth clearly visibleHands, hair, or other objects near the mouth
An aspect ratio chosen early for your publishing platform (9:16 is common for short vertical videos)Busy or crowded backgrounds

A clear subject and well-balanced composition give you a better starting point for natural movement.


3. Turn a Still Image into a Video

Watch this section in the full tutorial above: 01:44–02:22.

Step 1: Open the Video Tool

Select Video at the top of the page. You can generate a video from a text prompt alone or add images, video, or audio as reference material.

Fish Creative supports several ways to generate video:

  • Text to video.
  • Image to video.
  • First and last frame to video.
  • Reference image to video.
  • Video reference.
  • Audio reference combined with other inputs. For this tutorial, we will use image to video, starting with the character image we just created.

Step 2: Add Your Image or Other Reference Material

Select the image directly from your asset library. There is no need to download it and upload it again. Depending on the model and generation mode, you may also be able to add a first frame, last frame, reference images, reference videos, or reference audio.

Each type of reference serves a different purpose:

  • Single image: Establishes the character, setting, and starting composition.
  • First and last frames: Define how the video begins and ends.
  • Reference images: Help guide characters, objects, or the overall look.
  • Reference video: Provides a guide for action, camera movement, or pacing.
  • Reference audio: Provides sound and timing cues for models that support it.

Step 3: Write Your Video Prompt

An image prompt describes what is in the frame. A video prompt should focus on how the scene moves and changes. Useful details include:

  • The subject's actions.
  • Changes in expression or gaze.
  • Camera movement.
  • Changes in the background or environment.
  • Pacing and visual style.

Example prompt:

The person looks naturally into the camera, making small movements with their head and shoulders while maintaining a confident, friendly expression.
Movements are relaxed, smooth, and realistic.
The camera slowly pushes in, and the studio lighting stays consistent.
The video should feel natural and professional.

Tip: If you plan to add lip sync, start with simple, controlled movement. Sharp head turns, fast camera moves, and objects passing in front of the face can make lip sync more difficult.

Step 4: Choose a Video Model

Fish Creative supports a growing range of video models. Here are some of the available options:

  • MiniMax H3
  • Seedance 2.0
  • Seedance 2.5
  • Kling O3 Pro
  • Seedance 2.0 Mini
  • Kling 2.6 Pro
  • Hailuo 2.3
  • Veo 3.1
  • Kling V3 Pro
  • Seedance 2.0 Fast
  • Seedance 1.5 Pro

Tip: More models will be added over time, so check Fish Creative for the latest selection. Models may support different inputs, durations, and aspect ratios. Select a model first, then choose from the settings it offers.

Step 5: Set the Video Parameters

Depending on the model, you can adjust:

  • Aspect ratio: Options include 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, and Adaptive.
  • Duration: Supported lengths vary by model, currently ranging from roughly 4 to 30 seconds across the available options.
  • Resolution: Choose a resolution suited to a preview or your final output.
  • Number of outputs: Generate one video or several at once.

Tip: The Generate button displays the estimated credit cost and how many videos your remaining balance can cover at those settings.

Step 6: Generate and Review Your Video

Click Generate and wait for the video to finish. When you preview it, check that:

  • The face stays consistent.
  • Movements look natural and flow smoothly.
  • There are no obvious distortions.
  • Camera movement and pacing match your prompt.
  • The mouth stays clearly visible for most of the clip. You can use the result as a reference, recreate it, send it to Lip Sync, extend it, publicly share the generation record, or download it. If you are happy with the video, click Use for Lip Sync to continue.

4. Make Your Character Speak with Lip Sync

Watch this section in the full tutorial above: 02:23–02:44.

Fish Creative's Lip Sync tool works with both images and videos. Add an image or video of your character, pair it with audio, and generate a video with mouth movements that match the speech.

Step 1: Prepare Your Character Image or Video

You can:

  • Upload an image or video from your device.
  • Select something you have already created from your asset library.
  • Click Use for Lip Sync on an image or video result.

Step 2: Prepare Audio with Text to Speech

Lip Sync needs an existing audio track. You can upload a recording or create one with Fish Audio's Text to Speech tool.

To create audio in Fish Audio:

  1. Open Text to Speech.
  2. Type or paste the lines you want your character to say.
  3. Choose a voice from the Fish Audio voice library, or use a cloned voice you have created.
  4. Generate the audio and save it to your asset library.
  5. Return to Fish Creative and select that audio in Lip Sync.

Tip: You can also download the generated audio and upload the file to Lip Sync.

For this tutorial, try a short script:

Welcome to Fish Creative.
Start with a single image, turn it into a video, and make your character speak naturally.

Tip: For your first test, keep it to one or two sentences. Short audio clips make it easier to check the character's appearance, speaking pace, and lip sync.

Step 3: Add Your Visuals and Audio

Return to Fish Creative, select Lip Sync, and add your character image or video. Then choose audio from your asset library or upload an audio file.

Tip: There is also an optional prompt field. Leave it blank if you do not need additional instructions. If you want to specify something about the visuals, add a short description following the guidance in the interface.

Step 4: Choose Your Settings and Generate

With your inputs in place, set:

  • Aspect ratio.
  • Resolution.
  • Number of outputs. Check that you have selected the right character and audio, then click Generate.

Watch the finished video from beginning to end and check that:

  • Mouth movements follow the timing of the speech.
  • The face stays consistent.
  • The beginning and ending feel natural.
  • The video and audio are complete.
  • The character's movement fits the tone of the voice. You can use the result as a reference, recreate it, apply lip sync again, extend the video, publicly share the generation record, or download it.

Suggestions for Better Lip Sync

RecommendedAvoid
A front-facing character or a slight angleHair, hands, or objects near the mouth
Clear speech with little background noise, starting with a short script at a natural speaking paceSudden movements and fast head turns in video inputs

Tip: If the result is not working, try a different image, video, or audio track before adding lots of detailed instructions.


5. Three Projects to Try: Talking-Head Videos, AI Films, and Ecommerce Ads

Watch this section in the full tutorial above: 02:45–03:39.

These projects use the same image, video, and lip sync tools, but each calls for a slightly different approach. A talking-head video depends on a consistent character and a good vocal delivery. An AI film needs continuity between shots. An ecommerce ad needs to show the product accurately and communicate its selling points quickly.

Here is how to put the tools together for each one.

Project 1: An AI Presenter Video

Suppose you are making a 15-second vertical explainer about how to write a stronger opening for a short video. You want one presenter facing the camera and delivering a short script in a natural voice.

What you will need:

  • A character image facing forward or at a slight angle.
  • A script that takes about 15 seconds to read.
  • A voice from the Fish Audio library, a cloned voice, or your own recording.
  • A 9:16 output setting.

Suggested workflow:

  1. Create your character in Image, using a medium close-up with the person facing the camera.
  2. For a straightforward talking-head video with minimal movement, send the image directly to Lip Sync.
  3. If you want subtle breathing, a nod, or shoulder movement before the person starts speaking, generate a video with restrained movement first.
  4. Generate the script audio in Text to Speech and save it to your asset library.
  5. Add the image or video and the audio to Lip Sync, then generate the finished piece.

Character image prompt:

A young content creator sitting in a clean, modern studio.
Medium close-up, facing the camera with a relaxed, confident, approachable expression.
Soft side lighting, a slightly blurred background, and a photorealistic style.
Vertical 9:16 composition.

Sample script:

Do not start your short video by introducing yourself.
Tell viewers what they will get out of watching.
Give them a reason to stay in those first three seconds.

Tip: Keep the gaze, expression, lip movements, and voice working together. The character does not need to move much. A steady presence helps viewers focus on what is being said.

Project 2: An AI Short Film

Imagine a 20- to 30-second suspense scene: a detective enters an abandoned station on a rainy night and spots a mysterious figure in the distance. Plan this as a series of short shots instead of generating the whole sequence in one go.

Start with three shots:

  1. Establishing shot: An abandoned station on a rainy night. The camera slowly moves toward the entrance.
  2. Action shot: The detective enters the station from the side of the frame, stops, and looks up to survey the space.
  3. Dialogue close-up: A close-up of the detective delivering a key line.

Suggested workflow:

  1. Use Image to create a character reference for the detective, then an image of the station.
  2. Use As Reference to carry the selected character and setting into later images, helping the shots look consistent.
  3. Generate the three shots separately. For the establishing shot, use text to video or reference images. For the action shot, try a single image, first and last frames, or a video reference. Keep the detective facing forward or at a slight angle in the dialogue close-up.
  4. Generate the dialogue audio in Fish Audio, then apply lip sync only to the close-up where the character speaks.
  5. Download the shots and assemble them in your usual video editor: setting, action, then dialogue.

Character reference prompt:

A detective in their early forties, wearing a long, dark trench coat.
Short hair, a tired but alert expression, and rain droplets on the collar.
Cinematic, photorealistic character reference with cool blue tones.
Clearly defined facial features in front and three-quarter views.

Video prompt for shot 1:

An abandoned train station on a rainy night.
Puddles reflect dim lights as fine rain continues to fall.
A thin mist drifts slowly in the distance.
The camera moves smoothly from outside the station toward the entrance.
A suspenseful cinematic atmosphere with cool blue tones.

Sample dialogue:

I know you have been waiting for me.

Tip: A common problem with AI films is that each shot looks good on its own, but the characters, clothing, and locations do not match across cuts. Establishing a character reference and generating one shot at a time makes consistency easier than repeatedly describing the same character from scratch.

Project 3: An Ecommerce Product Ad

Suppose you are making a 15-second vertical ad for a facial serum. Viewers need to see the bottle and texture clearly, while a presenter briefly explains the main benefits.

Plan four shots:

  • 0–3 seconds: A close-up of the bottle so viewers immediately recognize the product.
  • 3–7 seconds: A person picks up the product, showing it in use.
  • 7–11 seconds: A close-up of the serum's texture or a product detail.
  • 11–15 seconds: A presenter speaks to the camera, followed by a closing frame featuring the product and brand.

Suggested workflow:

  1. Upload a clear product photo as a reference for its appearance.
  2. In Image, create a product still life, an image of someone holding the product, and a texture close-up. Once you have results you like, use As Reference to create more assets in the same style.
  3. Turn the still life and usage images into short videos. Keep product shots simple: a slow push-in, rotation, or change in lighting. Complex movement can distort the bottle or packaging text.
  4. Prepare a separate front-facing image or a steady video for the presenter, and generate audio for the product message in Fish Audio.
  5. Use Lip Sync to create the presenter segment, then download the clips, edit them together, and add captions and finishing touches.

Product image prompt:

A premium facial serum in a clear glass bottle on a pale stone countertop.
The front of the bottle faces the camera.
Soft morning light falls from behind and to one side, with a few water droplets and clean reflections around the product.
High-end skincare advertising photography, a clearly defined product silhouette, and space for copy.
Vertical 9:16 composition.

Product video prompt:

The camera slowly moves toward the serum bottle as water droplets slide gently down the glass.
The background lighting shifts softly.
The product remains steady and clearly defined, with restrained movement and the look of a high-end skincare ad.

Sample script:

Lightweight, fast-absorbing hydration in one step.
Use it morning and night for a simpler skincare routine.

Tip: Product recognition matters most in ecommerce content. Use clear reference images and give the product enough space in the frame. If packaging text, logos, or the bottle's shape must be exact, inspect the video frame by frame before publishing and use the actual product assets to correct any discrepancies in post-production.


6. Useful Buttons and When to Use Them

Watch this section in the full tutorial above: 03:40–04:32.

There is more to Fish Creative than the first click of Generate. The buttons beneath each result let you send an image or video to the next tool or keep refining a version you like. Here are the ones you will use most often.

Prompt Optimization: Add Detail to a Starting Idea

If your prompt is something brief, such as "a person in a café" or "a product ad video," prompt optimization can help you expand it. It adds details about the subject, setting, composition, lighting, action, or camera movement.

Use it in three steps:

  1. Write down the essentials that must stay the same, such as the character's identity, product name, required action, and aspect ratio.
  2. Use prompt optimization to add visual detail.
  3. Read the revised prompt before generating. Remove anything you do not need, and make sure the character, product, and actions still match your original idea.

Tip: Prompt optimization is useful when you know what you want but are not sure how to describe it. Treat the result as a draft to review before you generate.

Recreate: Try Another Version

Use Recreate when a result is close, but the expression, composition, action, or camera movement needs work. It lets you build on an existing generation without setting up the task again from a blank page.

Before recreating, work out what needs to change:

  • Wrong subject or setting: Revise the core description in your prompt.
  • Composition needs work: Specify the framing, camera position, and placement of the subject.
  • Too much movement: Ask for fewer actions and use words such as "subtle," "slow," and "steady."
  • You just want another variation: Keep the main settings and generate again.

Tip: Recreate works best for small adjustments to an idea you are already happy with.

As Reference: Build on a Result You Like

Click As Reference to add the current image or video to a new generation task. Use it to carry forward a character, product, setting, visual style, or movement you have already established.

Common uses include:

  • Placing the same character in different settings.
  • Creating several ad assets for the same product.
  • Using an existing shot or movement to guide a later video.
  • Building a series from one successful result.

Tip: A reference guides the next generation; it does not simply duplicate the original. Your prompt still needs to say what should stay and what should change. For example: "Keep the character's face and clothing consistent, and change the background to a street at night."

Image Upscaling: Recover Detail with Topaz HD

Use image upscaling when you are happy with the subject and composition but need more resolution or finer detail. Fish Creative uses Topaz HD to upscale individual images and recover detail.

Tip: This is useful before downloading a final image, creating a cover image, or moving into video generation. If the character's anatomy or the product's shape is already wrong, upscaling usually will not fix it. Recreate the image or revise the prompt first.

Generate Video: Animate an Image You Like

The Generate Video button beneath an image sends that exact result straight to the video tool, saving you a download and upload.

Tip: Once you are there, focus the prompt on action, camera movement, and changes over time. You do not need another lengthy description of the still image. Try: "The person raises their head slightly, the camera slowly moves closer, and the background lights gradually brighten."

Use for Lip Sync: Give Your Character a Voice

You can select Use for Lip Sync on both image and video results. With an image, the tool creates a talking video from the image and audio. With a video, it synchronizes mouth movements with the audio while working with the existing motion.

Tip: Going straight from an image to lip sync works well for a steady talking-head video. Starting with video is useful when you want breathing, nodding, or body movement. In either case, make sure the face is clear and the mouth is unobstructed.

Extend Video: Continue an Existing Clip

Use Extend Video when you like a clip's action, camera movement, and style but need it to run longer. Extending lets you continue the existing scene and movement.

Tip: Check the last frame before extending. A distorted character, fast movement, or a subject about to leave the frame may cause problems in the continuation. A clip that ends in a steady, clear frame is usually easier to extend smoothly.

Asset Library and History: Find and Reuse Your Work

Your creations are saved automatically in the asset library and generation history. From there, you can find previous images, videos, and audio and add them directly to new tasks.

Tip: Keep a few key versions: your original reference, the first usable result, the final high-quality output, and the steady video you use for lip sync. That makes it easier to revise one stage without starting the whole process over.

Public Sharing and Download: Check the Final Result

Download saves the result to your device so you can edit it, use it in a design, or publish it. Publicly sharing a generation record lets others see what you have created. Use that option only when you are ready to make the content public.

Tip: Before downloading, check the aspect ratio, resolution, character and product details, and lip sync. Decide whether the video still needs captions or branding in post-production.

Credit Cost and Remaining Generations: Check Before You Generate

The Generate button shows the estimated credit cost for your current settings and how many images or videos your balance can cover at those settings. As you change the resolution, duration, or number of outputs, check how the estimate changes before submitting.

Tip: This helps you plan test runs and final outputs. Use early generations to work out the composition and movement, then choose your final settings once you are happy with the direction.


Start Creating with Fish Creative

Fish Creative brings your creative tools together in one workspace. In this tutorial, you used image generation, video generation, and lip sync to go from a single prompt to a character that speaks with a Fish Audio voice.

Your first project can be simple: one clearly framed character, a little movement, and a sentence or two. Once you are comfortable with the process, try reference images, first and last frames, video references, or more ambitious camera work.

Start Creating your own work with Fish Creative →

Questions Fréquemment Posées

Can I Generate Images and Videos from Text Alone?
Yes. The Image tool accepts text prompts, and the Video tool supports text to video. You can also add image, video, or audio references for more control.
Can I Use a Still Image for Lip Sync?
Yes. Lip Sync accepts both images and videos. If you want more body movement in the finished result, you can turn the image into a video first, then apply lip sync.
Do All Models Have the Same Settings?
No. Supported inputs, aspect ratios, durations, and resolutions vary by model. Check the options shown after selecting your model.
Do I Need to Save Generated Results Manually?
Your results are saved to the asset library, where you can reuse them across tools. You can also download them to your device.
Frank Lu

Frank Lu

Product, Operation & Data Engineer

Lire plus de Frank Lu

Créez des voix qui semblent réelles

Commencez à générer un son de la plus haute qualité dès aujourd'hui.

Vous avez déjà un compte ? Se connecter