Oct 7, 2026Guide

Image Generation MCP: Images, Video & Lip Sync with Fish Audio

Image Generation MCP: Images, Video & Lip Sync with Fish Audio

Fish Audio's MCP server now generates and edits images, creates videos, syncs lips to new audio, and turns portraits into talking avatars — from Claude Code, Claude.ai, Cursor, Codex, Windsurf, and Devin. One sign-in, dozens of models, and a price quote before every image, video, lip sync, or avatar job.

October 2026 | Fish Audio MCP now supports image generation, video generation, and lip sync


Ask your coding agent for a product shot, a five-second teaser, or a talking-head update, and you'll usually hit a wall. Some agents can't create images at all, others offer a single built-in image model, and video is rarely an option. The usual fix is an image generation MCP server, but most of them connect a single model and ask you to create and paste a separate API key for every provider.

Fish Audio takes a different approach. The same MCP connection your agent already uses for text-to-speech, transcription, and voice search can now generate and edit images, create videos, sync lips to new audio, and turn a portrait into a talking avatar. You sign in once, pick from dozens of image, video, and lip sync models, and approve a quote before any image, video, lip sync, or avatar job is charged.

Because the voice tools live on the same server, your agent can also do something no single-purpose MCP can: write a script, voice it with Fish Audio, and hand that audio straight to a talking-avatar model — in one conversation.

Already connected to the Fish Audio MCP server? No update needed. Ask your agent to "list the Fish Audio media models" and the new tools are ready to use.

New to Fish Audio? Sign up free → and connect with your agent in under a minute.


What's New in the Fish Audio MCP Server

Start with what the update adds. With it, the Fish Audio MCP server works as an image generation MCP, a video generation MCP, and a lip sync tool in one. The new media tools sit alongside the existing audio tools and draw from the same credits as the Fish Audio web app.

Why one connection instead of one MCP per model?

  • No API keys to collect. You sign in with your Fish Audio account. There is no provider key to request, paste into a config file, or rotate.
  • One balance. Every model draws from your Fish Audio plan credits, so switching from an image model to a video model doesn't mean topping up another account.
  • Switching models is a sentence, not a setup. Ask for a different model in plain language, and the agent looks up its options and quotes the new job.

Here's what your agent can do now:

CapabilityWhat you give itAvailable models include
Text-to-imageA promptGPT Image 2, GPT Image 2.5 (Flare and Sunburst), Seedream 4.5 and 5.0, Nano Banana models
Image editing and upscalingAn image + instructionsThe models above, plus Topaz HD for upscaling
Text-to-video and image-to-videoA prompt, optionally a starting imageSeedance, Kling, Veo 3.1, MiniMax H3, WAN 3.0, Gemini Omni Flash
Lip sync (video-driven)A video + new audioSync LipSync v3, Sync LipSync v2 Pro, PixVerse Lip Sync, Veed Lip Sync
Talking avatars (image-driven)A portrait + audioCreatify Aurora Avatar, HeyGen Avatar 4, Veed Fabric 1.0 Avatar, Infinitalk Avatar

Model lineups change over time. Your agent always works from the live list, so it sees new models as soon as they're added.

Image generation, editing, and upscaling are available on the Free plan. Video generation, lip sync, and talking avatars require Paid Plus Plan or higher.


Generate, Edit, and Upscale Images

Describe the image, and the agent picks a model, reads its options, such as aspect ratio, resolution and quality, then quotes the job. To edit, give it an image and an instruction like "keep the subject, change the background to a warm sunset." Images usually finish in under a minute, and you can request several at once depending on your plan.

A kitten on a soft blanket, generated from a one-line prompt through the Fish Audio MCP server in Codex

Generate Video from Text, Images, or Reference Video

The video tools turn the MCP server into a video generation MCP for any agent that supports remote MCP servers. Start from a text prompt, animate a still image, or — on models that support it — guide the result with reference media. Aspect ratio, resolution, and duration options vary by model, and the agent reads them before quoting. Videos can take a few minutes; the agent checks progress on its own and returns a link when the file is ready.

Picture from a short AI-generated video of a kitten drinking water, created with Seedance 2.0 through the Fish Audio MCP server

Lip Sync and Talking Avatars

Both capabilities make someone on screen speak a new audio track. The difference is what you start from: an existing video or a single image.

Video-driven: re-sync existing footage. Give the model a video of someone speaking and a new audio track, and it re-times the mouth movements to match. It's built for updating footage you already have: changing a line in a product video, recording a fix without reshooting, or localizing a talking-head clip into another language. These models work from video and audio alone and don't take a text prompt.

Image-driven: animate a portrait. Give the model a single portrait and an audio file, and it animates the face, head, and expression to match the speech. Creatify Aurora Avatar also accepts an optional prompt to control framing, background, and lighting, while Infinitalk Avatar requires a prompt and produces a more expressive look.

Use material you have the rights to. Before every lip sync or avatar job, the confirmation message reminds you to use only footage, portraits, and voices you're permitted to use.

Each of these tools is useful on its own. The bigger change is what happens when they work together.


From Script to Talking Avatar in One Conversation

This is where having voice and video on the same MCP server pays off. Most avatar tools assume you show up with finished audio. With Fish Audio, your agent can produce the audio itself and pass it straight to the avatar model.

Here's the flow:

  1. Write the script. Paste it in, or let the agent draft it.
  2. Voice it. The agent calls text_to_speech with a voice from the Fish Audio voice library, your own voice clone, or a voice you created with Voice Design. Add audio tags like [excited] or [whispering] to shape the delivery. Text-to-speech is billed by the length of the text and charged as soon as the speech is generated, without a separate quote.
  3. Pick a face. Upload a portrait, or generate one with an image model in the same session.
  4. Animate it. The agent sends the portrait and the speech it just generated to an avatar model. There's no need to download the audio and upload it again — the finished speech is passed in directly.
  5. Approve and receive. You see a quote for the avatar video, confirm it, and get a link to the finished file.

We tested this exact path: a six-second clip generated with text_to_speech was accepted directly as the avatar's audio input, and the quote reflected its exact duration — no download and re-upload.

Diagram of a script becoming a talking-avatar video through the Fish Audio MCP server: script, text-to-speech, portrait, avatar model, finished video

Copy this into your agent to try it:

Use Fish Audio to make a 9:16 talking-avatar video:
1. Turn this script into speech with a warm female English voice:
   "Hi, welcome to Fish Audio. Today I'll show you lip sync and talking avatars."
2. Use my uploaded portrait with HeyGen Avatar 4 at 720p and have her speak that audio.
3. Quote the price first. Generate after I confirm, then send me the link.

The same idea works for lip sync. Point the agent at an existing talking-head video and give it a new line:

Use Fish Audio to give my uploaded talking-head video a new line:
1. Generate speech with my cloned voice: "Today's new arrivals are 20% off. Link in the comments."
2. Use Sync LipSync v3 to match the lips in the video to the new audio.
Quote first, and generate after I confirm.

Talking avatars and lip sync require a paid Plus plan or higher. See plans →

Reuse Any Result as the Next Input

Chaining isn't limited to audio. Any finished result in your workspace can feed the next job without leaving the conversation: generate an image, then ask the agent to "turn the image you just made into a four-second video with a slow push-in." The agent passes the image along by reference — nothing to download, nothing to re-upload.

For files on your computer, the agent creates a temporary upload slot and uploads the file for you. If your agent can't upload files directly, it gives you a link where you can drag the file in from your browser.

All of this happens inside the agent itself. That matters more than it might seem: on their own, most AI agents can't create video or talking avatars, and even those that generate images are tied to a single built-in model. Connecting an MCP server is what closes that gap.


Why AI Agents Need MCP for Image and Video Generation

Comparison of a manual multi-app workflow versus an AI agent generating images and video directly through one MCP connection

Today's AI agents are remarkably capable. They can plan a campaign, write the script, and build the page it lives on. When it comes to visuals, though, support is uneven. Some assistants, like Claude, don't generate photos or illustrations the way image generation tools do. Others, like ChatGPT and Codex, include image generation, but it's tied to one provider's model. Video, lip sync, and talking avatars are usually out of reach.

That gap creates a familiar routine. You write a prompt in your agent, paste it into a separate image or video app, wait, download the file, and bring it back into your project. Every revision repeats the loop. The alternative many developers reach for is an MCP server per model, which means a separate provider API key and config entry for each one.

The Model Context Protocol (MCP) is the standard way to give an agent new tools, and it's what lets Fish Audio remove those extra steps and open up a choice of models. Your agent calls the generation tools directly, the results come back as links, and the next step, whether a revision, an animation, or a voiceover, happens in the same conversation. Image generation in Claude Code, Cursor, or any other supported agent becomes one more tool call.

Getting there takes a single command.


How to Set Up the Image Generation MCP

The server lives at https://api.fish.audio/mcp and uses streamable HTTP. Sign-in happens in your browser with your Fish Audio account — no API key, no environment variables. Pick your agent below.

Looking for the Docs MCP? Our earlier guide to llms.txt, MCP, and Agent Skills covers docs.fish.audio/mcp, a read-only server that lets coding agents look up Fish Audio API documentation. The server in this post is different: it signs into your account and creates content. You can run both side by side.

Claude Code

claude mcp add --transport http fish-audio https://api.fish.audio/mcp

Then run /mcp inside Claude Code to sign in.

Claude.ai

Go to Settings → Connectors → Add custom connector and enter https://api.fish.audio/mcp. Claude walks you through the sign-in.

Cursor

Open the command palette (Cmd/Ctrl+Shift+P) → Open MCP settings → Add custom MCP, and add:

{
  "mcpServers": {
    "fish-audio": { "url": "https://api.fish.audio/mcp" }
  }
}

The cursor prompts you to sign in on first use.

Codex CLI

codex mcp add fish-audio https://api.fish.audio/mcp
codex mcp login fish-audio

The login command opens a browser window. Check the connection with codex mcp get fish-audio.

Windsurf

Go to Settings → Cascade → MCP Servers → View raw config (~/.codeium/windsurf/mcp_config.json) and add:

{
  "mcpServers": {
    "fish-audio": { "url": "https://api.fish.audio/mcp" }
  }
}

Devin CLI

devin mcp add fish-audio --transport http --scope user --url https://api.fish.audio/mcp --scopes mcp
devin mcp login fish-audio --scopes mcp

Confirm it's connected with the devin mcp list. Because Devin works in your shell, it can also save the results locally — in our test, it downloaded a finished video to the Downloads folder on request.

Sign In and Choose Your Team

Every client opens the same Fish Audio sign-in page. You approve the connection once and choose which team it works in; the client renews access automatically after that.

A connection is tied to the team you choose. It can only see that team's workspaces and use that team's credits. To work in a different team, add the server again and choose that team at sign-in.

Connected? Try your first prompt: "Create a 16:9 watercolor image of a lighthouse at dawn. Show me the price first, then generate it." Start free with Fish Audio →


How Each Generation Works: Quote, Confirm, Generate

Once connected, your agent can run paid models on your behalf. That raises a fair question: what stops it from spending more than you meant to? The answer is built into the server, not left to the agent's judgment.

  1. Discover. The agent lists available models and reads the one it needs — supported options, defaults, and input limits.
  2. Quote. Before any image, video, lip sync, or avatar job, the agent requests an estimate. Estimates are free. Plan limits are checked here too, so if a request needs a higher plan, you find out before you're asked to approve anything.
  3. Confirm. The agent shows you the price and waits. Nothing is charged until you approve.
  4. Generate. The agent submits the exact request you approved. The server prices it again at submission; if the price has changed, the job is rejected and the agent must quote again. The charge never exceeds the amount you approved.
  5. Deliver. The agent checks progress every few seconds and returns a link to the finished file, along with what was charged.

Voice tools are billed differently. Text-to-speech and speech-to-text don't go through the quote step. They're charged as soon as they run, with text-to-speech billed by the length of the text.

Two more safeguards run in the background:

  • Failed jobs are refunded automatically. If a generation fails, the credits come back without you asking.
  • Retrying the same request won't charge twice. Every paid request carries a unique key. If the agent resends the same request with the same key after a timeout, the server returns to the original job instead of starting and charging for a new one.

Many agents add their own layer on top: Codex, for example, asks for your permission before each tool call.


Under the Hood: The MCP Media Tools

For developers, these are the media tools the agent calls behind the scenes:

ToolCharges credits?What it does
list_my_workspacesNoLists workspaces in the connected team
list_media_models / get_media_modelNoLists models and reads one model's options and input limits
create_media_uploadNoCreates a one-hour upload slot for an image, video, or audio file
estimate_image_generation / estimate_image_edit / estimate_video_generationNoValidates a request and returns a quote
generate_image / generate_image_edit / generate_videoYesSubmits the approved job
get_generation_status / get_generation_resultNoTracks progress and returns the finished files and billing

Lip sync and talking avatars run through generate_video. It recognizes the job from its inputs: a video plus audio means lip sync, and a portrait plus audio means an avatar. Audio tools such as text_to_speech, speech_to_text, search_voices, create_voice_clone, and the Story Studio tools sit in the same server.

You never have to call any of these tools yourself, though. In practice, you describe the result you want, and the agent chains the right calls together.


Example Prompts by Use Case

Here are a few starting points, grouped by what you're making:

Use caseWhat you sayWhat the agent does
Product and app assets"Generate a 1:1 image of a ceramic coffee mug on a marble counter in soft morning light. Give me a few variations and show the price first."Picks an image model, quotes the batch, generates after you confirm
Product and app assets"Upload screenshot.png and replace the background with a clean gradient for our app store listing. Estimate first."Uploads your file, quotes the edit, applies it after you confirm
Marketing and social video"Make a 9:16, five-second video of sneakers splashing through a puddle in slow motion. Quote the credits, then generate."Quotes a text-to-video job and returns a video link
Marketing and social video"Turn the image you just generated into a short video with a slow push-in."Reuses the image directly as the starting frame
Updating and localizing video"Translate this script into Spanish, voice it with my own voice clone, then lip sync demo.mp4 to the Spanish audio with Sync LipSync v2 Pro. Quote before generating."Translates the script, generates the speech, uploads the video, and quotes the lip sync job
Presenters and avatars"Generate a portrait of a friendly presenter in a bright studio, then have her read release-notes.md as a 9:16 talking-avatar video with Creatify Aurora Avatar."Creates the portrait and the speech, then quotes the avatar video
Account"How many Fish Audio credits do I have left?"Checks your balance

In coding agents like Claude Code, Codex, and Devin, these tools work alongside the shell. The agent can batch-generate images for every product in a folder, upload local recordings, or download finished files into your project.


Plans, Credits, and Limits

Before you try these, it helps to know how credits work and which plan each capability needs.

Which Credits It Uses

Media generation through MCP uses your Fish Audio web plan credits, the same credits you spend on the web app. It does not use Developer API credits, which are a separate balance for the Fish Audio API. The two are purchased in different places, so it's worth knowing which one to top up.

To add web credits or change plans, go to Team → Plans. Since each MCP connection is tied to the team you chose at sign-in, generations draw from that team's credits.

Which Plan You Need

What your agent can generate depends on your plan. Plan limits are checked when the agent requests a quote, so a job your plan doesn't cover is flagged before you're ever asked to approve it.

  • Image generation, editing, and upscaling are available on the Free plan, which includes a daily image allowance.
  • When generating images from text, each text-to-image request can ask for up to 2 images on Free, 4 on Plus, and 8 on Pro and Max. Image edits return one image per request.
  • Video generation, lip sync, and talking avatars require Paid Plus, Pro, or Max plan.
  • Video, lip sync, and avatar jobs return one output per request.

If you're starting on Free, image generation is the easiest way to try the workflow. When you're ready for video, lip sync, or talking avatars, upgrading to Paid Plus plan or higher unlocks them.

Prices

Every model is priced differently, and the cost of a job also depends on the settings you choose, such as resolution, duration, or the number of images. Rather than relying on a price list, ask your agent for a quote: it's free, and it reflects the current price. However, note that text-to-speech is the exception: it's billed by the length of the text and charged as soon as it runs.

Because quotes cost nothing, you can compare a few models or settings for the same idea and pick the one that fits your budget before the job is charged.

Uploads

Some jobs start with your own files: an image to edit or animate, a video to lip sync, or audio for an avatar. The agent handles the upload for you, within these limits:

  • Images: JPEG, PNG, or WebP, up to 20 MB
  • Video: MP4, MOV, WebM, or MKV, up to 1 GB
  • Audio: MP3, WAV, M4A, AAC, OGG, or FLAC, up to 50 MB

Individual models can have tighter limits. For example, Creatify Aurora Avatar accepts 1 to 60 seconds of audio. The agent reads each model's limits before quoting, and if a provider rejects input after submission, the failed job is refunded. Results you've already generated in the same workspace don't need to be uploaded again; the agent passes them along directly.

Timing

Images usually finish within a minute. Videos, lip sync, and avatars can take several minutes. You don't need to keep checking: the agent tracks progress on its own and returns a link as soon as the file is ready.


Final Thoughts: Your Agent Is Now a Media Studio

AI agents can already plan a launch, draft the landing page, and ship the code. Until now, visuals were where most of them stopped: a single built-in image model at best, with video, lip sync, and talking avatars handed back to a person and a separate tool. This update closes that gap. One image generation MCP gives your agent voice, images, video, lip sync, and talking avatars, with no API keys to manage and a quote before every media job.

The bigger shift is how these pieces connect. A script becomes a voice, the voice becomes a talking avatar, and an image becomes a video, with each step handing its result to the next instead of files moving back and forth between apps. Voice is what makes a video feel like someone is actually speaking to you, and it sits at the core of what Fish Audio builds. Putting it in the same conversation as images and video means your agent can carry an idea from a few lines of text to a finished clip.

As new models are added, they appear in your agent's live model list, with nothing to reconfigure. The easiest way to see what that unlocks is to connect the server and ask for something you'd normally open three different tools to make.

New to Fish Audio? Sign up free → Connect the MCP server in under a minute and start with image generation, included in the free plan.

Ready for video, lip sync, and avatars? See plans → Plus and above unlock video generation, lip sync, and talking avatars through MCP.

Read the MCP server docs →

Frequently Asked Questions

Can Claude Code generate images?
Not natively, but it can with an MCP server. Add the Fish Audio MCP server with one command, sign in, and Claude Code can generate and edit images, create videos, and produce lip sync and talking-avatar videos.
Is there a video generation MCP for Claude Code and Cursor?
Yes. The Fish Audio MCP server handles text-to-video, image-to-video, lip sync, and talking avatars in Claude Code, Cursor, Codex, Windsurf, Claude.ai, and Devin. Video features require a Plus plan or higher.
Do I need an API key to use the Fish Audio MCP server?
No. You connect by signing in with your Fish Audio account in your browser. There is no API key or environment variable to configure.
Does it use my Developer API credits?
No. Generations through MCP use your web plan credits, the same as the Fish Audio web app. Developer API credits are a separate balance and aren't affected. You can manage web credits under [Team → Plans](https://fish.audio/app/team/?tab=plans).
Can the agent spend credits without asking me?
Not for media generation. Every image, video, lip sync, and avatar job is quoted first, and nothing is charged until you approve. The server re-checks the price at submission and rejects the job if it has changed, so the charge never exceeds the amount you approved. However, text-to-speech and speech-to-text work differently: they're charged as soon as they run, with text-to-speech billed by the length of the text.
What happens if a generation fails?
The credits are refunded automatically. If the agent retries a request after a timeout, the server recognizes the duplicate and won't charge twice.
Can I use audio generated with Fish Audio text-to-speech for lip sync or avatars?
Yes. Speech your agent generates with the Fish Audio MCP server can go straight into lip sync or avatar job in the same workspace, without downloading and re-uploading it.
Which AI agents work with the Fish Audio MCP server?
Any client that supports remote MCP servers over streamable HTTP with OAuth sign-in. Setup steps for Claude Code, Claude.ai, Cursor, Codex CLI, Windsurf, and Devin CLI are above, and the [MCP server documentation](https://docs.fish.audio/overview/mcp) stays up to date as more clients are added.
Sabrina Shu

Sabrina Shu

Sabrina is part of Fish Audio's support and marketing team, helping users get the most out of AI voice products while turning launches, updates, and customer insights into clear, practical content.

Read more from Sabrina Shu

Create voices that feel real

Start generating the highest quality audio today.

Already have an account? Log in