Image Generation MCP: Images, Video & Lip Sync with Fish Audio
Fish Audio's MCP server now generates and edits images, creates videos, syncs lips to new audio, and turns portraits into talking avatars — from Claude Code, Claude.ai, Cursor, Codex, Windsurf, and Devin. One sign-in, dozens of models, and a price quote before every image, video, lip sync, or avatar job.
October 2026 | Fish Audio MCP now supports image generation, video generation, and lip sync
Ask your coding agent for a product shot, a five-second teaser, or a talking-head update, and you'll usually hit a wall. Some agents can't create images at all, others offer a single built-in image model, and video is rarely an option. The usual fix is an image generation MCP server, but most of them connect a single model and ask you to create and paste a separate API key for every provider.
Fish Audio takes a different approach. The same MCP connection your agent already uses for text-to-speech, transcription, and voice search can now generate and edit images, create videos, sync lips to new audio, and turn a portrait into a talking avatar. You sign in once, pick from dozens of image, video, and lip sync models, and approve a quote before any image, video, lip sync, or avatar job is charged.
Because the voice tools live on the same server, your agent can also do something no single-purpose MCP can: write a script, voice it with Fish Audio, and hand that audio straight to a talking-avatar model — in one conversation.
Already connected to the Fish Audio MCP server? No update needed. Ask your agent to "list the Fish Audio media models" and the new tools are ready to use.
New to Fish Audio? Sign up free → and connect with your agent in under a minute.
What's New in the Fish Audio MCP Server
Start with what the update adds. With it, the Fish Audio MCP server works as an image generation MCP, a video generation MCP, and a lip sync tool in one. The new media tools sit alongside the existing audio tools and draw from the same credits as the Fish Audio web app.
Why one connection instead of one MCP per model?
- No API keys to collect. You sign in with your Fish Audio account. There is no provider key to request, paste into a config file, or rotate.
- One balance. Every model draws from your Fish Audio plan credits, so switching from an image model to a video model doesn't mean topping up another account.
- Switching models is a sentence, not a setup. Ask for a different model in plain language, and the agent looks up its options and quotes the new job.
Here's what your agent can do now:
| Capability | What you give it | Available models include |
|---|---|---|
| Text-to-image | A prompt | GPT Image 2, GPT Image 2.5 (Flare and Sunburst), Seedream 4.5 and 5.0, Nano Banana models |
| Image editing and upscaling | An image + instructions | The models above, plus Topaz HD for upscaling |
| Text-to-video and image-to-video | A prompt, optionally a starting image | Seedance, Kling, Veo 3.1, MiniMax H3, WAN 3.0, Gemini Omni Flash |
| Lip sync (video-driven) | A video + new audio | Sync LipSync v3, Sync LipSync v2 Pro, PixVerse Lip Sync, Veed Lip Sync |
| Talking avatars (image-driven) | A portrait + audio | Creatify Aurora Avatar, HeyGen Avatar 4, Veed Fabric 1.0 Avatar, Infinitalk Avatar |
Model lineups change over time. Your agent always works from the live list, so it sees new models as soon as they're added.
Image generation, editing, and upscaling are available on the Free plan. Video generation, lip sync, and talking avatars require Paid Plus Plan or higher.
Generate, Edit, and Upscale Images
Describe the image, and the agent picks a model, reads its options, such as aspect ratio, resolution and quality, then quotes the job. To edit, give it an image and an instruction like "keep the subject, change the background to a warm sunset." Images usually finish in under a minute, and you can request several at once depending on your plan.
Generate Video from Text, Images, or Reference Video
The video tools turn the MCP server into a video generation MCP for any agent that supports remote MCP servers. Start from a text prompt, animate a still image, or — on models that support it — guide the result with reference media. Aspect ratio, resolution, and duration options vary by model, and the agent reads them before quoting. Videos can take a few minutes; the agent checks progress on its own and returns a link when the file is ready.
Lip Sync and Talking Avatars
Both capabilities make someone on screen speak a new audio track. The difference is what you start from: an existing video or a single image.
Video-driven: re-sync existing footage. Give the model a video of someone speaking and a new audio track, and it re-times the mouth movements to match. It's built for updating footage you already have: changing a line in a product video, recording a fix without reshooting, or localizing a talking-head clip into another language. These models work from video and audio alone and don't take a text prompt.
Image-driven: animate a portrait. Give the model a single portrait and an audio file, and it animates the face, head, and expression to match the speech. Creatify Aurora Avatar also accepts an optional prompt to control framing, background, and lighting, while Infinitalk Avatar requires a prompt and produces a more expressive look.
Use material you have the rights to. Before every lip sync or avatar job, the confirmation message reminds you to use only footage, portraits, and voices you're permitted to use.
Each of these tools is useful on its own. The bigger change is what happens when they work together.
From Script to Talking Avatar in One Conversation
This is where having voice and video on the same MCP server pays off. Most avatar tools assume you show up with finished audio. With Fish Audio, your agent can produce the audio itself and pass it straight to the avatar model.
Here's the flow:
- Write the script. Paste it in, or let the agent draft it.
- Voice it. The agent calls
text_to_speechwith a voice from the Fish Audio voice library, your own voice clone, or a voice you created with Voice Design. Add audio tags like[excited]or[whispering]to shape the delivery. Text-to-speech is billed by the length of the text and charged as soon as the speech is generated, without a separate quote. - Pick a face. Upload a portrait, or generate one with an image model in the same session.
- Animate it. The agent sends the portrait and the speech it just generated to an avatar model. There's no need to download the audio and upload it again — the finished speech is passed in directly.
- Approve and receive. You see a quote for the avatar video, confirm it, and get a link to the finished file.
We tested this exact path: a six-second clip generated with text_to_speech was accepted directly as the avatar's audio input, and the quote reflected its exact duration — no download and re-upload.
Copy this into your agent to try it:
Use Fish Audio to make a 9:16 talking-avatar video:
1. Turn this script into speech with a warm female English voice:
"Hi, welcome to Fish Audio. Today I'll show you lip sync and talking avatars."
2. Use my uploaded portrait with HeyGen Avatar 4 at 720p and have her speak that audio.
3. Quote the price first. Generate after I confirm, then send me the link.
The same idea works for lip sync. Point the agent at an existing talking-head video and give it a new line:
Use Fish Audio to give my uploaded talking-head video a new line:
1. Generate speech with my cloned voice: "Today's new arrivals are 20% off. Link in the comments."
2. Use Sync LipSync v3 to match the lips in the video to the new audio.
Quote first, and generate after I confirm.
Talking avatars and lip sync require a paid Plus plan or higher. See plans →
Reuse Any Result as the Next Input
Chaining isn't limited to audio. Any finished result in your workspace can feed the next job without leaving the conversation: generate an image, then ask the agent to "turn the image you just made into a four-second video with a slow push-in." The agent passes the image along by reference — nothing to download, nothing to re-upload.
For files on your computer, the agent creates a temporary upload slot and uploads the file for you. If your agent can't upload files directly, it gives you a link where you can drag the file in from your browser.
All of this happens inside the agent itself. That matters more than it might seem: on their own, most AI agents can't create video or talking avatars, and even those that generate images are tied to a single built-in model. Connecting an MCP server is what closes that gap.
Why AI Agents Need MCP for Image and Video Generation
Today's AI agents are remarkably capable. They can plan a campaign, write the script, and build the page it lives on. When it comes to visuals, though, support is uneven. Some assistants, like Claude, don't generate photos or illustrations the way image generation tools do. Others, like ChatGPT and Codex, include image generation, but it's tied to one provider's model. Video, lip sync, and talking avatars are usually out of reach.
That gap creates a familiar routine. You write a prompt in your agent, paste it into a separate image or video app, wait, download the file, and bring it back into your project. Every revision repeats the loop. The alternative many developers reach for is an MCP server per model, which means a separate provider API key and config entry for each one.
The Model Context Protocol (MCP) is the standard way to give an agent new tools, and it's what lets Fish Audio remove those extra steps and open up a choice of models. Your agent calls the generation tools directly, the results come back as links, and the next step, whether a revision, an animation, or a voiceover, happens in the same conversation. Image generation in Claude Code, Cursor, or any other supported agent becomes one more tool call.
Getting there takes a single command.
How to Set Up the Image Generation MCP
The server lives at https://api.fish.audio/mcp and uses streamable HTTP. Sign-in happens in your browser with your Fish Audio account — no API key, no environment variables. Pick your agent below.
Looking for the Docs MCP? Our earlier guide to llms.txt, MCP, and Agent Skills covers
docs.fish.audio/mcp, a read-only server that lets coding agents look up Fish Audio API documentation. The server in this post is different: it signs into your account and creates content. You can run both side by side.
Claude Code
claude mcp add --transport http fish-audio https://api.fish.audio/mcp
Then run /mcp inside Claude Code to sign in.
Claude.ai
Go to Settings → Connectors → Add custom connector and enter https://api.fish.audio/mcp. Claude walks you through the sign-in.
Cursor
Open the command palette (Cmd/Ctrl+Shift+P) → Open MCP settings → Add custom MCP, and add:
{
"mcpServers": {
"fish-audio": { "url": "https://api.fish.audio/mcp" }
}
}
The cursor prompts you to sign in on first use.
Codex CLI
codex mcp add fish-audio https://api.fish.audio/mcp
codex mcp login fish-audio
The login command opens a browser window. Check the connection with codex mcp get fish-audio.
Windsurf
Go to Settings → Cascade → MCP Servers → View raw config (~/.codeium/windsurf/mcp_config.json) and add:
{
"mcpServers": {
"fish-audio": { "url": "https://api.fish.audio/mcp" }
}
}
Devin CLI
devin mcp add fish-audio --transport http --scope user --url https://api.fish.audio/mcp --scopes mcp
devin mcp login fish-audio --scopes mcp
Confirm it's connected with the devin mcp list. Because Devin works in your shell, it can also save the results locally — in our test, it downloaded a finished video to the Downloads folder on request.
Sign In and Choose Your Team
Every client opens the same Fish Audio sign-in page. You approve the connection once and choose which team it works in; the client renews access automatically after that.
A connection is tied to the team you choose. It can only see that team's workspaces and use that team's credits. To work in a different team, add the server again and choose that team at sign-in.
Connected? Try your first prompt: "Create a 16:9 watercolor image of a lighthouse at dawn. Show me the price first, then generate it." Start free with Fish Audio →
How Each Generation Works: Quote, Confirm, Generate
Once connected, your agent can run paid models on your behalf. That raises a fair question: what stops it from spending more than you meant to? The answer is built into the server, not left to the agent's judgment.
- Discover. The agent lists available models and reads the one it needs — supported options, defaults, and input limits.
- Quote. Before any image, video, lip sync, or avatar job, the agent requests an estimate. Estimates are free. Plan limits are checked here too, so if a request needs a higher plan, you find out before you're asked to approve anything.
- Confirm. The agent shows you the price and waits. Nothing is charged until you approve.
- Generate. The agent submits the exact request you approved. The server prices it again at submission; if the price has changed, the job is rejected and the agent must quote again. The charge never exceeds the amount you approved.
- Deliver. The agent checks progress every few seconds and returns a link to the finished file, along with what was charged.
Voice tools are billed differently. Text-to-speech and speech-to-text don't go through the quote step. They're charged as soon as they run, with text-to-speech billed by the length of the text.
Two more safeguards run in the background:
- Failed jobs are refunded automatically. If a generation fails, the credits come back without you asking.
- Retrying the same request won't charge twice. Every paid request carries a unique key. If the agent resends the same request with the same key after a timeout, the server returns to the original job instead of starting and charging for a new one.
Many agents add their own layer on top: Codex, for example, asks for your permission before each tool call.
Under the Hood: The MCP Media Tools
For developers, these are the media tools the agent calls behind the scenes:
| Tool | Charges credits? | What it does |
|---|---|---|
| list_my_workspaces | No | Lists workspaces in the connected team |
| list_media_models / get_media_model | No | Lists models and reads one model's options and input limits |
| create_media_upload | No | Creates a one-hour upload slot for an image, video, or audio file |
| estimate_image_generation / estimate_image_edit / estimate_video_generation | No | Validates a request and returns a quote |
| generate_image / generate_image_edit / generate_video | Yes | Submits the approved job |
| get_generation_status / get_generation_result | No | Tracks progress and returns the finished files and billing |
Lip sync and talking avatars run through generate_video. It recognizes the job from its inputs: a video plus audio means lip sync, and a portrait plus audio means an avatar. Audio tools such as text_to_speech, speech_to_text, search_voices, create_voice_clone, and the Story Studio tools sit in the same server.
You never have to call any of these tools yourself, though. In practice, you describe the result you want, and the agent chains the right calls together.
Example Prompts by Use Case
Here are a few starting points, grouped by what you're making:
| Use case | What you say | What the agent does |
|---|---|---|
| Product and app assets | "Generate a 1:1 image of a ceramic coffee mug on a marble counter in soft morning light. Give me a few variations and show the price first." | Picks an image model, quotes the batch, generates after you confirm |
| Product and app assets | "Upload screenshot.png and replace the background with a clean gradient for our app store listing. Estimate first." | Uploads your file, quotes the edit, applies it after you confirm |
| Marketing and social video | "Make a 9:16, five-second video of sneakers splashing through a puddle in slow motion. Quote the credits, then generate." | Quotes a text-to-video job and returns a video link |
| Marketing and social video | "Turn the image you just generated into a short video with a slow push-in." | Reuses the image directly as the starting frame |
| Updating and localizing video | "Translate this script into Spanish, voice it with my own voice clone, then lip sync demo.mp4 to the Spanish audio with Sync LipSync v2 Pro. Quote before generating." | Translates the script, generates the speech, uploads the video, and quotes the lip sync job |
| Presenters and avatars | "Generate a portrait of a friendly presenter in a bright studio, then have her read release-notes.md as a 9:16 talking-avatar video with Creatify Aurora Avatar." | Creates the portrait and the speech, then quotes the avatar video |
| Account | "How many Fish Audio credits do I have left?" | Checks your balance |
In coding agents like Claude Code, Codex, and Devin, these tools work alongside the shell. The agent can batch-generate images for every product in a folder, upload local recordings, or download finished files into your project.
Plans, Credits, and Limits
Before you try these, it helps to know how credits work and which plan each capability needs.
Which Credits It Uses
Media generation through MCP uses your Fish Audio web plan credits, the same credits you spend on the web app. It does not use Developer API credits, which are a separate balance for the Fish Audio API. The two are purchased in different places, so it's worth knowing which one to top up.
To add web credits or change plans, go to Team → Plans. Since each MCP connection is tied to the team you chose at sign-in, generations draw from that team's credits.
Which Plan You Need
What your agent can generate depends on your plan. Plan limits are checked when the agent requests a quote, so a job your plan doesn't cover is flagged before you're ever asked to approve it.
- Image generation, editing, and upscaling are available on the Free plan, which includes a daily image allowance.
- When generating images from text, each text-to-image request can ask for up to 2 images on Free, 4 on Plus, and 8 on Pro and Max. Image edits return one image per request.
- Video generation, lip sync, and talking avatars require Paid Plus, Pro, or Max plan.
- Video, lip sync, and avatar jobs return one output per request.
If you're starting on Free, image generation is the easiest way to try the workflow. When you're ready for video, lip sync, or talking avatars, upgrading to Paid Plus plan or higher unlocks them.
Prices
Every model is priced differently, and the cost of a job also depends on the settings you choose, such as resolution, duration, or the number of images. Rather than relying on a price list, ask your agent for a quote: it's free, and it reflects the current price. However, note that text-to-speech is the exception: it's billed by the length of the text and charged as soon as it runs.
Because quotes cost nothing, you can compare a few models or settings for the same idea and pick the one that fits your budget before the job is charged.
Uploads
Some jobs start with your own files: an image to edit or animate, a video to lip sync, or audio for an avatar. The agent handles the upload for you, within these limits:
- Images: JPEG, PNG, or WebP, up to 20 MB
- Video: MP4, MOV, WebM, or MKV, up to 1 GB
- Audio: MP3, WAV, M4A, AAC, OGG, or FLAC, up to 50 MB
Individual models can have tighter limits. For example, Creatify Aurora Avatar accepts 1 to 60 seconds of audio. The agent reads each model's limits before quoting, and if a provider rejects input after submission, the failed job is refunded. Results you've already generated in the same workspace don't need to be uploaded again; the agent passes them along directly.
Timing
Images usually finish within a minute. Videos, lip sync, and avatars can take several minutes. You don't need to keep checking: the agent tracks progress on its own and returns a link as soon as the file is ready.
Final Thoughts: Your Agent Is Now a Media Studio
AI agents can already plan a launch, draft the landing page, and ship the code. Until now, visuals were where most of them stopped: a single built-in image model at best, with video, lip sync, and talking avatars handed back to a person and a separate tool. This update closes that gap. One image generation MCP gives your agent voice, images, video, lip sync, and talking avatars, with no API keys to manage and a quote before every media job.
The bigger shift is how these pieces connect. A script becomes a voice, the voice becomes a talking avatar, and an image becomes a video, with each step handing its result to the next instead of files moving back and forth between apps. Voice is what makes a video feel like someone is actually speaking to you, and it sits at the core of what Fish Audio builds. Putting it in the same conversation as images and video means your agent can carry an idea from a few lines of text to a finished clip.
As new models are added, they appear in your agent's live model list, with nothing to reconfigure. The easiest way to see what that unlocks is to connect the server and ask for something you'd normally open three different tools to make.
New to Fish Audio? Sign up free → Connect the MCP server in under a minute and start with image generation, included in the free plan.
Ready for video, lip sync, and avatars? See plans → Plus and above unlock video generation, lip sync, and talking avatars through MCP.
Frequently Asked Questions
Can Claude Code generate images?
Is there a video generation MCP for Claude Code and Cursor?
Do I need an API key to use the Fish Audio MCP server?
Does it use my Developer API credits?
Can the agent spend credits without asking me?
What happens if a generation fails?
Can I use audio generated with Fish Audio text-to-speech for lip sync or avatars?
Which AI agents work with the Fish Audio MCP server?
Sabrina is part of Fish Audio's support and marketing team, helping users get the most out of AI voice products while turning launches, updates, and customer insights into clear, practical content.
Read more from Sabrina Shu