Natural Japanese TTS wasn't enough for Pictoria's AI characters
For Pictoria, the voice is part of the character itself. A model can produce accurate, natural Japanese and still fall short once a character needs to sound excited, emotional, playful, or distinct.
The challenge becomes more apparent in longer-form content. Small differences in pitch, pacing, and emotional delivery can change how a character comes across, making consistency just as important as naturalness.
Pictoria needed a way to give each character a voice that could hold its personality across different interactions and content.
Pictoria built Picto VOICE to choose the right model for every character
Rather than rely on a single TTS provider, Pictoria built Picto VOICE as a multi-model voice system. The team could evaluate different models and match each one to the requirements of a particular character or production.
That approach raised the bar for every model in the stack. Fish Audio needed to offer something meaningfully different, not simply another capable Japanese TTS option.
Fish Audio stood out where character expression was a priority
Pictoria evaluated Fish Audio alongside other voice generation technologies and found three areas where it stood out: expressive character voices, Japanese prosody and intonation, and flexibility in character voice design.
The difference was particularly apparent in how a voice carried emotion across a single utterance. Instead of simply reading a line, the model could give a character a distinct emotional register and personality.
Fish Audio also worked closely with Pictoria's team on the specific demands of Japanese AI character voice, turning the integration into an ongoing collaboration rather than a simple vendor relationship.
Fish Audio is consistently pushing the frontier of the high-quality voice generation expected in the Japanese market. We believe voices with this level of emotional nuance and character expression can only be created by a team that genuinely loves character-driven content.

Hayato Akedo
CEO, Pictoria
Fish Audio powers the characters where expression matters most
Within Picto VOICE, Pictoria uses Fish Audio for AI character use cases where character-level expression is the deciding factor. One key application is the voice production pipeline for Pictoria's original AI character IP, developed and operated through its own publishing channels.
The integration has delivered positive reception across these projects, with users consistently responding to three qualities: natural Japanese voices, stronger emotional expression, and consistency across long-form character content.
The result: stronger character voices, from production to audience
After integrating Fish Audio into Picto VOICE, Pictoria saw meaningful improvements in Japanese response speed and emotional expressiveness across its AI character projects. The voices also received strong positive feedback across Pictoria’s original character IP, particularly for sounding natural, carrying emotion, and staying consistent across hours of content.
That gives Pictoria a practical advantage in production: its team can match each character with the voice technology that best fits the experience, while using Fish Audio when expressive Japanese performance is the priority.