Japanese AI character voice for Picto VOICE.
How Pictoria integrates Fish Audio into Picto VOICE — their internal multi-vendor stack — for character-level expressiveness across Japanese AI character productions.
How Pictoria integrates Fish Audio into Picto VOICE — their internal multi-vendor stack — for character-level expressiveness across Japanese AI character productions.
Across hours of long-form character content — the zero-drift audio reliability required to maintain the suspension of disbelief for Japanese AI character IPs.
Hayato Akedo
CEO of Pictoria
"Fish Audio is consistently pushing the frontier of the high-quality voice generation expected in the Japanese market. We believe voices with this level of emotional nuance and character expression can only be created by a team that genuinely loves character-driven content. Thank you for building such an outstanding TTS technology."
Chief Scientist of Fish Audio
Shijia Liao

"Japanese AI characters are the use case that demands the most from a voice model. Pronunciation accuracy isn't enough. The voice has to carry character: the specific emotional register, the moment of pause, the warmth that makes someone feel like a person rather than sound like one. Pictoria sets that bar higher than almost any team we work with, and being part of Picto VOICE pushes our model further every release."

Pictoria is a Japanese AI character company building both AI character projects with voice talent and original AI character IPs developed in-house. At the center of every Pictoria production sits Picto VOICE — an internal voice system that combines multiple Japanese TTS technologies and selects the right one per character.
Voice isn't a feature of the AI character experience. For Pictoria, voice is the AI character experience.
Japanese AI character voice has a bar most TTS models can't cross. Naturalness and pronunciation accuracy are table stakes; the deciding factors are emotional expressiveness across a single utterance, accurate Japanese prosody and pitch-accent, and voice consistency across long-form character content.
Generic TTS readouts pass benchmark tests but break the suspension of disbelief that AI character experiences depend on. For Pictoria, the two persistent gaps in available Japanese TTS for AI characters were emotional expressiveness and natural Japanese prosody at the character level — both essential for the company's growing roster of original AI character IPs.

Rather than commit the entire character voice surface to a single provider, Pictoria built Picto VOICE as an internal orchestration layer that integrates multiple Japanese voice AI technologies. Each character and use case selects the right voice generation component from the stack.
This architecture protects Pictoria against vendor lock-in, lets the team evaluate new Japanese voice AI models continuously, and ensures the best technology — for that character, for that scene, for that emotional register — is always the one in production. Fish Audio earned a place in Picto VOICE alongside the other voice technologies the team integrates.
Pictoria evaluated Fish Audio as one of several voice generation technologies under consideration for Picto VOICE. Three things drew the team in:
Expressive range of character-style voices — particularly the ability to carry emotional nuance across a single utterance. The move that separates a character voice from a narrator.
Expressive Japanese TTS quality at the level of prosody and intonation, not just phoneme accuracy.
Flexibility in character voice design — the ability to shape voices that fit a specific character brief rather than picking from a fixed library.
What stood out beyond the model itself was the responsiveness of the Fish Audio team to Japanese AI character voice needs — the kind of partnership that turns a vendor relationship into a working collaboration with the research team behind the model.

Inside Picto VOICE, Fish Audio is used particularly for AI character use cases where strong character-level expressiveness is the deciding factor. A
representative deployment is Pictoria's AI character voice production pipeline for original Pictoria AI character IPs — the in-house IPs the company
develops and operates on its own publication channels.
Integrating Fish Audio into Picto VOICE produced strong impact and positive reception across these AI character projects.
User reception of the Japanese AI character voices produced through Picto VOICE has been positive across three consistent themes:
Products used: Fish Audio S2 Pro · Text-to-Speech
Meaningful improvements in Japanese response speed and emotional expressiveness within Picto VOICE.
Strong positive reception across original Pictoria AI character IPs.
Active collaboration on the Japanese character voice frontier; quantitative metrics held confidential per IP agreements.
Three things separate Japanese character TTS from generic Japanese TTS: emotional expressiveness across a single utterance, accurate Japanese prosody and pitch-accent rather than phoneme-only accuracy, and voice consistency across long-form character content. Pictoria's Picto VOICE selects voice technologies on these criteria for each character, with Fish Audio earning a place for character-level expressiveness.
Talk to our team about character voice, prosody control, and integration into multi-vendor stacks like Picto VOICE. We come prepared.