Reverie AI builds entertainment that responds to the user
Traditional entertainment follows a fixed timeline. The script is written, the performance is recorded, and every viewer sees the same sequence.
Reverie AI is building something different. Its interaction models can take in text, voice, video, touch, gestures, choices, and on-screen actions, then determine what happens next. A user can talk to a character, influence a scene, change an ending, or open a branch that did not exist a moment earlier.
That means the experience is no longer fully defined before the user enters it. Characters and scenes have to respond to what happens in the moment.
A responsive world cannot wait for the voiceover
In a fixed piece of content, the voice can be recorded after the script is finished. In an interactive world, the next line may not exist until the user does something.
Reverie needed voice that could begin before the full response had finished generating, follow direction within a line, and maintain the same character across scenes and branches. It also needed the economics to work when every interaction could generate new speech.
The goal was not simply to make a generated voice sound human. It was to make voice behave like part of a live world.
Reverie AI wanted control over the character, not a finished voice experience
Reverie evaluated voice providers against four requirements: fast streaming, expressive control, persistent character identity, and multilingual support. The team needed to direct a character in the moment rather than choose a single mood for an entire response.
Fish Audio met those requirements in one voice layer. Speech streams over WebSocket as it generates, expression can be directed inline with natural language, and reusable voice models allow characters to maintain a consistent vocal identity across an experience.
Fish Audio lets characters respond in the moment
With Fish Audio, Reverie can start streaming speech before the full response has finished generating. That keeps the gap between a user's action and a character's reaction short enough for the interaction to feel continuous.
The team can direct individual moments inside a line, using cues for expressions such as whispering, laughing, or excitement. Rather than applying one voice setting to an entire scene, Reverie can shape the performance around what is happening.
Reusable voice models keep those characters recognizable as the experience moves between scenes, sessions, and branches.
One voice layer for worlds that can go anywhere
Reverie's experiences are designed to reach global audiences, so voice needs to work across languages without creating a separate production pipeline for every market.
Fish Audio supports 83 languages on its current production model, allowing Reverie to carry its characters into new markets while keeping the voice layer within the same architecture.
The economics matter too. Fish Audio's pay-as-you-go API is priced at $15 per million UTF-8 bytes, giving Reverie voice infrastructure that can scale with usage as more interactions generate speech.
Voice becomes part of the world
For Reverie, the shift is from voice as playback to voice as participation.
Characters can react in the moment instead of delivering pre-recorded lines. Their voices can change with the scene while maintaining a recognizable identity. The same voice infrastructure can support the different formats Reverie is building, from roleplay companions and virtual pets to interactive adventures and games.
Fish Audio gave Reverie the voice layer to make those experiences responsive without turning voice production into a separate pipeline.