Tool overview
Fish is listed under Audio AI tools.
What is Fish?
Fish Audio offers expressive text-to-speech, voice cloning, voice design, speech-to-text, real-time streaming, creator tools, and developer APIs. Its S2-Pro model supports more than 80 TTS languages, while transcription supports more than 100 languages and dialects.
Best for
Creators and developers building expressive multilingual audio products
Who is it for?
Decision note
Strong for creators and developers needing expressive multilingual voice, transparent API pricing, and both hosted and self-hosted model options.
Key features
Expressive multilingual S2-Pro text-to-speech
Voice cloning and prompt-based voice design
Speech-to-text with speaker and emotion tags
REST, WebSocket, Python, and JavaScript APIs
Use cases
Create narration and character voices
Add speech to conversational agents
Transcribe podcasts and meetings
Build multilingual audio applications
Pros
- Includes an ongoing free creator tier
- API uses transparent pay-as-you-go pricing
- Publishes source and weights for Fish Speech
Limitations
Voice cloning and generated speech require clear authorization, disclosure, and misuse controls. Language quality varies by model and voice, and commercial self-hosting of Fish Speech requires a separate written license.
Pricing details
Billing options
Pricing note
The creator Free tier costs $0. Plus is $15 monthly or $11 per month when billed annually. The developer API has no subscription minimum and charges $15 per million UTF-8 bytes for TTS and $0.36 per audio hour for transcription.
Supported languages
- 80+ TTS languages
- 100+ transcription languages and dialects
Integrations
REST API
WebSocket API
Python SDK
JavaScript SDK
OpenAPI schema
Please log in to join the discussion.