Fish Audio’s Best Prompts Just Got Searchable

Fish Audio is a solid text-to-speech and voice-cloning model, but figuring out how to actually prompt it has meant digging through scattered forum posts, half-finished GitHub threads, and guessing at emotion tags until something sticks. This Redditor, u/Creepy_Sun_7638, got tired of that and built a searchable Fish Audio prompt library with real audio samples attached to every entry, so you can hear exactly what a prompt produces before you copy a single word of it. That last part is the actual upgrade here. Most prompt libraries for AI tools are just text dumps: a list of strings with no way to know what they sound like until you paste them in yourself and burn a generation credit finding out. This one flips that. You listen first, copy second, which turns prompt selection from a guessing game into an actual comparison.

The twist

While testing prompts for the library, this contributor noticed something that goes against the usual advice: Fish Audio’s S2 emotion tags work better short than long. A tag like “angry, shouting” outperforms a full paragraph describing the emotional state in detail. Most people assume more description equals better control, the same logic that works for image models where a longer, denser prompt usually pulls more of what you want out of the model. Here, less wins. Stack too many descriptive words onto an emotion tag and the model seems to average them out into something flatter instead of sharper. A single sharp tag reads as one clear instruction; a paragraph reads as noise the model has to sort through. That is the kind of detail you only find by actually testing hundreds of prompts and comparing the output side by side, which is exactly what this library lets anyone do without doing the testing themselves. It is the difference between reading a tutorial that guesses at best practices and reading one built entirely from observed results.

How to use it

  • 🔍 Open the Fish Audio Prompt Library and filter by model, language, voice, or emotion, since narrowing by language alone saves you from wading through entries that will not match your accent or delivery needs
  • 🎧 Play the sample audio for a few entries before picking one, so you know what you’re actually getting rather than relying on the text description of the tone
  • 📋 Copy the exact prompt text, since the wording and structure matter more than people expect. Word order, punctuation, and even where the emotion tag sits in the string can change the output
  • 🎙️ Drop it into Fish Audio and adjust the emotion tag first if the result feels off, before touching anything else in the prompt. It is the single variable most likely to fix a flat or overacted read
  • 🛠️ Found a prompt that works great? Submit it through the project’s GitHub issues so the library keeps growing, and note what voice or model you tested it on so the next person can filter to match

Pro tip

Keep emotion tags to two or three words max. The moment you start writing full sentences for tone, the model seems to lose the signal instead of sharpening it. Treat emotion tags like keywords, not stage directions. If you are used to prompting text models where extra context almost always helps, this is the habit you have to unlearn here.

A second one worth stealing: use the filters before you write anything from scratch. If someone already found a prompt that nails the voice and emotion you’re after, there’s no reason to reinvent it. Search by the closest match first, tweak the emotion tag second, and only start from a blank prompt if nothing in the library gets you close. This alone will save more time than any amount of manual trial and error, since most of the combinations people actually need have probably already been tested by someone else in the library.

Worth noting, the creator flagged that mobile view is still a bit rough, so browsing on a laptop or desktop will be a smoother experience for now, and the whole thing is open source on GitHub if anyone wants to help polish it. This is early and clearly still growing, which is exactly why community submissions matter here. The library is only as good as what people add back to it, so every submitted prompt raises the value for the next person who searches it.

For anyone who has fought with Fish Audio’s prompting quirks and had no real reference point, this closes that gap. Go listen to a few samples, steal a prompt that fits your project, and if you land on something good, send it back so the next person doesn’t have to guess either.

I wanted to see how other people were actually prompting Fish Audio, so I made a prompt library for it
by u/Creepy_Sun_7638 in PromptEngineering

Scroll to Top