TTS Arena Docs

Voting

What makes a vote count, and the rules that keep the board fair.

Casting a vote

  1. Type a line, or hit Random for one from the prompt pool.
  2. Two anonymous models - A and B - synthesize it.
  3. Listen to both, then pick the one that sounds more human. One choice, no skips.
  4. The identities are revealed. If the line came unchanged from Random, the public ratings update.

You need to listen to enough of each clip before voting unlocks - this keeps votes grounded in the audio rather than reflexive clicks.

Requirements

  • Sign in with Hugging Face. Voting is tied to your account so each vote counts once. Accounts must be at least 30 days old.
  • English only, for now - it's the language all models support. Multilingual is on the roadmap.
  • Prompts are capped at 1,000 characters.
  • Don't vote on a model you're connected to. If you work for, contract for, or are otherwise affiliated with a provider, don't vote in battles that include their model. That applies to personal accounts too. Providers agree to this in the integrity policy; it binds their people as well.
  • Only clean votes on first-use Random prompts move the public leaderboard. Typed custom prompts are still useful for side-by-side listening, but they do not affect ratings.

Keeping it fair

Votes run through an anti-abuse system so the board reflects real preferences:

  • Behavioral signals score each vote for risk (timing, patterns, device and network signals). High-risk votes are recorded but shadow-excluded - they never move the public ratings.
  • A captcha appears once per browser, and again if risk rises.
  • A background sweep looks for coordinated rings - accounts sharing an IP or device that pile onto one model, whether they signed in together or merely voted together - and retroactively excludes their votes.
  • The sweep also runs a statistical test for recognition. For every voter and every model, it compares how often that voter picks the model against how often everyone else does, counting only the battles the model actually appeared in. Liking a model puts you a few points above the crowd. Picking it almost every time it appears, when the crowd is split, is not a preference - it means you can tell which one it is, and a blind test where you can't be blind isn't a vote. The comparison excludes the voter's own votes from the baseline, so nobody can define the norm they are measured against.

Whenever votes are excluded, the leaderboard is recomputed from the clean set, so a rating never keeps the lift it was given.

None of this affects honest voting - it's invisible unless your activity looks automated or coordinated. What we collect to do it, and why, is on the privacy page.

On this page