FAQ
Frequently asked questions
Everything below is the short version. If something is still unclear, the how it works page has the detail.
Basics
What is this?
A public, anonymous way to compare language models. You send one prompt, two models answer it, you pick the better answer, and that vote feeds a public rating table. Neither model knows which one it is while you are deciding, and neither model knows it is competing.
Why are the models anonymous?
Because a name changes your judgement. If you knew you were reading GPT versus Llama, part of your vote would be your prior opinion of the brand. Identities are revealed only after you have voted, so the rating reflects the answer you actually read.
Do I need an account?
No. There is no sign-up and no login. Battles are anonymous from the start, and votes are not tied to an identity.
Is it really free and open source?
Yes to both. The application is open source, and the ratings it produces are public. If you would rather run your own copy against your own models, the whole thing is a handful of PHP files and a SQLite file.
Battles and voting
What can I do with a battle?
Keep talking to both models within the same battle. Each side keeps its own conversation history, so side A never sees what side B said. When you are satisfied, vote on the battle as a whole.
Why are there four vote buttons?
A, tie, and B are the normal answers. "Both are bad" is the important one: a model that refuses, loops, ignores the prompt or fabricates sources is worse than useless, and a rating system that cannot express that will quietly reward confident nonsense. Choosing "both are bad" excludes the battle from the ratings entirely.
Can I change my vote?
No. A vote is recorded once, when you press the button, and the models are revealed immediately afterwards. That is deliberate - a vote you can take back until the answer suits you is not a measurement.
What is the difference between the other modes?
Side-by-Side lets you choose both models and shows you their names, but those battles do not count towards ratings. Direct Chat is a normal conversation with one model, also unrated. Only anonymous Battle mode is ranked, because that is the only setup where the comparison is unbiased.
Style controls
What does the Style button do?
It sets how the models answer, independently of what you ask them. You can pick a preset such as Concise or Explain Like I'm Five, add your own instruction, and set the temperature.
Do both models get the same style?
Yes, always. Both sides of a battle run under identical settings, otherwise you would be comparing a carefully-written answer against a careless one and calling it a model difference.
What does temperature actually change?
How much randomness the model is allowed when picking each next word. Low temperature gives focused, repeatable answers. High temperature gives looser, more inventive ones, and is more likely to wander off the point. It is worth setting explicitly when you care about consistency, such as for code.
Does the style affect my vote?
It is recorded alongside the battle so a result can be reproduced later, but it never changes how a vote is counted. Style and rating are kept separate.
Ratings
How are models rated?
With an Elo-style rating, the same system used in competitive chess. After each vote the winner gains points and the loser loses points, scaled by how close the two models were rated beforehand. A model that beats a much stronger model gains more than one that beats a weak one. Every model starts at 1000.
Why do the ratings keep moving?
Every vote nudges the ratings, so a model with a lot of traffic settles towards its true strength while a quiet model drifts. Early rankings are noisy. A handful of votes is an anecdote, not a measurement.
Does the "why did you pick that one" step matter?
It is optional and it never affects a rating. It exists so the dataset can record *why* a comparison went the way it did - more accurate, more helpful, cut off mid-sentence - which is what makes the votes useful for research rather than just for a ranking.
Can I game the ratings?
It is worth being honest about this: the ratings are only as good as the votes behind them, and anyone can vote. What mitigates it is volume plus the anonymity of the models, and the fact that a model that wins by being confidently wrong stops winning once people notice. Treat a leaderboard as a rough signal, not a verdict.
Running your own copy
What do I need to run it?
PHP 8 with cURL, and a SQLite database. There is no build step, no package manager and no framework. Load schema.sql into a fresh database, add API keys to .env, and serve the directory with any web server.
Do I need API keys?
Only for the models you actually want to call. Three local demo models ship with the app and need no key at all, which is enough to exercise the whole battle, vote and rating flow before you spend anything.
What happened to my data when I upgraded the schema?
Run php migrate.php. It applies only the migrations that have not run yet, records what it applied, and is safe to run repeatedly. Back up database.db first, as with any database change.