chataiopensource

Blog

Notes from the arena

How the ratings work, what changed, and the reasoning behind it. Short posts, written when the reasoning is still fresh.

Both responses were bad, so neither gets credit

A vote that can only say "A" or "B" quietly rewards confident nonsense. We added a third answer, and it excludes the battle from the ratings entirely.

For a long time the only two options were "A is better" and "B is better". That looks neutral, but it is not. Forced into a choice between a good answer and a useless one, people pick the good one — and the useless one absorbs a loss it deserved. The rating system never learns that the model refused, looped, or invented a citation.

So there is now a third button: both are bad. Selecting it records the battle, excludes it from the ratings, and stores nothing that could move a number. No winner, no loser, no Elo change.

The important part is the exclusion. If a "both bad" battle still moved the ratings, it would be strictly worse than leaving the option out: you would be averaging a refusal into a model’s score instead of treating it as missing data.

The follow-up question is the obvious one — do people actually use it? The answer is yes, and mostly on the cases you would hope: refusals, answers that ignore the actual question, and responses that stop mid-sentence. Without the button those battles were being recorded as wins for whatever the other model happened to say.

Both models answer under the same style, on purpose

The composer lets you set a temperature and a system prompt. Both sides of a battle always get identical settings — otherwise you are grading prompt engineering, not models.

The composer now has a Style control: a preset such as Concise or Explain Like I’m Five, an optional instruction of your own, and a temperature dial. It is a small feature with one large constraint attached to it.

In a battle, both models receive the same preset, the same instruction, and the same temperature. Not similar — identical. The reason is that the moment the two sides differ in any setting other than the model itself, the vote stops measuring the models.

Imagine side A gets "be brief" and side B gets nothing. Side A will often produce the more useful answer for a chatty question, and the vote will record that as a property of the models. It is not. It is a property of the setup, and it will not reproduce.

The settings are stored alongside the vote so a result can be reconstructed later, but they never touch the rating itself. Style and score are separate columns for a reason: the day style starts moving numbers, the leaderboard stops answering the question it claims to answer.

Temperature is worth setting explicitly when it matters. For code, low is almost always better — you want the answer the model considers most likely rather than one of several plausible ones. For brainstorming, the opposite.

What the numbers on the leaderboard actually mean

Elo, starting ratings, why early numbers are noisy, and how to read a rating without over-reading it.

Every model starts at 1000. After a vote, the winner gains points and the loser loses them, and the size of the change depends on the gap beforehand. Beating a much stronger model is worth more than beating a weak one, which is the property that makes the system self-correcting rather than just cumulative.

Two things follow from that, and both are worth internalising before you trust a ranking.

First, early numbers are noise. A model with three votes has a rating, but not a measurement. The number converges as volume arrives, and a small lead between two rarely-battle models means almost nothing. A large gap between two models with thousands of votes between them means a great deal.

Second, the rating is a summary of one question: "given these two answers, which did this person prefer?" That is a real and useful signal, but it is narrower than "which model is best". It says nothing about cost, latency, context length, or whether the model can hold a tool call. Treat the leaderboard as one input to a decision, not the decision.

The rating is also only as honest as the votes behind it, and anyone can vote. Anonymity of the models is the main thing limiting the damage — you cannot vote for a brand, only for a text — but volume is the real protection.

The name is the bias

Why identities stay hidden until after the vote, and what that buys and what it costs.

If you know you are reading a model from a company you admire, part of your vote is your prior opinion of that company. The effect is large, it is well documented, and it is invisible from the inside — which is what makes it dangerous.

So the two models stay anonymous until you have voted. Reveal is a separate step, and it happens after the number is written.

What this buys is a rating that reflects the answer rather than the reputation. What it costs is that you cannot use brand familiarity as a tiebreaker, and the first few votes on a brand-new model are slow while people build an impression from scratch. We think that is the right trade, but it is a trade.

It is also why the unranked modes exist. Side-by-Side and Direct Chat show you the model names the whole way through, because the point of those modes is to work with a model you have chosen rather than to produce an unbiased comparison. Same models, different question, and the leaderboard only ever hears about the first one.

Self-hosting the arena in one file and a database

No build step, no package manager, no framework. What that buys, and what it costs.

The application is plain PHP with a SQLite file behind it. You can read every line that decides how a vote is counted, which is the property that matters most for something whose output is a number people make decisions from.

The cost is real. There is no ORM, no test framework in the usual sense, and no bundler, so the front end is hand-written JavaScript and the CSS is one large file. If you are used to a modern toolchain it will feel uncluttered in a way that is initially pleasant and occasionally annoying.

The driver is a genuine constraint rather than a preference. The host this was built on has the sqlite3 binary but not the pdo_sqlite extension, so a small PDO-shaped shim talks to the CLI. It is about two hundred lines, it implements only what the app actually calls, and it works.

Upgrading the schema is one command: php migrate.php. It applies only what has not run yet, records what it applied, and is safe to run twice. Back up the database first, as with any database change.