Rankings
How the World Class 100 works, and why the votes come in pairs
The World Class 100 ranks invented characters by pairwise votes rather than stars. The full method: Elo, Glicko, scheduled recomputes, and what is not live yet.
7 min read

A ranked list is worth reading only if you can see how it was made, and almost none of them show you. So: the World Class 100 is our ordered list of the characters in Profiles, all of them invented by us and every image generated, ranked by readers choosing between two characters at a time rather than awarding marks out of ten. Results are rated against each other rather than counted in isolation, and the tables are recomputed on a schedule, so nothing climbs because somebody sat refreshing it at three in the morning. Reader voting is not open yet. The section near the end says what stands in its place. Everything between here and there is the method.
What the World Class 100 is
A list of a hundred fictional people, in order, decided by readers.
The fiction is the first fact rather than the small print at the bottom. Every character is written by us, every face is generated, and no character depicts, is modelled on or is named after a real identifiable person. That is also the most interesting thing about the format. A hundred-name list of who is most beautiful is a familiar and slightly grubby genre; this is the only version of it with nobody on the receiving end. No one is judged, nothing is sold, and there is no one to contact because there is no one there.
Which is why the list stops where it stops. We do not rank real people by appearance, and we do not rank businesses, and the editorial standards say so in writing rather than leaving it to good intentions. There are national and regional cuts as well as the global hundred, and the same method runs all of them.
Why a score out of ten cannot do this
Every other site does it with stars, and stars collapse.
The five and seven point scale comes from Rensis Likert, in 1932, and its failure modes have been documented for almost as long as it has existed. Respondents cluster towards one end. They avoid the extremes, so a scale of ten is really a scale of four. Averages drift upward over time, because people who bother to rate something at all are usually people who liked it. And nothing anchors the scale: one reader’s 7 is another reader’s 9, and neither was ever told what a 7 is.
Then there is the arithmetic of a hundred. A hundred items have 4,950 distinct pairs, which is why nobody is asked to rank a hundred things directly. Told to place character forty-one against the other ninety-nine in one ordering, a reader manages the top few, the bottom few, and guesses the middle.
A comparison needs a reference point. A lone number does not have one. A second character does.
Why a vote between two
“Which of these two” is a question a person can answer in a second and be honest about, and it is the oldest trick in comparative measurement rather than anything we came up with.
The formal version is the Bradley-Terry model, published in 1952 by Ralph Bradley and Milton Terry: in a contest between two, the odds that one beats the other are the ratio of their two ability parameters. Ernst Zermelo had anticipated the same idea in the late 1920s, and it has been rediscovered independently since, which usually means an idea was the obvious one rather than the clever one.
It also explains why asking a crowd works. In 1907 Francis Galton published “Vox Populi” in Nature, having collected 787 usable estimates of the weight of an ox at a livestock show. The animal weighed 1,198 lb. The median of the guesses was 1,207 lb, and the mean of the same set 1,197 lb. That is not a demonstration that crowds are wise about everything; it has been hauled out to prove far more than it can carry. It shows the narrower thing we need: independent judgements, aggregated, can be startlingly good. Independent is the load-bearing word, and we will come back to it.
How a rating moves
Beating something highly rated should be worth more than beating something nobody rates highly. A tally of votes cannot tell the difference. A rating system is built on it.
The principle comes from Arpad Elo, a physics professor at Marquette University, whose system the United States Chess Federation adopted in 1960 and FIDE in 1970. A rating there is not a trophy cabinet but a prediction: what is likely to happen when this one meets that one. Each result then moves both numbers by an amount that depends on how surprising the result was. A heavy favourite that wins as expected gains almost nothing, and the character it beat loses almost nothing. An upset moves both, and moves them a long way.
That is why the list cannot be farmed by volume. A character put repeatedly against weak opposition converges on the rating that opposition justifies and then stops climbing, no matter how many pairs it wins.
It also handles the awkward thing about pairwise preference, which Condorcet described in 1785: majority preferences between pairs can cycle. A beats B, B beats C, C beats A, and there is no consistent order to be had. This is a real property of the method, not a bug we will pretend away. Ratings survive it because they score rather than order absolutely. A cycle leaves three characters sitting at very similar numbers, which is an accurate description of three characters readers cannot decide between.
Why the tables are recomputed on a schedule
Refreshing is not voting.
Votes are gathered continuously and the tables rebuilt in batches, not the instant one lands. The interval will be published with the table it produced. Three things follow. A batch can be inspected before it is published, so a coordinated push shows up as a shape in one recompute rather than disappearing into a slow live drift. Reversing a bad batch costs one recompute. And a table that was computed has a date on it, which an always-live counter can never have: live means the number is unrepeatable and nobody, including us, can say what it read last Tuesday.
This is also where Galton’s word comes back. Aggregation works when judgements are independent. Organised voting is precisely the failure of independence, and batching is the cheapest defence against it that does not involve us deciding which votes we liked.
Where the list cannot be trusted
The honest section, and shorter than we would like.
Coverage is uneven while the character set is small. Some pairs occur often and some hardly at all, and a rating is only as good as the pairs that happened.
Early ratings are unreliable, full stop. A character that has been in a handful of comparisons has a number, and that number means very little. The proper answer to this is a measure of its own: Mark Glickman’s Glicko system, devised in 1995 as an improvement on Elo, adds a ratings deviation, a stated measure of how reliable a rating currently is. A new or rarely compared entrant is marked uncertain rather than dumped in mid-table with the same confidence as everything around it. That is the treatment thin coverage deserves, and the direction we are building in. The threshold for marking one provisional will be published with the method that runs it.
The regional and national cuts are thinner by construction: they divide the same votes among smaller fields. Read them as weaker claims than the global hundred.
What is not live yet
Reader voting is not open, and the Profiles section is not yet populated.
This page describes the method the World Class 100 will be built with. What stands in its place in the meantime is an editorial ordering: made by us, against a written rubric, exactly as the City Index scores are made and labelled as an editorial judgement wherever it appears. It is not survey data and it is not a vote count, and nothing on the site will call it either.
When voting opens, it opens on the method above, and the tables will say which of the two they came from.
Reading an ordered list
Publishing the method before the list is the wrong way round commercially and the right way round otherwise. A list arrives, it gets screenshotted, and the method never gets read. Here the method is the page.
When the hundred arrives, you will know what a high number means: not that a character was liked a lot, but that she kept winning against opposition that was itself winning. The other tables in Rankings work the same way, on cities instead of characters, and the vocabulary for all of it is in the Encyclopedia. Argue with the order all you like. That is what it is for.
Frequently asked questions
- Who is ranked in the World Class 100?
- Invented characters, and nobody else. Every character in Profiles is written by us, every image is generated, and none of them depicts a real identifiable person or is offered or available for anything. We do not rank real people by appearance and we do not rank businesses, which is set out in our editorial standards and is not negotiable.
- Why pairwise votes instead of a score out of ten?
- Because a lone number has nothing to anchor it. Rating scales of the kind Rensis Likert introduced in 1932 cluster towards one end, drift upward over time, and mean different things to different people, so one reader's 7 is another reader's 9. A choice between two characters asks a question a person can answer without calibrating anything, and the comparison carries its own reference point.
- Can readers vote in the World Class 100 today?
- No. Reader voting is not open yet and the Profiles section is not populated. What stands in its place is an editorial ordering made by us against a written rubric, the same way the City Index scores are made, and it is labelled as editorial wherever it appears. When voting opens, the method on this page is the method that will run.
- Why are the tables recomputed on a schedule rather than live?
- Because a live counter rewards persistence rather than preference, and it cannot be dated. A batch recompute can be inspected before it is published, which makes a coordinated voting campaign visible as a shape in one batch rather than invisible as a slow drift, and it means every table carries the date it was computed on.