How Rate My Username's scoring actually works
Behind the number: the rubric, the consistency problem, and what I'd redesign first.
The question I get the most about Rate My Username isn't "is it accurate" — people already know it's not meant to be a scientific instrument. It's "how does it actually decide the number." Fair question, and the honest answer took longer to land on than I expected.
Starting with four axes, not one
The first version of the tool returned a single score with a single joke attached, and it felt flat almost immediately. A username that's funny but generic and a username that's original but forgettable both landed on similar numbers, which made the result feel random instead of considered. Splitting the score into four axes — vibe, originality, aura, and memorability — fixed that, because now two very different usernames could score the same overall total for genuinely different reasons, and the breakdown showed why instead of just asserting a verdict.
The actual problem: consistency, not cleverness
Writing a prompt that produces one funny verdict is easy. Writing one that produces a funny verdict for the thousandth username, in the same tone, without repeating the same three jokes, while staying roughly calibrated against every score that came before it, is a completely different problem. Early on the same username could score a 54 one run and a 79 the next, which is unacceptable for anything people are going to compare against a friend's result or a leaderboard. Most of the actual engineering time on this feature went into tightening that variance, not into making any individual joke funnier.
What narrowed the variance
Three things helped more than anything else. First, giving the model explicit numeric anchors for each axis — concrete examples of what a 20, a 50, and a 90 look like — instead of leaving "originality" undefined and hoping for consistency. Second, separating the scoring step from the verdict-writing step, so the number gets decided before the joke gets written, instead of asking for both at once and letting the funniest-sounding line drag the score around with it. Third, running repeat tests on the same handful of usernames across sessions and logging how much the total moved, which turned "it felt more consistent" into something I could actually measure and compare between prompt versions.
What I'd change next
The rubric still leans harder on wordplay and rhythm than on cultural reference recognition, mostly because reference-spotting is the least reliable part of the whole system. A future pass will probably split "originality" into two separate signals instead of one blended axis, since right now a username can earn originality points for two very different reasons that currently get averaged into one number instead of shown separately.