What a 1–10 match score should actually mean
A number is only useful if it can be low. Most AI scores can't be.
24 July 2026
Hand a language model a job description and a resume, ask it for a match score out of ten, and you will get a 7. Try a wildly unsuitable job and you'll get a 6. Try a genuinely excellent one and you'll get an 8. The scale has collapsed into a narrow band around "seems fine", and every downstream decision built on it is now noise.
This is the single hardest part of AI job matching, and it has almost nothing to do with model quality.
Why everything comes back a 6 or 7
There's no anchor. "Rate this out of ten" doesn't say ten of what, compared to what. Absent a reference point, the safest answer is near the middle — and the model is optimised to be safe.
Job ads are written to sound appealing. They're marketing documents. A model reading one in isolation is reading the best possible case for the role, which pushes every score up.
Each job is judged alone. A human ranking twenty postings does it by comparison — this one's better than that one. Score each independently and you lose the comparison that made the ranking meaningful.
Hedging is rewarded. Confidently scoring a plausible-looking job a 3 is a strong claim. Models trained toward helpfulness and away from overconfidence drift toward the middle whenever they're uncertain, and reading a job ad is uncertain by construction.
Why a collapsed scale is worse than no scale
If everything scores 6 to 8, the number carries almost no information but looks like it does. You start trusting a 7 over your own reading, then discover the 7s include roles demanding twice your experience. The number's precision is fake, and fake precision is more expensive than an honest "here are twenty jobs, go look".
The test for any scoring system is simple: can it return nothing? A matcher that never has a quiet day has stopped discriminating and started padding.
What fixes it
Anchor the scale to the person's field. A scale calibrated for backend engineers will rate every sales role a 2, and vice versa. Before scoring anything, work out what field this person is actually in from their resume, then build the rubric around it — with worked examples of what a 9 looks like, what a 5 looks like, and what a 2 looks like for them. Concrete anchors are what stop the drift to the middle.
Take hard constraints away from the model entirely. Geography, experience bounds, blocklists — these aren't judgment calls. Enforce them in plain logic before the model is consulted. Otherwise a sufficiently exciting job description will talk a model into deciding that your "remote only" was more of a preference.
Score against a full profile, not a keyword bag. Roles, skills in priority order, experience, locations, work mode, salary floor, hard rules. The richer the picture, the less room there is for a generic "seems relevant" answer.
Make it justify the number. Requiring reasons and red flags alongside the score does two things: it constrains the score toward something defensible, and it gives you the evidence to overrule it. A score you can't interrogate is a score you have to obey.
Then set a threshold and honour it. Ours is 7. Eight and above is flagged urgent, 7 is a strong match, below 7 is never sent, and there's a daily ceiling so a busy Tuesday can't turn into a feed. Some mornings that means nothing arrives — which is the correct answer when nothing good was posted, and the main reason we trust the number at all.
What it still gets wrong
It's a language model reading a job advert, and job ads lie by omission. The salary band that isn't there, the "fast-paced environment" doing heavy lifting, the team that has hired for this role three times this year — none of that is in the text, so none of it is in the score.
That's why every alert we send shows its reasoning and its red flags rather than just a number. You're meant to sanity-check it in five seconds, not obey it. A matcher that asks for trust it hasn't earned is doing the same thing as a collapsed scale: dressing up uncertainty as precision.
Related: Why keyword job alerts fail · what the model is actually asked · the FAQ, including what happens when a score is wrong.
See what it scores your matches
Four minutes to set up. Free during the beta, one click to delete.