AI lead scoring has an obvious failure mode: run it without telling the model what you actually want, and everything comes back a six or a seven. The scores are technically responsive and completely useless, because a score that does not discriminate is not a score.
The fix is not a better model. It is a better description of what good looks like.
The ICP profile does the work
Scoring compares each contact against your ideal customer profile. If that profile is vague, the comparison is vague, and the output clusters in the middle.
A weak profile: "B2B SaaS companies, decision makers."
That describes an enormous population. Almost any contact partially matches, so almost every contact scores mid-range.
A useful profile names four things:
- **Title** — the specific roles that buy, not the department they sit in.
- **Industry** — narrow enough that one message works across it.
- **Company size** — as a headcount band, because headcount is published consistently and revenue is not.
- **Signals** — the free-text field that does the heavy lifting, and the one most people leave empty.
Signals are where you put the things a dropdown cannot express: "recently hired a first RevOps person", "runs outbound but has no data function", "multiple offices in different countries", "uses a competitor we can displace". These are the qualifiers a good rep applies intuitively, and writing them down is what lets the model apply them too.
Why signals matter more than filters
Firmographic filters are already applied at collection time. If you filtered Leads Finder to VP-level marketers at 50-200 person SaaS companies in Germany, every result already matches those criteria — so scoring them on the same criteria adds nothing. Everything scores identically because everything is identical on those axes.
Scoring earns its cost by evaluating what the filters could not. That is almost entirely the signals field.
This is the single most common reason scoring output looks flat: the profile duplicates the filters instead of extending them.
Run it in the right position
Scoring costs two credits per row, which makes it the cheapest AI operation available and by far the most economically important, because of where it sits in the sequence.
The correct order:
- 1.Collect with tight filters.
- 2.**Score** at two credits per row.
- 3.Enrich or personalise only what survives.
Message Writer costs five credits per row. LinkedIn Profiles costs thirty. Running either before scoring means spending the expensive operations on contacts you are about to discard.
Concretely, on 200 contacts: score everything (400 credits), then personalise the top forty (200 credits) — 600 credits total. Personalising all 200 without scoring is 1,000 credits for a worse campaign, because the effort is spread across people who were never going to convert.
Teams that feel their credit balance is too small are usually running expensive operations in the wrong order rather than running too many searches.
Reading the distribution, not the scores
The useful diagnostic is not any individual score. It is the shape of the distribution.
**Everything clusters at 5-7.** Your profile is too vague, or it duplicates your collection filters. Add signals that describe things the filters cannot.
**Everything scores high.** Your filters are already doing the qualification and scoring is redundant here — or your profile is so broad that everything matches. Either way, save the credits.
**Everything scores low.** Your collection filters are wrong. The list does not match the profile you described, which is worth knowing before you contact any of them.
**A genuine spread.** Working as intended. Take the top slice.
That third case is the valuable one and people misread it as a scoring failure. It is not — it is scoring correctly telling you the list is wrong.
Where the threshold sits
For most teams, six is the deprioritisation line and the top twenty per cent is what gets personalised.
The reasoning is economic rather than statistical. Forty well-researched, specifically-opened messages consistently outperform two hundred generic ones, and cost less to produce. Scoring is what identifies which forty, so that personalisation effort concentrates where it can convert.
There is no universal correct threshold. Start at six, look at your reply rates by score band after a few hundred sends, and move it.
Intent scoring is a different thing
Worth separating, because the two get conflated.
**Lead Scoring** answers "does this person match who we sell to?" It compares a contact against your ICP. Firmographic fit.
**Intent Detection** answers "is this person in-market right now?" It reads a post and scores buying signal. Timing.
They are orthogonal. A perfect ICP match with no current need scores high on one and low on the other. Someone actively shopping who does not match your profile is the reverse.
The strongest lists are people who score well on both, which is why teams running intent monitoring alongside firmographic prospecting outperform teams doing either alone.
Sanity-check the output
Read the top ten and the bottom ten before you act on a scored list. Not all of it — twenty rows.
If the top ten do not look obviously better than the bottom ten to an experienced rep, the profile needs work. This takes two minutes and catches the flat-distribution problem before it costs a campaign.
Also worth knowing: AI output has failure modes and this one is not exempt. It works from published text, so someone with a sparse profile may score low despite being an excellent fit, simply because there was little to evaluate. Scoring is a prioritisation aid, not a verdict.
The honest summary
Lead scoring is not intelligence you buy. It is judgment you already have, written down precisely enough that it can be applied at volume.
Teams that get flat, useless output have almost always skipped the writing-down part. The twenty minutes spent articulating what actually makes a good customer is what determines whether the scoring is worth its two credits per row — and that work pays off well beyond this one feature.
