Lead scoring your reps actually trust.
Most scoring models are a point system somebody guessed at once, and reps figure out within a month that the number doesn't predict who buys. Here's how to build one from your own closed-won history instead, weight it, and check it every quarter.

A lead score is supposed to tell a rep which record to call first. Too often it sits on the record unread instead, because somebody built it once, by guessing which fields "felt" important and assigning points to each one, and never went back to check whether the score actually separated the deals that closed from the ones that didn't. A rep learns that within a few weeks of working the list, and once they've learned it, the field becomes decoration.
The short version
- A lead score reps have learned to ignore is worse than no score: it occupies the one field a rep would otherwise use their own judgment on, and it hides the leads that are actually likely to close among the ones that just look busy.
- The fix is building the model backward from your own closed-won and closed-lost records, not forward from a list of fields that sound like they should matter.
- HubSpot (Marketing Hub or Sales Hub Professional and Enterprise) and Salesforce (Sales Cloud Einstein, included in Performance and Unlimited, an add-on on Enterprise) build models from your CRM data. Pipedrive's Scores tool (Premium and Ultimate) is rule-based, built by hand, and today only scores deals, not leads or contacts.
- Peer-reviewed research on why reps skip marketing leads found that as reps gain experience, they respond less to being told to follow up and more to the quality of the lead prequalification process itself. A model a rep doesn't trust gets the same treatment as no model.
- A score with no scheduled recheck decays the same way any other CRM rule does. The deals it was built on and the fields it weights both drift, and nobody notices until a rep points out the score doesn't match reality anymore.
What is lead scoring?
Lead scoring is a number assigned to a lead, contact or deal that's meant to represent how likely it is to close, so a sales team can work the highest-probability records first instead of working the list in the order it arrived. In most CRMs it's built from two kinds of input: fit (who they are: company size, job title, industry) and engagement (what they've done: pages visited, emails opened, forms submitted). A record's score is either the sum of manually assigned points across a set of rules, or a probability produced by a model trained on your own historical data.
Both approaches can work. Both also fail the same way in practice: the rules or the model get built once, from an assumption about what a good lead looks like, and nobody schedules a date to compare the score against what actually closed. MQL vs SQL covers who owns a lead once marketing and sales agree it's qualified. This is about the number that's supposed to help you decide when it's qualified in the first place.
Why do reps stop trusting a lead score?
Because most scores are built from a guess about what should predict a close, not from a check of what actually did, and a rep who works the list every day finds out which is which faster than the model gets rebuilt.
This isn't a new problem, and it isn't only about scoring software. A 2013 study in the Journal of Marketing looked at why sales reps don't pursue a large share of the leads marketing hands them, using data from 461 sales reps across four firms. Its central finding, in its own words: as reps' experience increases, "their responses to managerial tracking of lead follow-up and marketing lead volume decrease," while "responses to the quality of the lead prequalification process increase." Put plainly: telling an experienced rep to work the list, or sending them more leads, moves the needle less than actually building a qualification process the rep believes is accurate. A number that doesn't hold up under a rep's own experience gets the same response as a manager's reminder email: acknowledged, then ignored.
A lead score is a prequalification process wearing a single-field costume. If it was built on an assumption instead of a check, it inherits the exact problem that paper describes, just automated.
How do you build a lead score from your own closed-won data?
- Pull your last 50 to 100 closed-won and closed-lost records
Export the full property set for the last quarter or two of closed deals, both won and lost, not just the wins. You need the lost column to know what a bad fit or a stalled deal actually looked like at intake, which is the half of the comparison a points-based model built from a features list never has. It fails as a starting point if your CRM doesn't record why a deal was lost, in which case the lost-reason field is the first thing to fix, before scoring anything.
- Find which fields actually separated the two groups
For every fit field (company size, industry, job title, region) and every engagement field (form type, pages visited, email opens, time to first reply), compare the distribution across your won pile against your lost pile. A field that looks identical in both groups is not a scoring criterion, however intuitive it seems. A field where the won group clusters tightly and the lost group is scattered is. It fails if you skip straight to weighting fields you already believed mattered, which is the exact step that produces a score nobody can defend when a rep asks why a record they know is dead scored higher than one they know is hot.
- Weight only what showed a real gap, and drop the rest
HubSpot's own scoring tool is built around this structure: group your criteria (for example, "engagement with sales" and "engagement with marketing"), give each group a point ceiling, and give each rule inside the group its own points. In HubSpot's own worked example, a 100-point score splits into a 60-point sales-engagement group (a started call worth 15, a booked meeting worth 10, a completed meeting worth 20) and a 40-point marketing-engagement group (a CTA click worth 6, an email open worth 2, a link click worth 5). A contact who books a meeting, opens two emails and clicks three links scores 10 in the first group and 19 in the second, for 29 out of 100. It fails if every field gets roughly equal points because nobody wants to argue about which matters more. That instinct is exactly what step 2 exists to overrule with a real comparison.
- Add negative weight for what predicted a loss, not just positive points for what predicted a win
A field that's common in your lost pile and rare in your won pile is a signal too, and most hand-built scores only add points, never subtract them. HubSpot's own documentation gives this example directly: add points for strong marketing engagement, but subtract points for a recent unsubscribe; add points for a good industry and size fit, but subtract points for a lead in a region you don't operate in. It fails if the model can only go up. A score that only accumulates treats a lead who unsubscribed last week the same as one who didn't, which is precisely the kind of gap a rep notices on the first bad call.
- Set the score threshold from your own SLA, not a round number
Once the model exists, the cutoff for "call this one first" should come from how many leads your reps can actually work at the response speed you've committed to, not from picking 50 or 70 because it sounds like a sensible midpoint. If your team can hit a same-day call on 30 records a week, the threshold is whatever score produces roughly 30 qualifying records a week from your actual lead volume, adjusted from there. It fails if the threshold is set once and never revisited against changing lead volume, because a fixed cutoff against a growing top of funnel quietly waters down what "high score" means.
Why last quarter's model won't fit next quarter

HubSpot's own scoring tool has a decay mechanism built in for exactly this reason: you can set an individual event's score to shrink over time, on a 1, 3, 6 or 12-month interval, so a form fill from eight months ago doesn't carry the same weight as one from yesterday. That handles one kind of drift, the age of a single signal. It doesn't handle the other kind: your product changes, your ICP shifts, a channel that used to bring in your best-fit leads starts bringing in tire-kickers, and the fields that separated wins from losses last quarter stop separating anything.
The fix is the same five steps above, run again on a schedule rather than once. Pull the last quarter's closed-won and closed-lost, re-check which fields still show a real gap, and re-weight. A field that mattered in Q1 and stopped mattering in Q3 is not a bug in the model, it's the market telling you something, and a score built once in January and never rechecked will not tell you when that happens.
Manual rules or a predictive model: what's the actual difference?
Both HubSpot and Salesforce offer a second option beyond hand-built rules: a model trained on your own historical data that produces a probability instead of a rule-based point total. The trade is the one every predictive model makes: less manual weighting, less ability to explain any single score.
HubSpot's predictive lead scoring (Marketing Hub or Sales Hub Enterprise) uses machine learning to output a "Likelihood to close" score: the percentage probability a contact converts to a customer within 90 days, plus a "Contact priority" tier (Very High, High, Medium, Low, Closed Won) that splits contacts into even 25% bands by that score. It's trained on a wide input set HubSpot publishes in full: page views, site visits, social clicks, days since last visit, email opens and clicks, form submissions, logged notes, days since last contact, whether the contact has a phone number, whether the email domain is a free one like Gmail, and firmographic data about the contact's company and about your own account. HubSpot is explicit about the limit of this approach: it calls the method "blackbox machine learning," meaning the inputs and the output are known but exactly how one turns into the other isn't, so you can see that a contact scored 22% and you cannot get a rule-based explanation of why.
Salesforce's Einstein Lead Scoring works the same way in principle: it's included with Sales Cloud Einstein on Performance and Unlimited Editions, and available as an add-on on Enterprise. Setup is a guided flow in Salesforce Setup: choose a conversion milestone (lead converts to an account and contact, or a lead creates an opportunity), choose whether to score all leads together or split them into up to 35 segments by a field like country or source, and choose which lead fields the model considers. If a segment doesn't have enough of your own converted-lead history to build a reliable model, Salesforce falls back to what it calls a "global model," built from anonymized data pooled across Salesforce customers, and switches to your own model once you've accumulated enough conversions for it to outperform the global one.
Pipedrive's Scores feature is the odd one out, and it matters which kind it is before you compare it to the other two. It's rule-based, not predictive: you build a score by hand, in groups labeled Highly Positive (+25 points), Slightly Positive (+10) and Negative (-10), using AND/OR conditions on deal and activity fields, the same way Pipedrive filters work. It's available on Premium and Ultimate plans, is managed by deal admin users, and as of this writing, the target entity is deals only. There is no lead or contact scoring in the product today.
Where a lead score breaks even when it's built correctly
A model built from real closed-won and closed-lost data is not the same thing as a model that will stay right, and it isn't a substitute for a rep's own read on a live conversation.
Small sample sizes produce confident-looking scores that aren't. If your last two quarters only closed 15 deals, the "fields that separated wins from losses" step above is working with a sample too small to trust. A field that happened to appear in 4 of 5 wins by chance will look like a strong signal and isn't. Salesforce's own fallback (using a pooled global model when a segment lacks enough conversion history) is an acknowledgment of exactly this problem; a hand-built score has no equivalent safety net unless you build one, by keeping the weighting simple and revisiting it more often while your volume is low.
A score describes the record, not the conversation that hasn't happened yet. A lead can score high on every fit and engagement field and still tell a rep on the first call that the project is shelved for the year. The score is a triage tool for where to spend the first call, not a replacement for what the rep learns on it.
A free email domain or a missing field isn't always a bad-fit signal. HubSpot's predictive model includes "whether the contact's email is a free mail domain" as an input, and Pipedrive's rule builder makes it easy to write a rule that subtracts points for exactly that. Founders and buyers at very small companies genuinely do run business communication through Gmail sometimes. Treat any single field, including this one, as one input among several rather than a disqualifier on its own.
What to check this week
Pull ten records that scored high last quarter and see how many actually closed. If the honest number is well below what the score implied, the model is describing what someone believed six months ago, not what's closing now.
Ask a rep, not a dashboard, whether they trust the score. If the answer is "I don't really look at it," you don't have a lead-scoring problem, you have an unused field, and the fix is rebuilding it from step 1 above, not adding more criteria to a model nobody opens.
Check whether your score can go down as well as up. If every rule only adds points, an unsubscribe, a bounced email or a lead moving out of your ICP region is invisible to it, and step 4 above is the gap to close first.
This connects to the same enforcement problem covered in pipeline stage exit criteria and in who's allowed to change the routing logic: a scoring model, like a stage rule, only stays accurate if someone owns rechecking it, and the fields it depends on are the same ones CRM hygiene tracks the decay rate of.
Common questions
Sources
HubSpot Knowledge Base, Overview of the lead scoring tool. Describes score groups, group and score limits, the worked 100-point example (60-point sales-engagement group, 40-point marketing-engagement group, scoring 29 points from a booked meeting, two email opens and three link clicks), positive and negative points, and score decay by 1/3/6/12-month interval. Page states last updated 2 September 2026. Fetched and read 29 September 2026 (2026)
HubSpot Knowledge Base, Determine likelihood to close with predictive lead scoring. Describes the Likelihood to Close and Contact Priority properties, Enterprise-only availability, the full published input list (page views, site visits, social clicks, email activity, CRM interactions, firmographic data), and states the model uses "blackbox machine learning" where inputs and outputs are known but not how one produces the other. Page states last updated 11 January 2026. Fetched and read 29 September 2026 (2026)
Salesforce Help, Enable Einstein Lead Scoring. States required editions (Sales Cloud Einstein, included in Performance and Unlimited, extra cost on Enterprise), the guided setup flow (conversion milestone, up to 35 lead segments, included fields), and that when a segment lacks sufficient conversion data, Einstein uses "a global model" built from "anonymous data from many Salesforce customers" until enough of the org's own data accumulates. Fetched and read 29 September 2026 (2026)
All 5 sources and how they were checked
Pipedrive Knowledge Base, Scores in Pipedrive. States availability on Premium and higher plans, that scores currently target deals only, the three criteria categories (Highly Positive +25, Slightly Positive +10, Negative -10), AND/OR condition logic matching Pipedrive's filters, and that scores are managed by deal admin users. Page states last updated 3 September 2026. Fetched and read 29 September 2026 (2026)
Sabnis, G., Chatterjee, S.C., Grewal, R., & Lilien, G.L., The Sales Lead Black Hole: On Sales Reps' Follow-Up of Marketing Leads. Journal of Marketing, 2013, Vol. 77, Issue 1, pp. 52-67. Abstract states, from data on 461 sales reps across four firms: "as sales reps' experience increases, their responses to managerial tracking of lead follow-up and marketing lead volume decrease; responses to the quality of the lead prequalification process increase." Abstract fetched and read 29 September 2026 (2013)
HubSpot, Salesforce and Pipedrive names and logos are trademarks of their respective owners, shown only to identify the products discussed. Salestruct is not affiliated with, sponsored by, or a reseller of any of them; we build sales systems on all three, and this page is written from what a sales team using them will actually encounter.
If you want a second pair of eyes on your own setup, Salestruct runs a free diagnostic.
Read next.
All articlesMQL vs SQL: definitions and handoff criteria
Both terms are usually explained as slide definitions. In your CRM they are fields, and only one mainstream platform ships them out of the box. Here is what each one actually is, in five systems, with the thresholds that move a lead across.
CRM pipeline stages: give every one an exit criterion
Your CRM ships with stages named after things your rep did, and it multiplies your forecast by a number attached to each one. Booking a presentation does not make a deal 60% likely to close, but that is what the default pipeline asserts.

