Skip to main content
All articles

CRM deduplication: your CRM checks one field.

HubSpot matches contacts on the email property. Pipedrive wants a name plus a phone, an email or an organisation. They draw the line in different places and they agree on one thing: neither checks deals, which is the duplicate that reaches your forecast.

Toni MedicToni MedicSalestructSeptember 9, 20268 min readCRM & Pipeline
The short version
Illustration of two identical filing cards each showing the same portrait silhouette, with a magnifying glass hovering over one small matching region of each card.
Two records for the same person. The check only ever looks at the small region under the lens, and everything outside it is allowed to differ.

Most people assume their CRM is quietly preventing duplicates in the background. It is, and the check is narrower than almost anyone realises.

HubSpot's documentation is precise about it: when a new contact is added, "HubSpot will look for a matching value in the Email property." For companies, it "looks at the primary values for the Company domain name property." One property per object. That is the automatic check, in full.

The short version

  1. Contacts are automatically deduplicated on the email property. Companies on the primary domain. Deals and tickets have no automatic deduplication described at all.
  2. So the same person with a work address and a personal address is two contacts, permanently, and neither record knows about the other.
  3. The same opportunity entered twice by two reps is two deals, and your pipeline total counts both.
  4. Deduplication is a prevention problem before it is a cleanup problem. Merging is irreversible, so the cheap work is at the point of creation.
  5. Four rules stop most of it: one route in, a required company link on every deal, a rule for personal addresses, and a monthly count rather than a quarterly clean.

What is CRM deduplication?

CRM deduplication is the process of finding records that describe the same real thing and collapsing them into one. It happens at two moments: automatically, when a record is created, and manually, when somebody reviews suggested pairs later.

The automatic half is the one worth understanding, because it is the only part that runs without anyone deciding to do it.

What the automatic check actually covers

Read the property list again, because the gaps are the story.

Contacts match on email. Not on name, not on phone, not on company. So anna@company.com and anna.k@company.com are two people as far as the system is concerned. A prospect who replies from their phone using a personal address is a second contact. A person who changes jobs and comes back is a second contact.

Companies match on primary domain. So a group with three trading names on three domains is three companies, and any account-level number you report is split across them without warning.

Deals and tickets are not automatically deduplicated. HubSpot's page describes no automatic check for them, and lists record IDs only as a way to "manually deduplicate contacts, companies, deals, tickets". This is the one that costs money. Two reps working the same account, or one rep re-entering after a sync failure, produce two deals, and both sit in the forecast.

Pipedrive draws the line in a different place, which is worth seeing side by side because it tells you what to go and check in your own account rather than in mine.

HubSpotPipedrive
PeopleThe email property alone. Creating a second contact on the same address is blockedSame name, plus one of: same phone, same email, or same organisation
CompaniesThe primary domain aloneSame name and same address, so two offices of one group stay separate
DealsNo automatic check describedNot covered by the duplicates feature at all

Two different philosophies. HubSpot matches on one strong identifier and refuses the create. Pipedrive wants a name plus a corroborating detail and flags rather than blocks, which catches the person who used a personal address and misses the person who changed their name.

They agree on exactly one thing, and it is the expensive one. Neither automatically deduplicates deals. Whatever you run, the duplicate that reaches your forecast is the duplicate nobody is watching for.

Illustration of a narrow doorway exactly one cube wide with a single cube passing through it, while several differently sized shapes drift past unobstructed on either side of the frame.
The check is exactly one shape wide. Anything arriving in a different shape passes it on either side and lands as a second record.

Cleanup is the expensive half, and it is irreversible

There is a manual duplicates tool, and it is genuinely useful, but two things about it change how you should think about the work.

Merging cannot be undone. Once two records are merged you are not getting the original pair back, so a bulk merge run by somebody in a hurry is a permanent decision made at speed.

The tool is a review queue, not a fix. It surfaces pairs for a human to compare property by property and choose what to keep. That is the right design and it is also the reason cleanup never finishes: the work is linear in the number of duplicates, so the only way to get ahead is to stop making them.

Which is why the four rules below are all about creation.

The fastest test on your own data

Search your contacts for your own name, then for the two or three people you email most. If any of them returns more than one record, your prospect data is worse than what you just saw, because nobody is watching that.

Four rules that stop duplicates being made

One route in, and everything else feeds it. Duplicates are mostly a symptom of several creation paths running at once: a form, an import, a rep typing manually, a sync from another tool. Pick one primary route per object and make the others go through it. Watch the quarterly import hardest, because it is the only path that creates thousands of records without a human looking at any of them.

Require a company on every deal. Since neither vendor checks deals, the association is the only thing that makes a duplicate visible at all. When both deals hang off the same company, anyone opening that company sees two open opportunities and asks why. When one floats unassociated, nobody ever sees the pair. Make it required before a deal can be created, the same enforcement as an exit criterion on a stage.

Decide what happens to personal email addresses. The single biggest source of contact duplicates, and almost nobody has a rule for it. Two workable answers: personal addresses go into a secondary email property on the existing contact, or they are not captured. Both work. Having no rule is what produces the drift.

Count monthly, do not clean quarterly. A quarterly cleanup is a project. A monthly count is a number. Open the duplicates tool, write down how many pairs it shows, and look at it again next month. A rising count tells you a creation path has broken, which is the one thing a cleanup can never tell you, because it destroys the evidence.

Common questions

Usually on one property per object rather than on a similarity score across the record. HubSpot documents that a new contact is matched by looking for a matching value in the Email property, and a new company by the primary Company domain name property. Anything that differs outside that single property, such as a personal email address or a second trading domain, produces a separate record that the system does not flag.
In HubSpot, no. Its deduplication documentation describes automatic matching for contacts and companies only, and lists deals and tickets solely under manual methods using record IDs. That makes duplicate deals the most expensive kind, because two records for the same opportunity both sit in the pipeline and both are counted in the total.
No. Merged records cannot be reverted, which is why bulk merging done quickly is a permanent decision made at speed. Treat deduplication as a prevention problem first: fix the creation paths that produce duplicates, then work the review queue deliberately rather than clearing it in one session.
Because a cleanup removes the symptom and leaves the cause, which is almost always more than one route into the system: a form, a manual entry, an integration sync and a periodic import all creating records independently. The import is the one to check first, since it is the only path that can create thousands of records without a human looking at any of them.
Monthly, and count rather than clean. Open the duplicates tool, write down how many pairs it shows, and compare it next month. A rising count tells you a creation path has broken, which is information a full cleanup destroys, because the cleanup resets the evidence you would have used to find the cause.

Sources

  1. Pipedrive, How does the Merge Duplicates feature identify duplicates in Pipedrive? People are matched on the same name plus one of "the same phone number, the same email address, or are part of the same organization"; organizations require an identical name and identical address, because "an organization can have multiple branches, offices, or franchises". The article describes contacts only and does not cover deals. Page states last updated 3 September 2026. Fetched and read 9 September 2026 (2026)

  2. HubSpot, Deduplicate records in HubSpot. States that a new contact is matched by "a matching value in the Email property" and a new company by "the primary values for the Company domain name property"; describes no automatic deduplication for deals or tickets, listing record IDs only as a manual method. Page states last updated 26 June 2026. Fetched and read 9 September 2026 (2026)

How these were checked

Both vendors' matching rules were read from their own pages rather than from a search summary. No third vendor is characterised here, and the comparison is deliberately two columns wide rather than presented as how CRMs behave generally.