Contact deduplication is the step that decides whether an event CRM integration produces numbers
you can defend. Before a registration can count toward a deal, a renewal, or a repeat-attendance
figure, you have to say with stated confidence that the registration record and the CRM record
are the same person. The method is to match on keys in a fixed priority order, sort every pair
into one of three confidence tiers, and keep the middle tier away from any automated merge.

The common shortcut is a lookup on email address. It works for many records and fails silently
on the rest.

## Why does an event CRM integration find the same person five times?

Because each system captured them at a different moment, for a different purpose, on a
different form.

- **Registration platform.** Captured at purchase, often with a personal email typed on a
  phone. The most current job title and the least stable identifier.
- **CRM.** Captured by a salesperson or an import. Work email, and a title that may be two
  roles out of date.
- **Marketing automation.** Captured at the first form they filled, years ago. Rich behavioral
  history, stale contact details.
- **AMS or membership database.** The authoritative member record, with a member ID that
  nothing else in the stack knows about.
- **Finance system.** Captured at invoice. Often the billing contact, sometimes the employer.

Each system recorded what it needed, and the cost lands downstream:
[event marketing attribution](https://eventiq.io/md/event-marketing-attribution), cohort analysis,
[attendee retention](https://eventiq.io/md/blog/attendee-retention), and renewal correlation all inherit the error
in the match before they inherit anything else.

In a Vendelux survey of more than 120 B2B marketing and events leaders in 2026, run by a vendor
selling to that same market, 90% said events influence deals that get no credit in their CRM.
A registration record never connected to the opportunity record is one way that happens, and
the [event data silos](https://eventiq.io/md/blog/event-data-silos) page covers the others.

## Which match keys, and in what order?

Work down the list, stop at the first key that produces a match at the confidence you require,
and record which key produced it. A match you cannot explain is a match you cannot defend when
a number is disputed.
| Priority | Key | Reliability | Where it fails |
| --- | --- | --- | --- |
| 1 | Stable person ID captured at registration: member ID, customer ID, SSO subject, or a badge ID linked to a person | Highest | Exists only if you designed for it |
| 2 | Exact email address, normalized: lowercase, trimmed, plus-tags stripped, Gmail dots collapsed | High | Work versus personal address, shared inboxes |
| 3 | Email local part plus organization domain, where the change of employer is known: j.smith at the old domain and at the new one | Medium-high | Common local parts |
| 4 | Last name plus first-name variant plus organization: Kate and Katherine, Bob and Robert, compound surnames | Medium | Common names at scale |
| 5 | Phone number, normalized to E.164 | Medium | Switchboards, shared mobiles |
| 6 | Last name plus postal code, or last name plus organization plus job function | Low | Review only, never an automatic merge |
| 7 | Fuzzy full-name similarity alone | Lowest | Candidate generation only, never a match |
Two rules. Never treat two low-reliability keys combined as strong: two weak signals about
common names produce confident nonsense at scale. And run keys in sequence, so that when a
record matches key 2 and key 4 to different people, the higher key wins and the conflict is
logged.

Normalization does more work than the algorithm. Expect to spend more time on it than on
matching logic.

## How do you assign confidence tiers?

Three tiers, because the middle tier is where the damage happens. A two-tier scheme forces
every ambiguous pair into an automatic merge or an automatic reject, and both are wrong.
| Evidence | Tier | Action |
| --- | --- | --- |
| Stable person ID matches | A | Link automatically |
| Exact email, single candidate | A | Link automatically |
| Exact email, two candidates | B | Review. Possible shared inbox |
| Email local part plus known organization change | B | Review |
| Name variant plus same organization plus same function | B | Review |
| Exact phone, mobile format | B | Review |
| Exact phone, switchboard format | C | No link |
| Last name plus postal code only | C | No link |
| Fuzzy name only | C | No link |
| Two Tier A matches that conflict | B | Escalate, do not link |
- **Tier A.** Link. Log the key used and the timestamp.
- **Tier B.** Hold as a candidate pair with its evidence. Route it to a named person. Record
  the decision and who made it. Never merged by an automated process.
- **Tier C.** No link. Keep for future evidence, and never count as matched.

The Tier B rule exists because of an asymmetry. A missed match
understates a number and is recoverable. A false merge puts one person's registration history,
spend, and communications against another's, and is very hard to unpick once downstream systems
have copied it.

Review capacity is the binding constraint, so triage Tier B by value: largest spenders, board
members, sponsor contacts, and at-risk renewals first.

## What breaks matching, and what do you do about it?

Four cases come up in most programs, and each needs a specific response more than it needs a
better algorithm.

**Work versus personal email.** Someone registers on a Sunday with a personal address and the
CRM holds their work address. The fix is at capture: ask for work email as a separate labeled
field, and ask for the employer, which makes keys 3 and 4 usable.

**Name changes.** Marriage, divorce, transliteration, a corrected misspelling. The old record
keeps the old name, so the person looks like two people, one of whom stopped engaging. Keep
prior names as alternate values, and treat first name plus organization with a changed surname
as Tier B.

**Shared and departmental inboxes.** events@, info@, and the assistant's address used for four
executives. An exact email match here is not a person match, and treating it as Tier A is how
one address ends up owning forty registrations. Keep a suppression list of shared-inbox
patterns and force matches on it to Tier B.

**Bookings by an agency, assistant, or employer.** The contact details belong to the booker,
and group bookings do this systematically. The fix is a required attendee-level field: name,
work email, and title per seat, not per transaction. Without it, a 40-seat corporate booking is
one identity in your data and forty people in the room.

A fifth case is common names at large organizations. Two people with the same name at one
large employer will match on key 4 every time, so send those to Tier B.

## How do you keep the match from decaying?

Matching never finishes. Details change, people move employers, and every event adds records.

**Capture a stable identifier at registration.** The highest-value change available, and a
form field more than a data project. Authenticate members so the registration carries the
member ID, and issue your own person ID to non-members on first contact.

**Persist the link, not the result.** Store resolved identities as a mapping table of system
IDs to one person ID, with key used, tier, timestamp, and decision maker, and never overwrite
source records. You can re-run with better rules without losing human decisions, and show how
each link was made when a figure is disputed.

**Re-run on a schedule.** Monthly for most programs, weekly during an active registration
cycle. New evidence can promote a Tier B pair or reveal that an existing link was wrong.

One boundary: resolving two records into one person does not merge their permissions. Treat
consent state as its own attribute, and take the rules to counsel.

## How do you know whether the matching is working?

The match rate will not tell you. It measures coverage and says nothing about correctness, and
it rises whenever you loosen the rules, which makes the data worse while the metric improves.

**Match rate = Records linked to a person ID ÷ Total records**

**Tier A share = Tier A links ÷ Total links**

**Review backlog = Tier B candidate pairs awaiting a decision**

**Precision = Correct links in sample ÷ Total links in sample**

**False-merge rate = Incorrect links in sample ÷ Total links in sample**

**Estimated recall = Correct links ÷ (Correct links + Missed matches, estimated from a hand-checked sample of unmatched records)**

Pull a random sample of 200 links, have someone who knows the audience check each by hand, and
record precision and the false-merge rate. Then check 200 unmatched records for people who
should have matched, and scale that miss rate to all unmatched records to estimate recall. Set
a false-merge tolerance before you measure, and state the achieved rate alongside any figure
that depends on the match.

## Example: 1,800 registrations across five hypothetical systems

Take an association annual meeting with 1,800 registrations, matched against a CRM of 34,000
contacts, a marketing platform with 41,000 records, an AMS with 12,400 members, and a finance
system with 1,140 invoices. All figures are hypothetical.
| Key | Registrations | Cumulative share | Tier |
| --- | --- | --- | --- |
| 1. Member ID present at registration | 740 | 41.1% | A |
| 2. Exact normalized email | 642 | 76.8% | A |
| 3. Local part plus known organization domain | 94 | 82.0% | B |
| 4. Name variant plus organization | 128 | 89.1% | B |
| 5. Phone, mobile format | 31 | 90.8% | B |
| 6 and 7. Low-confidence candidates only | 103 | 96.6% | C |
| No candidate at all | 62 | 100.0% | none |
| Total | 1,800 |  |  |
Read the result by tier: 1,382 records are usable without review, 253 need human decisions,
and 165 are effectively unmatched. Reporting "we matched 90.8%" counts the review
queue as resolved.

First diagnostic: 1,060 registrations carried no member ID, while the AMS holds 12,400 members.
If a plausible share of those are members whose ID was never captured, the highest-return fix
is authenticating members at registration.

Second: suppose hand review of the 253 Tier B pairs finds 181 correct, 44 incorrect, and 28
undecidable. 44 ÷ 253 = 17% of the queue would have been false merges, which is the argument
against auto-merging the middle tier. Confirmed links after review are 1,382 + 181 = 1,563, or
86.8% of registrations.

Third, check Tier A as well. Suppose a 200-link sample of Tier A
finds 3 incorrect, all exact email matches to shared addresses. That is a 1.5% false-merge rate
at the tier nobody reviews, about 21 wrong links across 1,382, fixed by a suppression list.

Finally, finance: 1,140 invoices for 1,800 registrations, because group bookings put several
seats on one invoice. Matching on payer email would turn every seat on a group invoice into one
identity.

The conclusion that survives is about capture. 1,060 ÷ 1,800 = 59% of registrations arrive
without a stable identifier, and every rule after key 1 compensates for a missing form field.

## What to do this quarter

- Add the stable identifier at registration: authenticate members, and issue your own person
  ID for non-members.
- Require name, work email, and title for every seat in a booking.
- Write the key priority list and the three-tier decision table, and get the operations owner
  and the IT director to sign the same version.
- Build the shared-inbox suppression list and force matches on those addresses to review.
- Stand up the mapping table with key used, tier, timestamp, and decision maker.
- Hand-check 200 links and 200 non-matches, publish precision, recall, and false-merge rate,
  and set next quarter's tolerance.

## Common questions

### Can we buy a tool that does contact deduplication for us?

Tooling helps with candidate generation, normalization, and queue management. It does not
decide your key priority, thresholds, shared-inbox list, or tolerance for false merges, and
those determine whether the output is trustworthy. Someone in your organization writes them.

### What is an acceptable match rate?

There is no portable figure, and chasing a high one is how false merges get introduced. Report
the tier breakdown, and hold yourself to a false-merge tolerance in place of a coverage target.

### Who should own this?

One named person, with decision rights written down: who sets key priority, who approves a
Tier B merge, who can change a threshold. Without that, two teams build two match sets and two
dashboards that disagree. See [event data management](https://eventiq.io/md/event-data-management) for the wider
set of rules.

## Where EventIQ fits

EventIQ replaces nothing. It connects on top of the platforms you already run: event platforms
(Cvent, Zoom, Swapcard), CRM (Salesforce, HubSpot, GoHighLevel), and marketing (Google Ads, Meta
Ads, LinkedIn Ads, Mailchimp, Google Analytics). Platforms with an API outside that list are
connected on request.

For this topic the product holds one key. Attendees are matched to CRM contacts by exact email,
and the result stays on each attendee record: matched, unmatched, or no email. Exact email is
the only key, so a personal address at registration against a work address in the CRM shows as
unmatched. [Salesforce](https://eventiq.io/md/integrations/salesforce) brings deals with stage, amount,
and close date, linked to an event through a campaign relationship a person confirms. HubSpot
brings contacts, companies, and deals, with no event link.

The product stops there. EventIQ does not deduplicate people across sources: a person who
arrives through two sources is counted twice in the figures built on those records. It does not
match on name, phone, or organization, sort pairs into tiers, or merge records. No screen shows
an overall match rate, so the tier breakdown and the sample check are yours to run. There is no
AMS connection, so the member ID in key 1 is a join you make yourself.

[Book a demo](https://eventiq.io/#early-access) to see the exact-email match result on each attendee record on a
sample event, in a 20-minute demo.

---

HTML version: https://eventiq.io/blog/event-identity-matching
