Customer data dashboard

CRM and ERP Duplicates: Fix the Records Without Recreating the Problem

Home›Blog›Data quality

Where duplicates come from, which source should be the reference, and how to fix them without the next sync recreating them.

Updated on 29 September 20268 min read

One customer, four versions

What happens when each tool decides on its own what is true.

Run the diagnostic

Key takeaways

  • A duplicate is a symptom. The cause is almost always a flow that recreates the record, not an isolated data-entry error.
  • There is not always a single source of truth. The CRM can be the reference for the sales relationship, the ERP for invoicing.
  • The matching key decides everything. Email and company name fail; a stable identifier such as the company registration number holds.
  • Merging before scoping means cleaning a database that will get dirty again within the week.
Your turn

Are your duplicates an incident or a system problem?

Four questions about what you see day to day. At the end, you get your risk level and the three priority actions to take.

A duplicate never looks like a problem. It looks like a detail: a reminder sent twice, an invoice addressed to the wrong contact, revenue that does not match from one dashboard to the next. Taken one by one, these discrepancies can be fixed by hand. Accumulated, they end up casting doubt on everything the CRM displays.

The usual reaction is to launch a big clean-up. It brings relief for a few weeks, then the duplicates come back, because the cause was never addressed. This article takes the opposite approach: understand where duplicate records come from, decide which source is the reference, and only then fix them — so that the fix holds.

The visible duplicate is only the last link

Two identical records are easy to spot. The risk starts earlier, when each tool names, identifies and updates the same customer according to its own rules.

The CRM may use the email address, the ERP an account number and the support tool a freely typed company name. A sync then passes values along without guaranteeing that they describe the same entity.

One record, four versions

What the same company becomes as it moves from one tool to another.

CRMACME Francecontact@acme.frCreated on 12 March
ERPACME FRCustomer 00281Name truncated during migration
SupportAcme SASNo identifierFree text, nothing links it
ReportingACME × 33 sourcesRevenue counted twice

Diagram — fictitious example.

Before merging two records, look for the flow that could recreate them tomorrow.

Where most duplicates come from

  • File imports. A trade-show list, a prospect file, an extract from an old tool: each import creates records without checking whether they already exist, unless a deduplication rule was configured beforehand.
  • Syncs without a shared key. When two tools exchange records with no shared identifier, the connector cannot know that “ACME FR” and “ACME France” are the same company. Whenever in doubt, it creates a new one.
  • Parallel data entry. Sales creates the account in the CRM, accounting in the ERP, support in its own tool. Each does its job well, and the company ends up with three records.
  • Organizational changes. Acquisitions, subsidiary mergers, changes of company name: the company changes its name, not always its record.

In all four cases, the duplicate record is not a human error to be fixed but the logical result of a process. As long as that process remains in place, it produces new duplicates every time it runs.

What duplicates really cost

The cost rarely shows up on a single line. It spreads: a sales rep calling a customer already handled by a colleague, a payment reminder sent to an outdated address, a campaign sent twice to the same decision-maker.

The heaviest cost lies elsewhere. When CRM figures and ERP figures diverge, management stops using them to make decisions, and parallel spreadsheets take over again. The system still exists, but no one trusts it anymore.

Define the source of truth, field by field

There is not always a single reference application. The CRM can own the sales owner while the ERP owns payment terms.

Which field can link these two records?

Same company, two systems, two entries.

FieldCRM recordERP recordVerdict
Emailj.martin@acme.frcompta@acme.fr✗ Two people
Company nameACME FranceACME FR✗ Two spellings
Registration no. (SIREN)812 445 907812 445 907✓ Identical

Only an identifier assigned to the company holds as a key: email changes with people, company name with data entry. Fictitious example.

Write the rule on one page

The reference rule rarely takes more than a page. For each key piece of data — company name, registration number, billing address, sales owner, payment terms — it specifies three things: which system is the reference, who is allowed to change it, and in which direction the value flows during syncs.

As long as this document does not exist, each connector applies its own logic, and it is often that of whichever writes last.

The special case of contacts

For companies, an official identifier such as the SIREN number settles the question. For people, there is no equivalent. The email address remains the best clue, but it changes when someone leaves the company, and the same person may use two.

Good practice is to combine several fields — email, name, associated company — and to reserve automatic merging for cases where they all match.

Deal with duplicates without recreating them

Cleaning comes after scoping. This sequence avoids arbitrarily picking a winning record or losing useful relationships.

  1. Scope List the objects, systems and teams involved.
  2. Measure Quantify certain, probable and to-be-reviewed duplicates.
  3. Decide Define which fields are kept, enriched or subject to validation.
  4. Fix Merge, redirect relationships and keep a record.
  5. Prevent Block inconsistent creations and monitor the flows.

Two points make the difference at the fixing stage. First, distinguish certain duplicates from probable ones: the former share a stable identifier and can be merged in bulk; the latter look alike without formal proof and require human validation.

Second, merging does not mean deleting. A well-run merge attaches to the retained record everything that depended on the others — opportunities, tickets, invoices, communication history — and keeps track of what was combined, so a discrepancy can be explained six months later.

Stop duplicates from coming back

A cleaned database only stays clean if the doors through which duplicates came in are closed.

At creation

A check that looks for an existing record before creating a new one — on the registration number for a company, on several fields for a contact — prevents most unnecessary creations, whether data entry is manual or automatic.

In the flows

Each sync must use the chosen matching key and refuse to create a record when it cannot find a certain match. A record awaiting validation is better than a silent duplicate. Imports follow the same rule.

In the organization

Data without an owner degrades. Naming, for each key piece of data, the person or team accountable for it, and reviewing newly created records every week, is often enough to spot a faulty flow before it has produced hundreds of records.

Choose the next action based on the risk

Three situations, three starting points. The diagnostic at the top of this article tells you which one is yours.

Three situations, three starting points

Fix

When

Duplicates are localized and their cause is known.

Action

Process the batch, then monitor new records.

Scope

When

Discrepancies affect several objects or several flows.

Action

Map the rules before any merge.

Audit

When

Figures and identities diverge between systems.

Action

Measure quality and assign responsibilities.

In most cases, the right starting point is a short assessment: measure the volume of probable duplicates, list the flows that create records, and identify the data that has no owner. That is how the projects we run at Mirakl or Orhatek begin: reconciling master data before touching the records.

This work takes a few days. It turns a vague impression — “our data isn’t reliable” — into a list of measured discrepancies, identified causes and decisions to make.

To go further, discover our CRM integration offer, our ERP integration support and our data expertise.

Identify duplicates, their origin and the rules that should prevent them.

Discuss your scope

A turnK consultant replies within 48 business hours.

the most common questions

How can I tell whether my duplicates come from a data-entry error or a flow?

Look at when and where they appear. If they arrive in batches after a sync or an import, a flow is recreating them; if they are isolated and scattered, data entry is more likely. In the first case, fixing the records is not enough: you need to fix the flow.

Which field should be used to match records between the CRM and the ERP?

A stable identifier assigned to the company, such as the SIREN number in France. An email address identifies a person and disappears when they leave; the company name varies with data entry. For contacts, combine several fields rather than relying on just one.

Should you choose a single reference application for all data?

No, the source of truth is defined field by field. The CRM can be the reference for the sales owner and the ERP for payment terms. What matters is that the rule is written down and applied by the syncs.

How do you stop duplicates from coming back after the clean-up?

By blocking inconsistent creations at the source and monitoring the flows. Check whether a record exists before any creation, enforce the matching key in syncs and regularly review new records.