Every CRM starts clean. Then the imports arrive: a spreadsheet from an event, an export from the old system, a list a partner shared, web forms with no validation, and three salespeople who each have their own way of typing a phone number. A year later, your lead list has "NA" in the company field, names in capital letters, countries spelled five different ways, and phone numbers nobody can dial from the dialler.
Messy data is not only an aesthetic problem. It breaks filters, segments, personalisation, routing and reporting. This guide covers CRM data cleaning in practical terms: the most common problems in lead and contact data, how to fix each one safely, when paying for enrichment makes sense, and how to keep data clean once you have cleaned it.
Why CRM data cleaning matters
Bad data costs you in ways that are easy to miss because they are spread across the whole team.
- •Personalisation goes wrong. "Hi ABHISHEK" or "Hi NA" in an email subject tells the reader the message was automated and careless.
- •Segments leak. A filter for "Country is United States" misses every record that says USA, US or U.S.A.
- •Calls fail. A number without a country code cannot be dialled reliably from a different region, and WhatsApp needs full international numbers.
- •Reports lie. Lead source, region and company-size reports are only as accurate as the fields behind them.
- •Routing breaks. Territory and assignment rules based on country or state send leads to the wrong owner, or to nobody.
- •Trust erodes. Once salespeople stop trusting the CRM, they keep their own spreadsheets, and the data gets worse.
Good data hygiene is not a one-off project. But a proper first clean-up, done carefully, gives you a baseline you can maintain with a small monthly routine.
The most common problems in lead and contact data
Before you fix anything, it helps to know what you are looking for. Almost every CRM we look at has the same handful of issues.
| Problem | Example | What it breaks |
|---|---|---|
| Placeholder values | "NA", "N/A", "null", "-", "test" in real fields | Filters for blank fields, personalisation, reports |
| Inconsistent name case | "ABHISHEK SHARMA", "sara" | Email greetings, mail merges, professionalism |
| Unformatted phone numbers | "98765 43210", "(415) 555-0100", "+91-98765-43210" | Click-to-call, WhatsApp, duplicate detection |
| Country and state variants | USA, US, U.S.A., United States; CA, Calif., California | Segments, territories, regional reports |
| Missing website or company | Work email on file but company and website blank | Account matching, research, enrichment |
| Duplicates | Same person imported twice with different spellings | Double outreach, split history, inflated counts |
| Unknown email type | Personal, work and role addresses mixed together | Audience selection, deliverability |
How to fix each problem
Clear placeholder values
Placeholders are the most misleading data in a CRM because they look like data. A field that says "NA" is not blank, so a filter for "company is empty" misses it, and a mail merge happily prints it. The fix is to turn placeholders into real blanks. Build a list of the placeholder strings your data actually contains ("NA", "N/A", "null", "none", "-", "test" and similar) and clear them. Be careful with short values: a real person can be named "Na", so match whole-field values rather than substrings.
Fix name capitalisation
Names typed in all capitals or all lower case are common from forms and badge scanners. Converting them to proper case ("ABHISHEK SHARMA" to "Abhishek Sharma", "sara" to "Sara") fixes most of them. The important rule is to leave mixed-case names alone. If a name is already "McDonald", "de Souza" or "DeShawn", someone typed it deliberately, and an automatic rule will only make it worse.
Standardise phone numbers to E.164
E.164 is the international phone format: a plus sign, the country code, then the number with no spaces or punctuation, such as +919876543210. Storing numbers this way means every dialler, WhatsApp integration and duplicate check reads them the same way.
The tricky part is numbers without a country code. "9876543210" is probably Indian, but a rule should not assume that. The safe approach is to rewrite only numbers that already include a country code, and to handle local numbers separately. If you know your whole database is in one country, you can set that as a default for numbers without a code. If your business spans several countries, leave local numbers for a human to check.
Standardise countries and states
Pick one canonical form for every country and state and map the variants to it: USA, US and U.S.A. become United States; "CA" becomes California in a US context. City fields also collect metro-area names and abbreviations that are better split into city, state and country. Once the values are consistent, segments and territories work.
A related fix: if the country is blank but the phone number has a country code, the country can be filled from the phone. +91 means India. +1 covers both the United States and Canada, so you may need the area code or other fields to decide.
Fill missing websites and companies
A work email carries useful information. If a contact's email is neha@acme-labs.com and the website field is empty, the website is almost certainly acme-labs.com. This only works for work emails: gmail.com and yahoo.com tell you nothing about the employer, and role addresses such as info@ or hr@ are better left out of the rule too.
The company name can also be estimated from the domain (acme-labs.com suggests "Acme Labs"), but treat that as a guess. Domains are often abbreviations, old brand names or parent companies. If you use this rule, review the results before applying them.
Classify email addresses
Tag every email address as work, personal, role or invalid. This makes it easy to build a B2B campaign audience from work addresses only, keep role addresses out of personal sequences, and find records that need a better address. It also gives you a quick sense of list quality at a glance.
Deal with duplicates
Duplicates deserve their own guide, but the short version is: standardise first, then deduplicate. Two records for the same person often look different only because one has "+91 98765 43210" and the other "9876543210", or one says "ACME LABS" and the other "Acme Labs". Once formats are consistent, matching becomes far more reliable, and you will merge fewer records by mistake.
Preview, apply, undo: the discipline that makes cleaning safe
The biggest risk in CRM data cleaning is not leaving bad data in place. It is overwriting good data with a confident but wrong rule. A few principles keep bulk changes safe.
- Only fill blanks or fix format. A cleaning rule should never replace a real value with a different real value. If the country says India and the phone says +1, flag it for a person; do not change it.
- Preview before apply. Look at sample before-and-after rows for every rule. Five minutes of reading catches the edge cases no rule anticipated.
- Work from current values. If a salesperson edited a field after you previewed, the clean-up should respect their edit, not the old snapshot.
- Log every change at the field level. You want to know what the old value was, what the new value is, and which job changed it.
- Make undo easy. You should be able to reverse a whole clean-up job in one click, without restoring a backup of the entire database.
- Run big jobs in the background. A clean-up should not block the team from working in the CRM.
If your current process is "export to Excel, fix, re-import and hope", it misses most of these. Re-imports can overwrite fields that changed in the meantime, and there is no simple way back.
CRM data cleaning in Vedain Data Health
Vedain's Data Health module is built around the principles above. It scans your leads, contacts or accounts, gives you a health score as a percentage, counts problems by type, and shows sample before-and-after rows so you know exactly what each rule will do.
Free clean-up rules
These rules are included on every plan with no record limit:
- •Clear placeholder values such as "NA", "N/A", "null", "-" and test values, turning them into real blanks.
- •Fix name capitalisation for all-caps and all-lower-case names, leaving mixed-case names alone.
- •Standardise phone numbers that include a country code into international format (+919876543210). Local numbers are not guessed unless an admin chooses a default country for numbers without a code.
- •Fill country from phone number when the country is blank.
- •Standardise countries and states, and split metro areas into city, state and country.
- •Fill website from work email, skipping personal providers and role addresses.
- •Fill company from work-email domain, an optional rule that is off by default because it is an estimate you should review.
- •Classify email addresses as work, personal, role or invalid, for use in filters and campaign audiences.
Rules only fill blanks or fix format; they never overwrite good data. You preview before applying, fixes are worked out from each record's current values, and a field someone edits while the job runs is left alone. Every change is logged per field for 180 days, and any job can be undone from the job history; undo keeps later edits made by your team. Jobs run in the background. On our test environment, 100,000 records were scanned in under 20 seconds and a clean-up applied in under four minutes.
Enrichment with AI: when it is worth paying
Cleaning fixes what you already have. Enrichment adds what you do not. Some of the most useful fields for segmentation, such as company size, seniority and job function, are rarely filled by the lead and hard to fill by hand at scale.
Enrichment is worth paying for when the field will actually change a decision: which leads get routed to senior reps, which accounts are big enough for an enterprise sequence, which contacts are decision-makers. It is not worth paying for fields nobody filters on.
Vedain offers three AI enrichment options inside Data Health. AI company profile and AI title intelligence are billed together as AI enrichment: every workspace gets its first 500 records free once (a lifetime allowance, not monthly), then a single flat pack priced by the number of distinct records actually enriched. Inbox intelligence is billed separately from your Vedain wallet, per record actually updated:
| Enrichment | What it fills | Price |
|---|---|---|
| AI company profile | Company size, country and description, read from the company website | Part of AI enrichment: first 500 records free once per workspace, then a flat pack by records actually enriched — up to 2,000 ₹499 / $14.90 / AED 25; up to 5,000 ₹999 / $29.90 / AED 50; up to 10,000 ₹1,699 / $49.90 / AED 100; then +₹750 / $20 / AED 50 per extra 5,000 |
| AI title intelligence | Seniority and job function from the job title | Included in the same AI enrichment pack as AI company profile (not priced separately) |
| Inbox intelligence | New phone numbers and titles from the sender's own emails; people who have left; referrals to colleagues | ₹1.50 / $0.02 / AED 0.07 |
Inbox intelligence reads your connected mailbox. When a contact's own email signature shows a new phone number or title, it fills the blank. When someone writes that another person has left the company, that claim goes to a review queue where you can mark the person as left. When someone says "please contact my colleague", that becomes an "Add as lead" suggestion. AI output is checked against the email text.
A few safeguards keep costs predictable. You see an estimate before starting, you can set an optional spend cap per job, and records the AI cannot improve are free. Undo restores the previous values, though the charge is not refunded. Vedain uses its own platform AI for these features. See pricing for plan details.
Keeping your CRM clean
A clean-up without a maintenance habit only buys you a few months. These three habits keep the data from sliding back.
Import hygiene
- •Map columns carefully and reject files where the key columns are clearly shifted or merged.
- •Decide the canonical formats (phone, country, state) before importing, not after.
- •Run a clean-up scan straight after every large import, while you still remember where the file came from.
- •Verify email addresses on imported lists before the first campaign goes out. Vedain's Email Verification runs a free quick check on every lead and contact.
- •Record the lead source on every import so you can trace bad data back to its origin.
A monthly data hygiene routine
- Check the health score and compare it with last month.
- Scan leads, contacts and accounts and review the problem counts by type.
- Preview and apply the free clean-up rules.
- Review the "left the company" queue and the referral suggestions if you use inbox intelligence.
- Look at records created in the last month by source, and fix the source that produced the most problems.
- Re-check email addresses on any segment you plan to mail heavily in the coming month.
Once the first big clean-up is done, this is a short routine. The point is consistency, not perfection. If you want to try the routine on your own data, you can create a free Vedain workspace and run a scan.
Track a health score
A single number, reviewed regularly, keeps data quality on the agenda. In Vedain, the "Data health score" appears on the dashboard, the Leads and Contacts lists show a dismissible banner when there is work to do, and the Setup Assistant includes a "Clean your data" step for new workspaces. Data Health is available to Admins by default and has its own permission, so you can grant it to a sales-ops person without giving them full admin rights.
Frequently Asked Questions
How often should I clean my CRM data?
Do a thorough clean-up once, then a light monthly routine, plus a scan after every large import. Teams that import often may want to run the free rules weekly.
Will automated cleaning overwrite data my team entered by hand?
It should not. Safe cleaning rules only fill blanks or fix format, and they work from each record's current values. In Vedain, every change is logged per field for 180 days and can be undone.
What is E.164 format?
E.164 is the international standard for phone numbers: a plus sign, the country code and the subscriber number with no spaces or symbols, for example +919876543210. It is the format most calling and messaging systems expect.
Should I deduplicate before or after cleaning?
After. Standardising names, phone numbers and countries first makes duplicates much easier to detect and reduces the chance of merging the wrong records.
Is AI enrichment accurate enough to trust?
Treat it like any other data source: useful, not infallible. Use it for fields that drive decisions, review samples, and keep the ability to undo. In Vedain, inbox claims about other people leaving a company go to a review queue rather than being applied automatically.
Scan your leads and contacts, see your health score, and apply the free clean-up rules with full undo.
Data Health is included on every Vedain plan.
Try Data Health free