Last updated: August 29, 2026
By Ben Argeband, Founder & CEO of Heartbeat.ai
Most “data quality” arguments between recruiting ops teams and data vendors happen because nobody agreed on what a field actually means. Someone says a phone number is “bad” when really it’s a main line being treated like a direct dial, or an email is “unverified” when it just hasn’t been sent to yet. This page is a working data dictionary for provider contact data fields — plain definitions, how recruiters actually use each one, and how to measure whether a field is doing its job.
What’s on this page:
Who this is for
Buyers and ops teams who need to understand what fields mean before they build routing rules around them. If you’re mapping provider data into an ATS/CRM, dialer, or sequencing tool, this is meant to cut down on rework and prevent bad automation decisions.
- TA ops and recruiting ops building field mappings and dashboards
- Agency ops standardizing outreach workflows across recruiters
- Data and RevOps teams integrating provider datasets into internal systems
Quick answer
- Core answer
- A provider contact data dictionary defines each field, its allowed values, and how recruiters use it to route outreach, enforce suppression, and measure outcomes.
- Key insight
- You can’t validate, suppress, or prioritize outreach if fields don’t have consistent definitions and timestamps attached to them.
- Best for
- Buyers and ops teams who need shared field semantics before building routing logic.
Compliance and safety note
This is intended for legitimate recruiting outreach only. Respect candidate privacy, opt-out requests, and applicable data laws. Nothing here is medical, legal, or compliance advice.
A simple framework: field-level trust
Every field you operationalize should have four things attached to it: a definition, an intended workflow use, a way to measure it, and a known failure mode. Skip any one of these and the field becomes a source of disagreement instead of a source of decisions.
- Define the field in plain language — what it is and what it is not.
- Constrain it with allowed values, formats, and null rules.
- Route it into an actual decision: call vs. email, prioritization, suppression, or assignment.
- Measure outcomes tied to the field, using denominators you can audit (per 100 dials, per 100 sent emails, per 100 delivered emails).
- Correct the definition or the workflow rule based on what the failure modes tell you.
Common mistakes teams make
- Mixing “type” with “quality.” “Mobile” is a type; “line tested” is a quality signal from a point in time. Conflating them makes reporting misleading.
- Assuming a field implies consent. A contact path existing in a dataset isn’t permission to contact it repeatedly or off-purpose.
- Overwriting multi-path contact data. Normalizing every phone into one “phone” field destroys your ability to route or QA anything.
- Not versioning definitions. If what a field means changes quietly, your dashboards start lying to you without anyone noticing.
Step-by-step: building the dictionary into your workflow
-
Inventory your contact fields and map each to a decision. List every field that affects outreach — phones, emails, locations, identifiers, specialty, suppression flags, timestamps — and write down the specific decision each one drives.
-
Standardize definitions and allowed values. Define format (E.164 for phone, standard constraints for email), allowed values, and null rules. If a field can be multi-valued — multiple locations, multiple phones — store it that way instead of overwriting.
-
Add provenance and freshness metadata. At minimum, store a source category and a last-updated timestamp for each contact path. Freshness is what lets you explain a performance dip without guessing.
-
Implement suppression as a first-class system. Maintain do-not-contact flags, channel-level opt-outs, timestamps, and reasons. Enforce suppression before any dial or send, across every tool involved — ATS/CRM, dialer, email platform.
-
Instrument outcomes so fields can be QA’d. Log outcomes with denominators you can audit: phone metrics per 100 dials, email metrics per 100 sent and per 100 delivered, identity metrics per 100 records for match and duplicate rates.
-
Close the loop weekly. Review the top two or three failure modes — wrong number, gatekeeper routing, bounces, opt-outs — and update routing rules, suppression, and field definitions accordingly.
Jump to a field group
- Identity fields (who is this?)
- Segmentation fields (who should work this?)
- Phone fields (how do we call?)
- Email fields (how do we email?)
- Location fields (where do they practice?)
- Suppression and preferences (who should we not contact?)
- Provenance and freshness (why trust it today?)
Field-level troubleshooting table
Use this to align definitions, workflow usage, and QA measurement across teams.
| Field / concept | Definition | How recruiters use it | Common failure mode | QA / measurement |
|---|---|---|---|---|
| NPI | National Provider Identifier — a unique identifier for covered health care providers in the U.S., used under HIPAA administrative simplification rules. | De-duplicate records; match across systems; anchor identity when names or locations vary. | Individual vs. organization NPIs mixed together; stale practice location attached to the record. | Match rate per 100 records; duplicate rate per 100 records after matching rules. |
| Provider name | Structured name fields (first/last/suffix) used for matching and personalization. | Personalize outreach; match to internal records. | Nicknames, initials, and suffix mismatches create duplicates. | Duplicate rate per 100 records; manual review queue volume. |
| Taxonomy | Standardized classification describing provider type and specialization. | Route to the right recruiter; tailor messaging; segment reporting by specialty. | Multiple taxonomies per provider; taxonomy too broad for routing. | Segment-level reply rate per 100 delivered emails. |
| Role / specialty label (internal) | Your internal normalized specialty label mapped from taxonomy and business rules. | Assignment, comp routing, pipeline reporting. | One-to-many mapping causes misroutes, especially for subspecialties. | Reassignment rate per 100 assignments; override frequency. |
| Direct dial | A phone number intended to reach the provider without going through a switchboard. | Prioritize for first attempts; use for time-sensitive outreach. | Routes to front desk; forwarding changed; number reassigned. | Connect rate per 100 dials; wrong-number disposition rate per 100 dials. |
| Main line | A practice or facility number that typically reaches a switchboard or front desk. | Gatekeeper-friendly outreach; confirm best contact path. | High gatekeeper friction; long holds; after-hours routing. | Connect rate per 100 dials; gatekeeper disposition rate per 100 dials. |
| Phone type | Enum describing the phone’s role — direct dial, main line, scheduling, fax, unknown. | Routing (who to call first) and suppression (never dial fax). | Fax mislabeled as voice; direct dial overwritten by main line. | Disposition mix per 100 dials by phone type. |
| Phone extension | Extension required to reach a person or department after dialing a main number. | Reduce wasted dials; improve gatekeeper routing. | Extension missing or outdated. | Connect rate per 100 dials, with vs. without extension. |
| Line tested | A phone number checked programmatically for basic callability at a point in time. | Prioritize callable numbers; reduce wasted dials. | Test is stale; callable but reaches the wrong department. | Connect rate per 100 dials segmented by last test date. |
| Phone carrier / line class (when available) | Carrier or line classification metadata used for routing and QA — not a guarantee of reachability. | Sequence design and troubleshooting dead-end numbers. | Ported numbers and carrier changes make this stale. | Connect rate per 100 dials by carrier/line class. |
| Timezone | Timezone associated with the practice location, used to schedule outreach windows. | Call-window planning; reduce after-hours attempts. | Provider practices across timezones; timezone inferred from billing address. | Answer rate per 100 connected calls by local time of day. |
| Best call window (operational) | An internal field derived from observed outcomes indicating when calls are likeliest to reach a human. | Schedule call blocks; route to recruiters working that window. | Overfitting to small samples; not refreshed as patterns change. | Answer rate per 100 connected calls by call window. |
| Email address | A professional email associated with the provider or their practice context. | Compliant email outreach; nurture when calling is low-yield. | High bounce rate; role-based inbox; spam filtering. | Deliverability rate per 100 sent; bounce rate per 100 sent; reply rate per 100 delivered. |
| Email type | Enum describing the email’s role — personal professional, practice, role-based, unknown. | Sequence design; routing to admin vs. provider. | Role-based inbox treated like a provider inbox. | Reply rate per 100 delivered by email type. |
| Verification method (email) | How you validated the email operationally — delivered, bounced, confirmed by reply. A workflow field, not a future guarantee. | Prioritize channels that have worked; suppress known bad paths. | Method not stored; “present” treated as “verified.” | Deliverability rate per 100 sent, segmented by verification method. |
| Last verified | Timestamp for when a contact path was last confirmed by an outcome — a connected call, delivered email, or reply. Distinct from last updated. | Prioritize fresher, outcome-confirmed paths; set refresh rules. | Confused with last updated; timestamp only reflects import date. | Outcome metrics segmented by last-verified buckets. |
| Confidence score (internal) | An internal, explainable scoring field used to rank contact paths for routing. | Decide which phone/email to try first; reduce wasted attempts. | Score treated as “accuracy” and never recalibrated as data ages. | Connect/deliverability outcomes by score band. |
| Practice location | Address or location metadata tied to a practice site — not necessarily where the provider lives. | Territory assignment; local market prioritization; call-window planning. | Provider works multiple sites; billing-only address used for routing. | Reassignment rate per 100 assignments; “wrong market” flags. |
| Multi-location count | Count of distinct practice sites tied to the provider record. | Decide whether to run multi-site outreach vs. single-path. | Collapsed locations hide the best contact path. | Connect rate per 100 dials, multi- vs. single-location. |
| Do-not-contact (global) | Flag indicating a record should not be contacted through any channel. | Hard suppression across all tools. | Duplicates bypass suppression; flag not synced to dialer/email tool. | Suppression enforcement rate per 100 attempted dials/sends. |
| Channel opt-out | Channel-specific suppression — do not email, do not call — with a timestamp. | Respect preferences; protect deliverability and brand. | Opt-out stored only in notes; not enforced automatically. | Opt-out recurrence rate per 100 contacted records. |
| Suppression reason | Enum describing why suppression exists — requested, bounced, wrong number, policy. | Auditability; faster remediation. | Free-text reasons break reporting. | Reason distribution per 100 suppressed records. |
| Source category | High-level label for where a field’s value came from — registry, practice site, feedback, outreach. | Trust weighting; debugging when a source starts decaying. | Source missing; mixed sources without traceability. | Outcome metrics segmented by source category. |
| Last updated | Timestamp for when the field value was last refreshed or changed in your system. | Prioritize fresher paths; explain performance shifts. | Timestamp reflects import date, not an actual refresh. | Outcome metrics segmented by last-updated buckets. |
Canonical metric definitions: Connect Rate = connected calls / total dials. Answer Rate = human answers / connected calls. Deliverability Rate = delivered emails / sent emails. Bounce Rate = bounced emails / sent emails. Reply Rate = replies / delivered emails.
Weighted checklist for scoring a dataset or integration
Score each item 0–2, multiply by weight, and total it. Use the notes column to record what you actually verified rather than what you assumed.
| Category | Check | Weight | Score (0–2) | Notes |
|---|---|---|---|---|
| Definitions | Every contact field has a written definition, allowed values, and null rules. | 5 | ||
| Multi-path storage | Phone and location fields support multiple values — no overwriting direct dial with main line. | 5 | ||
| Provenance | Fields include source category and last-updated timestamps. | 4 | ||
| Freshness vs. verification | Last updated and last verified are distinct and used in routing rules. | 4 | ||
| Phone usability | Phone fields distinguish direct dial vs. main line, plus line-tested status/date. | 5 | ||
| Email usability | Email fields support deliverability QA and suppression. | 4 | ||
| Identity | NPI supports de-duplication and matching, with rules for edge cases. | 5 | ||
| Segmentation | Taxonomy supports specialty routing, including multi-taxonomy handling. | 3 | ||
| Suppression | Global and channel opt-outs are enforced across all outbound systems. | 5 | ||
| Feedback loop | Structured dispositions feed back into suppression and routing weekly. | 4 |
The trade-off is straightforward: richer field semantics — multi-path phones, timestamps, source categories — mean more integration work up front, but they cut wasted outreach and make QA defensible instead of anecdotal.
Outreach templates
These assume professional, compliant recruiting outreach. Keep messages short, easy to decline, and check suppression before sending.
First-touch voicemail (direct dial)
“Hi Dr. [Last Name], this is [Name] with [Org]. I’m calling about a [role] opportunity in [market]. If you’re open to a quick, confidential conversation, call me at [number]. If you’re not interested, tell me and I’ll stop.”
Email (deliverability-first)
Subject: Quick question — [specialty] role in [market]
“Dr. [Last Name] — I recruit physicians for [Org]. Are you open to a brief call about a [role] position in [market] — schedule, comp model, call expectations? If yes, what’s the best time this week. If no, I won’t follow up.”
Gatekeeper-friendly main line script
“Hi, I’m trying to reach Dr. [Last Name] regarding a professional opportunity. What’s the best way to send a message for review — email or fax — and who should I address it to?”
Ops feedback loop (data correction and suppression)
“We attempted outreach for Dr. [Last Name] and got [wrong number/bounce/opt-out]. Please (1) suppress the invalid channel with a reason and timestamp, and (2) confirm the best professional contact path for future outreach.”
Common pitfalls
- Single-field normalization that destroys routing. Collapse all phones into one field and you lose the ability to prioritize direct dial vs. main line — and to explain outcomes afterward.
- Missing timestamps. Without last updated and last verified, you can’t tell “bad data” apart from “old data.”
- Suppression not enforced across tools. If opt-outs live only in notes, duplicates keep getting contacted.
- Confusing identity with reachability. NPI and taxonomy help match and segment records; they don’t guarantee you can actually reach the provider.
- Mini-case: duplicates bypassing suppression. A common failure is email-only suppression — Record A is suppressed for an email opt-out, but Record B, same NPI with a different email or phone, still gets dialed. The fix: normalize identity around NPI where applicable, enforce a global do-not-contact flag at the identity level, and apply channel opt-outs across every contact path tied to that identity.
How to improve results
Use canonical metric definitions so reporting is comparable
- Connect rate = connected calls / total dials (per 100 dials).
- Answer rate = human answers / connected calls (per 100 connected calls).
- Deliverability rate = delivered emails / sent emails (per 100 sent).
- Bounce rate = bounced emails / sent emails (per 100 sent).
- Reply rate = replies / delivered emails (per 100 delivered).
A QA sampling protocol that avoids invented numbers
- Pick a cohort. One specialty, one market, one week of outreach.
- Sample records. Pull a fixed-size sample you can realistically review weekly, including both phone and email where available.
- Log outcomes with structured dispositions. Wrong number, gatekeeper, voicemail, provider answered, bounced, replied, opted out.
- Compute metrics with denominators. Per 100 dials and per 100 connected calls for phone; per 100 sent and per 100 delivered for email.
- Segment by field values. Compare direct dial vs. main line, recent vs. older line-tested status, email type, last-verified buckets.
- Change one rule at a time. Adjust routing order, suppression rules, or freshness thresholds, then re-measure on the next cohort.
Field naming and storage conventions
- Use enums for types. phone_type = direct_dial | main_line | fax | unknown.
- Store multi-values explicitly. Arrays or child tables for phones, emails, and locations — not a single overwritten field.
- Separate timestamps. last_updated (system refresh) is not the same as last_verified (confirmed by outcome).
- Keep suppression structured. do_not_contact_global (boolean) plus channel-level opt-out fields with timestamps and reasons.
The three-layer contact path rule
This is a fast fix for field-level credibility problems in recruiting CRMs.
- Layer 1: routing number (practice main line) — for gatekeeper workflows and confirmation.
- Layer 2: priority number (direct dial) — for first attempts when present.
- Layer 3: learned best channel (from dispositions) — the last successful channel and timestamp, e.g. “provider answered on main line ext 214.”
Store all three and you stop overwriting good paths, and your QA becomes explainable — you can see which layer actually produced the outcome.
Legal and ethical use
- Legitimate recruiting only. Don’t repurpose contact fields for unrelated marketing.
- Respect opt-outs and preferences. Suppress quickly and across all systems.
- Minimize data. Store what recruiting operations actually need; avoid collecting sensitive personal details.
- Not legal advice. Align your process with applicable laws and your organization’s own policies.
Evidence and trust notes
Definitions are only useful if they’re measurable and auditable. For how Heartbeat.ai defines and evaluates accuracy and outreach metrics, see the accuracy and metrics definitions methodology.
Primary references for provider identifiers and classifications:
- NPI Registry (CMS)
- CMS: National Provider Identifier Standard
- National Uniform Claim Committee (NUCC) taxonomy
Related field explainers:
- What is a direct dial number? (recruiting use cases)
- Mobile vs VoIP: how to tell (and why it matters)
- Prescriptive authority data for recruiters (field definitions)
FAQs
What should a provider contact data dictionary include?
Field name, definition, allowed values/format, null rules, provenance (source category), freshness (last updated), intended workflow use, and a QA metric tied to outcomes.
How do ops teams compare two datasets with different field names?
Map both to shared concepts — identity, segmentation, phone, email, location, suppression, provenance — then compare using the same denominators: per 100 dials, per 100 sent emails, and per 100 delivered emails.
What’s the difference between direct dial and line tested?
Direct dial describes intended routing — reaching the provider without going through a switchboard. Line tested describes callability at a point in time. You can have one without the other.
Which metrics should we track for phone and email fields?
For phone: connect rate and answer rate. For email: deliverability rate, bounce rate, and reply rate.
How should we handle opt-outs and do-not-contact flags?
Store them as structured fields with timestamps and reasons, and enforce them across every outbound tool. Free-text notes are not a substitute.
Next steps
- Operationalize the directory: pick your top 20 fields and copy these definitions into your internal schema docs.
- Instrument QA: make sure your dialer and email platform export the denominators needed for the canonical metrics.
- Go deeper on key phone fields: start with direct dial and mobile vs VoIP.
- Implement in your workflow: create a Heartbeat.ai account to map fields into a recruiting workflow with suppression and QA built in.
About the Author
Ben Argeband is the Founder and CEO of Swordfish.ai and Heartbeat.ai. With deep expertise in data and SaaS, he has built two platforms used by sales and recruitment professionals. Ben’s focus is helping teams find direct contact information for hard-to-reach professionals and decision-makers. Connect with Ben on LinkedIn.