Outcome
Manual research and CRM entry fell from 15–25 minutes per clinic to 2–4 minutes of final review — roughly 250 hours saved for every 1,000 clinics processed, across more than 40 US metros and several UK regions.
Context
Provider research had become an operational bottleneck.
The work required a reliable view of healthcare organisations across more than 40 US metro areas and several UK regions. The useful facts existed, but they were scattered across maps, company databases, clinic websites and an existing CRM.
A researcher could assemble one record manually. Repeating that process across hundreds or thousands of organisations made coverage inconsistent, expensive and difficult to audit.
Problem
Finding a clinic was easier than deciding whether the record could be trusted.
Organisation information was spread across maps, company databases, public clinic websites and an existing CRM. Records were duplicated, incomplete or out of date. Researching each organisation manually did not scale across more than 40 US metro areas and several UK regions.
My role
I designed and built the pipeline end to end.
I defined the research workflow, connected the source APIs, built the deduplication and rejection rules, wrote the web crawler, designed the structured extraction schema and integrated the approved output with Attio.
I also built the operational safeguards: confidence scoring, source retention, checkpoints, rate limits, retries and field-level failure handling. This public case study intentionally excludes individual-level contact handling and confidential implementation details.
System
From discovery to a review-ready CRM record.
- Search for clinics through Google Places and Apollo across defined regions. Use field masks and pagination so large searches return only the data needed.
- Normalise website domains and merge duplicates from different sources. Compare them with existing Attio records before creating anything new.
- Reject parked domains, directory listings, veterinary practices and other false matches. Check website health concurrently across the remaining list.
- Enrich organisations with employee count, estimated revenue and technology data from Apollo.
- Map each clinic website with Firecrawl and a custom breadth-first crawler. The crawler handles standard links, JavaScript onclick targets and meta-refresh redirects.
- Use roughly 60 rules to rank pages likely to contain services, ownership details and organisation information before scraping them.
- Extract structured organisation facts from the selected pages. Each field includes its source and a confidence score. Retained analysis is anonymised and this case study omits any person-level detail.
- Score record completeness with code so reviewers can focus on records with missing or conflicting evidence.
- Write approved records to Attio through its API. If a field is invalid, remove that field, log the reason and retry the rest instead of losing the whole record.
- Use checkpoints, token-bucket rate limits and retry rules so large regional batches can resume after an API or network failure.
Key decisions
Accuracy came before automation.
The pipeline does not treat every search result as a healthcare provider. It rejects parked domains, directories, veterinary practices and other false matches before spending time on enrichment or extraction.
Each retained fact carries its source and a confidence score. The system prepares records for review rather than silently presenting uncertain data as fact. Invalid CRM fields are isolated and retried without discarding the rest of an otherwise useful record.
Outputs
The reviewer receives evidence, not just a populated row.
Each proposed organisation record includes normalised identity data, enrichment fields, relevant website findings, source links, confidence flags and a completeness score. Missing or conflicting evidence is made visible so review time is spent on genuine exceptions.
Results
More coverage with a smaller review burden.
Manual research and CRM entry takes 15 to 25 minutes per clinic. The pipeline reduces this to two to four minutes of final review, saving about 80% of the time.
Across 1,000 clinics, that equates to roughly 250 hours saved. Records arrive with organisation details, source links and confidence flags ready for review.
Lessons
The hard part was entity quality, not scraping.
Collecting more data did not automatically create a better record. Domain normalisation, duplicate detection, rejection rules and traceable sources created more value than adding another enrichment provider.
Large research jobs also need to be resumable. Checkpoints and isolated failures turned a fragile script into an operating system that can survive rate limits, malformed fields and temporary network problems without restarting an entire region.
Why it matters
The pattern applies wherever healthcare-market data is fragmented.
This project demonstrates how I approach a high-friction operational workflow: define what a trustworthy output looks like, automate the repeatable stages, preserve the evidence and keep a human focused on the decisions that remain uncertain.
Skills & tools
What it took to build this.
- Web crawling & discovery
- Entity resolution
- Deduplication
- Structured extraction
- Checkpointing
- Rate-limited retries
- Confidence scoring
- Resilient batch processing
- CRM integration (Attio)
- API integration at scale
- Focus
- Healthcare organisation discovery, research and CRM preparation
- Status
- Operating in production