Outcome
A five-stage pipeline cut a 500-paper dermatology review from 45–60 hours to roughly 8–12 hours of reviewer time — about 80% less handling time, with every claim still traceable to a retained, auditable source.
Context
Repeated evidence reviews were consuming the time needed for judgement.
Product and patient-facing decisions across eczema, acne, psoriasis and topical steroid withdrawal required repeated reviews of a large and changing evidence base. Each review involved collecting papers, finding full text, removing duplicates, applying consistent inclusion rules and turning the retained evidence into something that could be compared.
The goal was not to automate clinical judgement. It was to remove repetitive handling while making the evidence trail clearer for the person responsible for that judgement.
Problem
The sources did not arrive in a reviewable form.
A dermatology review can involve hundreds of papers from several databases. The same paper may appear more than once. Full text may be missing. Titles and abstracts use inconsistent terms. Public patient discussions add useful context but cannot be treated as clinical evidence.
Doing collection, screening, tagging and synthesis by hand was too slow for repeated research across eczema, acne, psoriasis and topical steroid withdrawal.
My role
I designed the workflow, evidence schema and review safeguards.
I built the five-stage Python pipeline for collection, deduplication, screening, analysis and report preparation. I defined the screening rubric, the structured evidence fields, the separation between academic and anecdotal sources and the checks that prevent draft synthesis from becoming an untraceable answer.
System
A five-stage path from search query to review-ready evidence.
- Run defined Boolean searches against Semantic Scholar and PubMed. Parse PubMed XML and collect titles, abstracts, authors, publication details, citation counts and open-access links.
- Collect public Reddit posts through a separate route for patient-reported experience. These records are labelled as anecdotal data and never merged with trial evidence.
- Use fuzzy title and author matching to remove duplicate papers returned by different services.
- Look for open-access full text through Unpaywall, Europe PMC and CORE. Use browser automation for accessible download pages and store files in authorised Google Drive storage.
- Run a first screening pass that scores relevance from 0 to 2 against a written rubric. Low-scoring papers are excluded with the reason recorded.
- Extract text from retained PDFs with PyMuPDF. A second pass records condition, symptoms, intervention category, treatment theme, evidence strength and reported effect in a fixed schema.
- Run a similar but separate screen over public patient reports to record common concerns and treatment experiences.
- Convert each paper into a numeric text representation called an embedding. UMAP reduces the number of dimensions so the data is easier to process. HDBSCAN then groups papers with similar text without requiring a fixed number of groups in advance.
- Calculate a silhouette score to measure how clearly the groups are separated and flag weak clusters for review.
- Write the screened records and labels to Airtable, then build thematic tables across condition, intervention and evidence strength.
- Generate draft treatment summaries only from the screened dataset. Every claim remains tied to the retained evidence and requires human review.
Key decisions
Different kinds of evidence stay different.
Public patient reports can reveal recurring concerns and language that formal research may miss, but they are not clinical evidence. The system therefore collects and labels them through a separate route and never merges them with trials or academic papers.
Automated screening is also treated as a first pass. Exclusion reasons are recorded, weak clusters are flagged and every generated summary remains tied to the screened source record for human review.
Outputs
The output is an auditable evidence workspace.
Reviewers receive a deduplicated set of papers, source and access links, inclusion decisions, structured labels, thematic groupings and draft summaries whose claims can be traced back to the retained material. Patient-reported themes remain visible in a distinct dataset.
Results
Automation reduced handling time without pretending to replace review.
Preparing and screening a 500-paper review manually would take 45 to 60 hours. Automated collection, deduplication and first-pass screening reduce the work to roughly eight to twelve hours of reviewer time.
That is about an 80% time saving. Every retained paper has a source record, inclusion reason and structured evidence fields, which makes the final review easier to audit.
Lessons
Traceability is more important than a fluent summary.
A convincing paragraph is not a useful clinical output if its source cannot be inspected. The most important design decisions were therefore the fixed schema, recorded exclusion reasons, source links and explicit boundary between automated preparation and human judgement.
Clustering also works best as navigation rather than truth. It helps a reviewer see themes across a large corpus, while weak or ambiguous groups remain visible instead of being forced into a neat taxonomy.
Why it matters
The same approach can make evidence-heavy workflows faster and safer.
This work demonstrates how I combine clinical context, structured data and practical automation: reduce repetitive work, preserve provenance and design the system around the person who remains accountable for the final decision.
Skills & tools
What it took to build this.
- Literature search design
- Fuzzy deduplication
- Screening rubric design
- PDF text extraction
- Structured evidence schema
- Embeddings & clustering (UMAP/HDBSCAN)
- Model-assisted synthesis
- Source traceability
- Human-in-the-loop review
- Separating academic vs anecdotal evidence
- Focus
- Eczema · Acne · Psoriasis · TSW
- Status
- Operating in production