Outcome
A scheduled monitor across 30+ communities cut weekly review time from 5–7 hours to about 20 minutes a day — roughly 4 hours saved every week, with urgent opportunities caught within one polling cycle instead of days later.
Context
Useful market signals were appearing faster than they could be reviewed.
People regularly discuss their productivity problems, current tools and buying decisions in public online communities. Those conversations can improve product decisions and reveal time-sensitive opportunities, but only if the relevant ones can be found without spending the day scanning feeds.
Problem
The signal was distributed across more than 30 communities.
Useful product feedback and sales conversations were spread across more than 30 Reddit communities. Checking them by hand took too long. Urgent posts, such as a person asking for a recommendation, could become stale within hours.
My role
I designed, built and deployed the monitoring system.
I defined the source and keyword strategy, built the collection and ranking pipeline, designed the structured language-model output, created the immediate-alert and digest formats and deployed the unattended job on a Linux server.
I also added the unglamorous parts that make scheduled work dependable: deduplication, short-lived caches, file locks, retries, validation, structured logs and failure alerts.
System
From public conversation to a ranked research queue.
- Poll more than 30 subreddits through the Reddit API. Focused communities are read broadly. General communities are filtered using competitor and category terms.
- Load each candidate post and its full public comment tree so the ranking uses the surrounding discussion.
- Remove posts already processed and discard cached items after 24 hours.
- Ask a language model for a strict JSON result containing High, Medium or Low relevance, an urgency flag and one short reason. Validate the returned fields before use.
- Combine that result with upvotes and comment count to produce a clear ranking score.
- Send urgent recommendation requests as immediate alerts. Cache the remaining items for digest emails at 9 AM and 9 PM UTC.
- Build a readable HTML email with source links, the reason each item was selected and response ideas for manual review.
- Protect unattended runs with file locks, exponential retry, structured logs and a failure email if the job stops.
- Deploy through a repeatable shell script that sets up an Ubuntu server, Python environment, restricted credentials and cron schedule.
Key decisions
Ranking uses context, not a keyword match alone.
A mention of a productivity term does not necessarily contain a useful signal. The system therefore reads the public comment tree around each candidate and combines a validated relevance classification with engagement data.
Urgency and importance are handled separately. Recommendation requests can trigger an immediate alert, while routine research is deliberately batched into two predictable reviews so the monitor does not create another source of distraction.
Outputs
Two outputs support two different decisions.
Urgent findings arrive as individual alerts while they are still timely. Everything else is grouped into a concise HTML digest containing the source, the reason it was selected and draft talking points for manual review.
Results
The system converted continuous scanning into scheduled review.
Scanning and sorting more than 30 communities would take five to seven hours each week. The two daily digests reduce that to about 20 minutes of review a day, saving roughly four hours each week.
Urgent recommendation requests arrive within one polling cycle. Routine findings are grouped into two predictable reviews instead of interrupting the day.
Lessons
An intelligence system should reduce noise, not redistribute it.
Collecting every possible mention would have produced a larger but less useful inbox. The combination of community-specific search rules, contextual ranking, deduplication and digest timing mattered more than raw collection volume.
Strict structured output also made the language-model step operationally useful. A small validated schema was easier to rank, log and recover than an open-ended narrative response.
Why it matters
The pattern transfers to any public market with too much weak signal.
This project shows how I turn an interrupt-driven research task into a reliable operating rhythm: collect broadly, rank transparently, escalate only what is urgent and keep a person responsible for the final response.
Skills & tools
What it took to build this.
- Scheduled research pipelines
- Classification & ranking
- Structured LLM output validation
- File locking
- Exponential retry
- Structured logging
- Failure alerting
- Linux server deployment
- Cron scheduling
- Credential management
- Focus
- Productivity-software research and time-sensitive public market monitoring
- Status
- Operating in production