How reports are generated
A detailed explanation of the HoloRadar research pipeline — from source collection to structured report output.
Source Selection
HoloRadar collects content from public platforms where genuine discussion happens. Sources are selected based on their relevance to the query topic and scored for quality before inclusion. Platform availability varies by subscription tier — the report's source distribution section discloses exactly which platforms contributed to any given report and how many items came from each.
Ranking
Collected items are ranked using three primary factors: recency (more recent posts score higher within the 30-day window), engagement (replies, shares, and interactions signal community relevance), and text quality (coherent, substantive posts score higher than low-effort or spam-like content). Ranking determines which items receive more weight in the synthesis step.
Duplicate Detection
The same story or opinion often surfaces across multiple platforms. HoloRadar applies cross-source deduplication to avoid over-representing a single piece of content that was shared widely. Items that are substantially identical in meaning are grouped and counted once rather than inflating cluster sizes with repetition.
Bot and Low-Quality Filtering
Automated, bot-generated, or coordinated inauthentic content is filtered before synthesis begins. Known patterns of low-quality content — including spam, automated reposts, and engagement manipulation — are excluded. This filtering is imperfect; the confidence score and limitations section of each report disclose when noise or ambiguity in a topic may affect reliability.
Evidence Clustering
Related discussions are grouped into evidence clusters based on shared meaning, topic, and recurring framing. A cluster represents a recurring theme in the discussion — not just a keyword match. Each cluster includes an item count and a representative paraphrase. Clustering makes patterns visible that would be invisible when reading individual posts. Clusters are not sentiment-scored; they represent discussion activity.
Representative Sources
Each report includes representative voices — paraphrased excerpts that capture the most recurring patterns within the evidence clusters. These are not verbatim attributed quotes. They are distilled representations of recurring language and framing observed across many independent sources. This distinction is disclosed in every report.
Supported Platforms
The research engine currently supports the following platforms (availability varies by tier): Reddit, X, YouTube, Hacker News, GitHub, Blogs and Articles, Forums, TikTok, Bluesky, Threads, LinkedIn, Pinterest, Instagram, Perplexity, and Polymarket. New platforms are added as the engine expands. The source distribution section of each report shows exactly which platforms were analyzed.
Confidence Score
Every report carries a confidence rating of high, medium, or low, accompanied by a plain-language explanation. Confidence reflects factors including total source volume, platform diversity, topic clarity, and clustering ambiguity. A high-confidence report means the topic is well-defined, well-represented across multiple platforms, and the clustering is clean. A low-confidence report means one or more of these conditions is compromised — the report is still useful but should be interpreted with that context.
Limitations
Every report discloses its limitations explicitly. Common limitations include: public discussion only (private conversations are not accessible), geographic and sector variation that is present but not broken out, the fixed 30-day window (historical periods are not covered), and the distinction between paraphrased representative voices and verbatim attribution. Limitations are not a weakness — they are a transparency feature that makes the report more trustworthy, not less.
A Note on Sentiment
HoloRadar does not produce sentiment percentages. The pipeline emphasizes engagement signals, evidence clustering, and confidence scoring rather than positive/negative classification. Sentiment classification on community discussion is inherently ambiguous — the same post can be sarcastic, nuanced, or context-dependent in ways that automated sentiment scoring distorts. Reports focus instead on what themes recur, how engagement is distributed, and what the evidence clusters contain.
Last updated: August 15, 2026
This methodology page describes the current state of the HoloRadar research engine. As the platform evolves, this page will be updated to reflect changes to source coverage, ranking factors, and report schema. If you have questions about how a specific report was generated, contact us at [email protected].