Methodology
Last updated 10 August 2026
RivalScout publishes numbers about other companies, so it owes you an account of where each one comes from. This page describes every source we read, how each metric is calculated, how a change is scored, and the point at which a figure stops being a measurement and becomes an estimate.
What we collect
Collection happens on your watchlist’s schedule. Every fetch identifies itself, obeys robots.txt and rate-limits itself per host — see the crawler page for the full behaviour.
From a competitor’s own site
- Services pages. Service names a competitor advertises, and promotional offers.
- Changelogs and release notes. Dated entries and version tags.
- Careers pages. Open roles and their departments.
- Blogs and newsrooms. Post titles, dates and publishing rate.
- Homepage and positioning pages. The published wording, compared run to run.
From public third-party sources
- Google News. Articles naming the competitor.
- Google reviews. Rating, review count and recent review text from the competitor's Google business listing, via DataForSEO.
- GitHub. Public releases and commit activity, where a repository is public.
- Google Trends. Relative search interest for the tracked terms.
- App Store and Google Play. Public review text and ratings, where an app exists.
From sources you add
Custom collectors read a JSON endpoint, an RSS feed or a named element on a page. They run under the same fetch rules as everything else, and a collector that fails five times in a row is deactivated rather than left to retry indefinitely.
What we replace vs. what we read
The marketing site names common tools so teams can see what RivalScout covers. Below is the honest mapping: every named tool, the category it sits in, and the RivalScout collector that covers it today.
- Full. RivalScout reads a comparable public source directly.
- Partial. RivalScout covers part of the category from public sources, or proxies the signal from a related public source.
- With your own key. The data comes from a search-data provider you connect, not from RivalScout's own crawl.
| Category | Tool | RivalScout source | Coverage |
|---|---|---|---|
| Page monitoring | Visualping / Distill.io / Hexowatch | site collectors: services page, changelog, careers, page | Replaced |
| Services intelligence | Prisync / Price2Spy | services extraction + LLM normalisation | Replaced |
| Company and market | Feedly AI | Google News collector | Replaced |
| Search and SEO | Semrush | search metrics via your own DataForSeo key | With your own key |
| Search and SEO | Ahrefs | search metrics via your own DataForSeo key | With your own key |
| Search and SEO | SpyFu | search metrics via your own DataForSeo key | With your own key |
| Company and market | Owler | Google News + careers page + custom collectors | Partial |
| Company and market | Crunchbase | Google News + careers page + custom collectors | Partial |
| Company and market | Exploding Topics | Google Trends collector | Partial |
| Social and sentiment | Brandwatch | Google reviews + App Store review sentiment | Partial |
| Social and sentiment | Sprout Social | Google reviews + GitHub activity | Partial |
| Social and sentiment | BuzzSumo | Google reviews + news coverage | Partial |
| Sales enablement | Crayon / Klue / Kompyte | win/loss deals panel + digest reports | Partial |
What we do not collect
- Nothing behind a login, a paywall or an access control of any kind.
- Nothing from a site whose robots.txt disallows our crawler.
- No review-site scraping. Sentiment does not use G2, Trustpilot or Capterra; it uses public discussion and public app-store reviews.
- No purchased personal data, and no personal profiles of a competitor’s staff.
Measured, modelled and estimated
Measured covers anything read directly off a public page: a published service name, an offer, a release date, an open role, a post and its score. If we show it, we stored the page it came from.
Modelled covers indices built from measured values — press momentum, social engagement, hiring index — where the number is a ratio against that competitor’s own trailing baseline rather than a quantity anyone publishes.
Estimated covers search data: traffic, rankings, keyword gaps, referring domains and domain authority. None of it is derivable from public pages. It appears only when your organisation connects its own search data provider key, and it is labelled as that provider’s estimate, never as a measurement from a competitor’s analytics.
Metric reference
These definitions are read live from the same table the product reads, so this list is exactly what your reports compute.
- Discussion sentiment (index)
Average sentiment of public discussion about this competitor over the last 7 days, on a -100 to +100 scale. Measures recent Google reviews and public App Store reviews collected this run. It is not a review-site rating and does not use G2, Trustpilot or Capterra.
Confidence: Classified by a language model on the text of each public item. Fewer than five items in the window, or fewer than three runs of history, is marked low confidence.
- Hiring index (index)
Open roles now, indexed against this competitor's own trailing 90-day average, with a departmental breakdown. 100 is normal.
Confidence: Counts roles listed on the careers pages we track. Fewer than three runs of history is marked low confidence.
- Press momentum (index)
Articles about this competitor in the last 7 days, indexed against its own trailing 90-day weekly average. 100 is normal.
Confidence: Needs at least 90 days of collection before the baseline is meaningful. Fewer than three runs of history is marked low confidence.
- Price position (%)
The competitor's lowest paid plan as a percentage of the median lowest paid plan across this watchlist. 100 is the middle of the set.
Confidence: Computed only where price extraction confidence is high. Competitors without a readable public price are excluded.
- Pricing volatility (count)
Number of price changes observed in the last 180 days.
Confidence: Only counts changes to prices published on public pricing pages.
- Review rating (rating)
The star rating on this competitor's Google business listing at the time of the run, as reported by DataForSEO. It is the listing's own lifetime average, not a value we compute.
Confidence: Read from the Google business listing through DataForSEO. With fewer than three runs behind it, we mark it as thin evidence.
- Review volume (count)
The number of reviews on this competitor's Google business listing at the time of the run, as reported by DataForSEO.
Confidence: Read from the Google business listing through DataForSEO. Movement between runs is the signal; the absolute total is the listing's lifetime count.
- Shipping velocity (count)
Releases and changelog entries published in the last 30 days.
Confidence: Only counts releases published where we can read them. A private changelog reads as zero.
Every metric carries a confidence marker. Thin history — fewer than three runs, or too few items in the window — is shown as low confidence rather than presented as settled.
How changes are scored
Each detected change gets a materiality score from 0 to 100, built from four components:
- Signal class. What kind of change it is, weighted as below and capped at 40 points.
- Magnitude. How large the move is relative to the previous value.
- Persistence. Whether the same direction has held across consecutive runs, so a one-off blip does not outrank a sustained trend.
- Proximity. Whether the competitor appears in deals you recently lost. If you have imported win/loss records, changes at competitors that are beating you score higher.
Default class weights:
- Service added or dropped — 40. An advertised service appeared or disappeared.
- Promotional offer changed — 32. A promotion started, changed or ended.
- Product release — 30. A release, changelog entry or tagged version shipped.
- Positioning or homepage change — 25. The homepage or another positioning page was rewritten.
- Hiring surge — 20. Open roles moved, in total or within one department.
- Press spike — 20. News coverage moved against its own recent baseline.
- Social spike — 15. Review activity moved against its baseline.
- Blog cadence — 10. Blog or content publishing rate changed.
- Other — 5. Anything the classifier does not recognise.
An owner can retune these in settings. Tuning expresses the ratio between classes, not a new ceiling: weights are normalised against the largest, so raising everything raises nobody.
How AI is used
Language models do three jobs: extracting structured values from messy pages, summarising a change in a sentence, and drafting recommendations. They do not decide what changed — detection is deterministic comparison between two stored snapshots.
Every claim in a digest must cite a specific measured change. A draft is validated against the run’s scored changes before publication, and any draft that invents a change, a number or a competitor is rejected. When a draft fails, the run publishes a deterministic digest assembled directly from the scored changes instead. An ungrounded digest is never published.
Reproducibility
Every fetch is stored as a compressed, hash-addressed snapshot, and every figure in a report links back to the snapshot it came from. Two people opening the same report see the same numbers, and a number from three months ago can still be traced to the page that produced it. Snapshots are retained for 180 days.
Limitations
- Collectors read published pages. When a site is restructured a collector can miss data until it is updated, and a gap is shown as a gap rather than filled in.
- Sentiment is a model’s reading of public discussion. It is not a survey, and a small number of loud posts can move it.
- Indices are relative to a competitor’s own history, so they mean little in the first 90 days of tracking.
- A competitor can change something we cannot see: an unadvertised service, a private beta, a deal negotiated in a room. Absence of a signal is not evidence of no change.
- Search estimates are the provider’s model, with the provider’s error.
Corrections
If a figure here looks wrong, tell us and we will trace it back to its snapshot: hello@rivalscout.io.