Methodology · Scoring

How Authority and Safety are scored

Two numbers, each backed by mathematics we publish in full: a Bayesian evidence engine over global popularity, decades of corroborated history, DNS and TLS posture, and threat intelligence. When the evidence is thin, we say “not enough signals yet” — we never dress up ignorance as a low score.

Five design principles

A score is a posterior, not a sum

Checklist scorers add points for whatever they happen to see, so a site nobody has measured looks identical to a site measured and found empty. AboutUs treats every signal as evidence updating a prior belief. Missing evidence widens the uncertainty interval — it never silently scores as zero, and it never silently scores as fine.

Influence is proportional to the cost of faking

A factor's maximum weight tracks what it costs an adversary to forge it. Global popularity rank and years of corroborated history cost years and large sums to fake, so they dominate. A TLS certificate and clean email records can be configured in an afternoon, so they are worth a few points at most — real, but bounded.

Gates for facts, sums for degrees

Hard evidence of harm — an active threat-feed listing, a registry hold — caps the safety band outright, no matter how tidy the site's hygiene is. Points cannot buy back a fact. Everything else is compensatory and bounded.

Every displayed point maps to a sentence

Each score decomposes into labeled reason lines whose points sum exactly to the number on screen, including an explicit calibration-adjustment line. If we can't explain a point, we don't award it.

Correlated signals fuse before they weigh

Tranco rank, Open PageRank, and corpus link counts all measure the same thing — prominence. They combine into one factor before weighting, so being popular never counts three times.

The observation model

Observed, absent, or unknown — never conflated

Every signal enters the engine in one of three states. Observed: the check ran and returned a value. Observed-absent: the check ran and confirmed the thing is not there — absence is information, and it scores. Unknown: the check failed, was rate-limited, or is not yet wired — it carries zero weight, cannot help or hurt the score, and widens the confidence interval instead. This distinction is why an attacker cannot improve a score by making our lookups fail, and why a site we simply haven't measured is never labeled a low-authority site. Every signal's weight also decays with staleness — a 45-day-old rank observation counts less than yesterday's.

Authority

Five factor families, weighted by the cost of faking them

FamilyWeightSignals (fused within the family)Cost to fake
Prominence52%Tranco global rank · Open PageRank link-graph score · deduped corpus inbound links — fused as corpus percentiles, two independent sources required for full evidence weight$10⁴–10⁶ and months of sustained traffic to fake
Verified continuity30%Corroborated verified-years plus infrastructure stability (same registrant/IP/nameservers/mail across a snapshot archive since 2006, distinctiveness-weighted) and operational continuityYears — and the drop-catch rule breaks shortcuts
Identity corroboration4%Company-registry match (Companies House, EDGAR, GLEIF) — display-first and hard-capped at 4 pointsA £12 shell company exists; that is why the cap exists
Security & email hygiene8%TLS validity and maturity · DNSSEC · MX/SPF/DMARC. Additionally, broken HTTPS caps Authority at 85 — a ceiling, not a deductionNearly free to configure — hence a small, bounded block
Prominence × continuity6%An interaction term rewarding sites that are both prominent and have been so for yearsRequires both of the expensive things at once

Verified-years: history that has to agree with itself

Continuity is measured as verified-years: the span over which independent history lines — the domain registry (RDAP), the AboutUs record kept since 2006, the first web-archive snapshot, and the first certificate in public CT logs — agree within ±2 years. One line alone counts at less than half strength; agreement is the point. And the drop-catch rule: if the registry shows a recent re-registration, verified-years resets no matter how old the archives are — a freshly bought domain's history vouches for the previous operator, not the current one. Registrant changes and visible content discontinuities likewise reset the span (they remain in the record as dated facts — a reset is not a penalty, it is honesty about who the history belongs to). Continuity is further corroborated by infrastructure stability — years of the same registrant, IP, nameservers, and mail in our snapshot archive since 2006, each weighted by how distinctive and hard-to-fake it is. How we measure continuity →

Aggregation, calibration, and confidence

Within each family, evidence combines as a mass-weighted posterior shrunk toward a low prior — the null hypothesis is “not notable,” and thin evidence stays close to it. Families combine by the published weights, and a fixed monotone calibration map (anchored so that google.com = 99, an established corroborated business ≈ 87, a typical indexed site ≈ 40, a parked domain ≤ 6) turns the fused value into the model score. The Authority number published on a profile is that model score expressed as a percentile — the share of scored sites that fall below it — so a typical site reads near 50 rather than near 17. The ordering never changes, and the evidence panel on every profile still shows the reason lines that sum to the underlying model score. Because most scored sites cluster low, the percentile compresses at the very top: the strongest sites on the web all publish at or near 100. Confidence is computed from total evidence mass: below 0.35the score is withheld entirely (“not enough signals yet”), and above it the score ships with an uncertainty interval that narrows as evidence accrues. There are no tier cliffs anywhere — every transform is continuous, because a cliff is a purchasable boundary.

Safety

A hazard model, not a checklist

Safety events group into six families — threat feeds, deception, commerce, infrastructure, transport, email hygiene. Within a family, the worst event dominates and the rest contribute a quarter of their weight (max-fusion: ten blacklists listing one IP is one fact, not ten). Family hazards sum into a total H, and Safety = 100 − 100·(1 − e−H), capped at 99. Old incidents decay with per-family half-lives — a delisted domain heals.

EventSeverity (0–0.99)
Active threat-feed listing0.92
Registry hold (operational status, mild)0.05
Typosquat of a high-authority brand0.50
Anomalous infrastructure change (corroborated)0.28
Commerce claims without matching identity0.22
No valid HTTPS today0.16
High-abuse TLD0.10
No MX and no DMARC0.04

New is not unsafe — the invariant, made precise

Domain age contributes zero hazard. A two-month-old legitimate site with clean checks scores 99. But age multiplies independently-observed hazards in the deception families (×1.35 under 30 days, ×1.2 under 90, ×1.08 in the first year): new ≠ unsafe, but new + phishing evidence is worse than old + the same evidence.

Honest bands

“Clear” is only claimable when a harm source was actually checked — good hygiene with zero threat checks reads “No known flags (partially checked).” Hard harm evidence gates the band to “Known risk” before anything else is considered. And tiny residual risk (a clean small business without corporate email records) displays as Clear with its evidence lines still shown — a mild footnote must not read as a warning forever.

What deliberately scores nothing

Restraint is part of the method

  • Domain age by itself — new is not unsafe; age only amplifies independently-observed deception evidence
  • Claim status — whether an owner has claimed a page is a workflow state, not evidence about the website
  • Breach-history mentions — shown as dated context in the record, never scored
  • The wiki's historical self-descriptions — displayed as dated history (Lane P), never as current fact
  • Popularity on AboutUs itself — self-reference is not evidence

Known limits

What this model cannot yet see

A site that serves our checks clean pages while serving victims something else (cloaking) defeats every content-derived negative; until fetch-diversity checks land, such a site reads “partially checked” — never “Clear.” A purchased aged domain that keeps its registrar and content continuity can survive the drop-catch rule. We monitor the score distributions themselves for drift, because any published score invites gaming — and we would rather document the limits than pretend they don't exist. Scores are statistical evidence summaries, not verdicts, warranties, or accusations; the evidence lines under every score are the product.

Authority evidence draws on Open PageRank link-graph data and the Tranco research ranking.