Documentation

Architecture notes

How the 39 tools are built, and what each answer is actually derived from

Overview

There are two execution models on this site and no third. Some tools ship their logic in the page and run entirely on your machine. The rest call a single Cloudflare Worker — a JavaScript isolate at the network edge — which performs the lookup, assembles the answer in memory and returns it in the same response. There is no database, no queue and no origin server behind it, so there is nowhere for a submitted value to be written even in principle.

Each tool's own page says which of the two it is. The rest of this document is what happens inside each path, in enough detail to check the work.

Tools that run entirely in your browser

The date and age calculators and the password entropy meter make no network request at all. The value you type is read by JavaScript already loaded in the page, the result is computed there, and the page never sends it anywhere. Open the network tab and there is nothing to see; disconnect and they keep working.

The date engine

Calendar arithmetic is fiddly rather than difficult, and the awkward parts are well known: leap days, months of unequal length, and the fact that "months between" means two different things to two different people. The engine computes in the proleptic Gregorian calendar, reports whole completed units, and states the counting convention alongside every result rather than leaving you to guess which one it used.

It is covered by an automated test suite that runs on every change, with cases for February 29 birthdays, month-end rollover, negative spans, and the difference between whole-unit and fractional counting. A calculator that is wrong in the edge cases is worse than no calculator, because you cannot tell from the output.

Password entropy

Strength is reported in bits of Shannon entropy over the character set actually used, which counts guesses rather than rewarding a capital letter and a digit bolted onto a dictionary word. The calculation happens in the page, so no password is transmitted — which is the only defensible way to build such a tool.

DNS and email record verification

The email tools read records that domain owners publish deliberately, over Cloudflare's DNS-over-HTTPS endpoint (cloudflare-dns.com/dns-query, application/dns-json). Nothing is sent to your mail server, no credentials are ever requested, and no record is cached or logged by us. These are public records readable by anyone, which is exactly why the checks work at all.

SPF — RFC 7208

The domain's TXT records are fetched and the one beginning v=spf1 is parsed. Every include, redirect, a, mx, ptr and exists mechanism it references is walked, counting DNS lookups as it goes:

The ten-lookup limit
RFC 7208 §4.6.4 caps a record at ten DNS-querying mechanisms. Exceed it and conforming receivers return permerror — the record stops working entirely. This is where most broken SPF records fail, and the count is reported whether or not it passes.
Policy strictness: -all vs ~all
Hardfail instructs receivers to reject unauthorised senders; softfail asks them to accept and mark. Both are valid; which one you want depends on how confident you are that the record is complete.
+all
Flagged as critical. It authorises every host on the internet to send as your domain, which is worse than publishing no record at all.

DKIM — RFC 6376

A DKIM public key lives at selector._domainkey.example.com, so a selector is required — there is no way to enumerate them from DNS, which is why the tool asks for one and suggests the common defaults (google, selector1, s1, k1, default). The record is parsed for k= (key type) and p= (the base64 public key), and the key length is derived from the decoded key so that a 1024-bit key can be distinguished from 2048-bit.

DMARC — RFC 7489

Read from _dmarc.example.com. The policy (p=none|quarantine|reject), the subdomain policy (sp=), the alignment modes (adkim=, aspf=), the percentage (pct=) and the report addresses (rua=, ruf=) are reported as published. p=none is not a failure — it is the correct first step, and the honest reading is "monitoring, not yet enforcing".

BIMI, MTA-STS and TLS-RPT

BIMI is read from default._bimi.example.com and checked for a logo URL and a Verified Mark Certificate; the SVG is validated against the Tiny P/S profile mailbox providers require and rendered through a standard image binding rather than by injecting its source, so a hostile SVG has nothing to execute. MTA-STS (RFC 8461) is read from _mta-sts.example.com and its policy file over HTTPS; TLS-RPT (RFC 8460) from _smtp._tls.example.com.

The link tools answer a narrower question than they are often assumed to: where does this actually go, and what do public records say about the place it lands? Four stages, in order.

1. Unwrapping the address

Tracking and cloaking layers are peeled off before anything else happens: base64-encoded path segments that decode to another URL, percent-encoded nested addresses, and ?url=-style redirect parameters. The unwrap runs up to five passes, because these layers are routinely nested, and it reports how many layers it removed rather than quietly normalising them away.

2. Following the chain

Each hop is fetched with redirect: 'manual' so the worker sees every Location header itself instead of letting the runtime follow them silently. The trace stops after ten hops, keeps a set of addresses already visited so a redirect loop is detected rather than exhausting the budget, and aborts any single hop that takes longer than four seconds.

Because a redirect can also be written into the page instead of sent as a status, a 200 response has its HTML scanned for <meta http-equiv="refresh"> and the trace continues to that address. This is the one point where page content is downloaded — and it is only pattern-matched as text. A Worker isolate has no DOM and no renderer, so nothing on the page executes, no image or script it references is requested, and no cookie it tries to set is stored.

3. Reading the public record

Once the final address is known, several independent sources are queried in parallel rather than in sequence — the answer arrives in roughly the time of the slowest one:

  • DNS metadata — MX, NS and TXT records for the registrable domain, over DNS-over-HTTPS.
  • Registration data — RDAP (RFC 7480) via rdap.org, which routes the query to the authoritative registry and returns machine-readable dates. This is where domain age comes from; it is read, not estimated.
  • Certificate data — the Certificate Transparency logs at crt.sh, falling back to a HEAD request against the host if the log query fails.
  • Reputation feeds — only when live feeds are enabled, and skipped explicitly otherwise, with "skipped" reported in the result rather than being passed off as "clean".

4. Scoring

The signals are combined into one number, described in full below.

How the risk score is calculated

The score is additive: each signal that fires contributes a fixed weight, the total is clamped to 100, and the verdict is a band on that total. Nothing is machine-learned, nothing is a black box, and every signal that fired is listed in the result with its own weight — so you can disagree with the arithmetic instead of trusting it.

These are all fifteen rules, with the weights the code actually uses:

Risk signals and their score weights
SignalWeightCondition
Live feed match+60An enabled reputation feed returns a block verdict for the final URL.
Inline credentials+50A username or password is embedded in the URL — the classic https://paypal.com@evil.tld masking trick.
Punycode / homograph+45The hostname is an internationalised name or mixes scripts, so it can render as a brand it is not.
Typosquatting+45The domain core is within a small edit distance of a known brand. The distance and the brand are both reported.
Subdomain brand impersonation+40A brand keyword appears in the subdomain of an unrelated registrable domain.
Raw IP as host+40The host is an IP literal, so there is no name to check and no certificate to match.
Newly registered domain+35RDAP reports a creation date less than 30 days ago.
Missing or invalid TLS+30No HTTPS, or no usable certificate for the host.
Algorithmic hostname+25Shannon entropy above 4.2 bits and hostname longer than 14 characters — both, because either alone has too many false positives.
No mail infrastructure+25A domain impersonating a brand that has no MX records at all.
Hostile DNS infrastructure+25An impersonating domain whose nameservers are free or dynamic-DNS providers, or whose TXT records carry high-entropy payloads.
Long redirect chain+20Three or more addresses in the chain.
Subdomain stacking+20Three or more subdomain levels, used to push the real domain out of view on a phone.
Brand keyword squatting+15Per matching brand keyword in the hostname, capped at +45 so keyword stuffing cannot dominate the score on its own.
Abused TLD+15The top-level domain is one of the registries with a disproportionate share of abuse.

The verdict bands

Five bands, not three — the middle of the range is where honest uncertainty lives, and collapsing it into "suspicious" throws away the distinction between a domain that is merely young and one that is actively impersonating a bank:

  • 0–15 Safe — nothing fired, or only something trivial.
  • 16–35 Low risk — one weak signal. Usually a young domain or an unusual TLD.
  • 36–60 Medium risk — worth reading the factor list before you click.
  • 61–85 High risk — several independent signals agree.
  • 86–100 Critical — a critical-weight signal fired, or many stacked.

Weights sum well past 100 in principle, which is deliberate: past a certain point the exact number stops mattering, so the total is clamped and the band is what is reported.

What these tools cannot tell you

Every tool here has a boundary, and a tool that hides its boundary is worse than one that has a narrow one. Plainly:

A clean result is not a safety guarantee
It means these specific checks found nothing. A brand-new phishing page on a well-aged domain with valid TLS and no lookalike name will score low, correctly, because none of the signals above are true of it. The score describes the address, not the intentions of whoever is running it.
A high score is not proof of malice
Legitimate sites do trip these signals. A startup on a two-week-old domain behind three redirects on a cheap TLD looks statistically identical to a campaign. That is why every factor is listed rather than just the total — the factor list is the part you can actually judge.
DNS answers can be stale
Records are cached across the internet for the duration of their TTL. A record you published minutes ago may not be visible yet, and one you deleted may still resolve. This is how DNS works, not a bug in the check — but it does mean "I just changed it" and "the tool is wrong" are easy to confuse.
The email checks read published policy, not delivery
A perfect SPF, DKIM and DMARC set-up says your records are correct. It does not prove a particular message will authenticate, because that depends on which server sent it and what happened in transit. Aggregate DMARC reports are what tell you that.
The calculators apply a stated convention
They are arithmetic, not advice. Where a convention is contested — how many months between two dates, how to count a corrected age — the tool names the convention it used, and that is the extent of the claim.

Data handling

The retention claim on this site is architectural rather than a policy promise. There is no database bound to the worker, no KV namespace, no queue and no log sink for submitted values; a domain, URL or header block exists as a variable inside a single request and is gone when the isolate finishes with it. Nothing needs to be trusted about our intentions, because there is nowhere for the data to go.

  • No accounts, so nothing to attach a history to.
  • No record of the domains, URLs, images, passwords, headers or dates you submit.
  • The in-browser tools transmit nothing at all — they work offline.
  • Reputation feeds are the one case where a URL reaches a third party, which is why they are opt-in and reported as skipped when they are off.

The privacy policy covers the rest, including the analytics and advertising the site does load, and what those see — which is page visits, not tool inputs.

Performance and caps

Running at the edge means the lookup happens in a data centre near you rather than in whichever region an origin server sits in. The practical limits worth knowing:

  • Ten hops maximum on a redirect trace, with loop detection and a four-second abort per hop.
  • Five passes maximum when unwrapping encoded or nested addresses.
  • Ten DNS-querying mechanisms is not our cap but RFC 7208's — SPF evaluation stops there because conforming receivers do.
  • External sources have their own timeouts, and a source that fails is reported as failed rather than silently treated as a pass.

Rate limiting exists to keep the tools available and is described in the terms. There is no public API; if you need one, say so.

Found something here that does not match what the tool does? That is the most useful mail we get — the contact page says what to include.