An IDN homograph attack requires zero technical access to your domain. The attacker doesn't need to compromise your DNS, your mail server, or your certificate. They simply register a different domain — one that displays identically to yours in every font that renders it. The Cyrillic letter "а" (U+0430) is visually indistinguishable from the Latin "a" (U+0061) at normal reading size. One character substitution can produce a domain that looks like apple.com, paypal.com, or amazon.com in any email client, browser URL bar, or chat message — pointing to a server entirely under attacker control.

1. The Technical Foundation: Internationalized Domain Names and Punycode

The legacy Domain Name System was designed for ASCII characters only — 26 lowercase letters, digits, and hyphens. To enable non-English speakers to register domains in their native scripts (Arabic, Cyrillic, Chinese, Devanagari, Greek, etc.), ICANN introduced Internationalized Domain Names (IDN) in 2003 through RFC 3490 and the subsequent Punycode encoding scheme (RFC 3492).

Punycode converts Unicode strings into an ASCII-compatible encoding (ACE) prefixed with xn--. A domain containing non-ASCII characters is stored in DNS and resolved using its Punycode representation, but displayed to users in its decoded Unicode form:

Unicode domain:   аpple.com  (Cyrillic 'а' U+0430 in first position)
Punycode form:    xn--pple-43d.com
DNS resolves:     xn--pple-43d.com → attacker's server IP
Browser displays: аpple.com        ← looks identical to apple.com

The attack works because the Punycode form — which is what actually routes the connection — differs entirely from the legitimate domain, while the rendered display form is pixel-perfect identical in most fonts at typical screen sizes.

2. Commonly Exploited Unicode Character Substitutions

Cyrillic is the most-used script for homograph attacks targeting Latin-alphabet domains because it has the highest density of visually identical characters. These are the highest-impact substitutions:

Target (Latin)Substitute CharacterUnicode Code PointScriptTargeted Brands
aаU+0430Cyrillic Small Letter Apaypal, amazon, apple, facebook
cсU+0441Cyrillic Small Letter Escoinbase, chase, citi
eеU+0435Cyrillic Small Letter Iegoogle, netflix
oоU+043ECyrillic Small Letter Ogoogle, coinbase, outlook
pрU+0440Cyrillic Small Letter Erpaypal, apple
xхU+0445Cyrillic Small Letter Ha(rare)
yуU+0443Cyrillic Small Letter U(rare)
iіU+0456Ukrainian Small Letter Imicrosoft, linkedin

Beyond Cyrillic, Greek (Μ U+039C for M), Armenian (ո U+0578 for n), and other scripts provide additional lookalike characters. ICANN maintains a reference list of confusable character pairs in its IDN implementation guidelines.

3. Browser Rendering Policies — Who Shows Punycode and When

Browser vendors have implemented policies to mitigate homograph attacks, but these policies vary by browser and have known gaps:

BrowserDisplay PolicyTriggers Punycode DisplayGaps
ChromeShows Punycode if domain mixes scripts OR if TLD is untrustedMixed Cyrillic+Latin, untrusted TLDAll-Cyrillic domains in trusted TLDs (.com) may still show decoded Unicode
FirefoxUses an allowlist of scripts per TLD; shows Punycode for violationsScript not in TLD allowlist, or mixed scriptsAllowlist not exhaustive; new TLDs may not have a policy yet
SafariShows Punycode for mixed scripts; uses Apple's internal confusable databaseMixed scripts, known confusablesConfusable database requires updates as new attack patterns emerge
EdgeInherits Chromium policySame as ChromeSame as Chrome

The critical gap across all browsers: an all-Cyrillic domain (no Latin characters mixed in) that is registered under a trusted TLD like .com may still render in decoded form in some browser/OS combinations. The browser policies target mixed-script labels, not all-Unicode labels. A domain where all characters are replaced with Cyrillic equivalents may not trigger Punycode display.

Email clients are worse. Most email applications render hyperlink display text exactly as the sender specifies — the underlying href is not always visible without hovering. Phishing emails routinely display paypal.com as link text while the href routes to рaypal.com (Cyrillic р). The hover tooltip may show the Punycode form in some clients but not others, and on mobile there is no hover.

4. Technical Detection — How Scanners Identify Homograph Domains

Step 1: Punycode Prefix Scanning

Any hostname label containing the prefix xn-- is immediately flagged for IDN analysis. If the URL was submitted in Unicode form (not Punycode), the scanner normalizes it using RFC 3492 Punycode encoding and checks each label.

Step 2: Mixed-Script Detection

The most reliable homograph indicator is a label that mixes characters from two different Unicode script families. A scanner extracts each label (the segments separated by dots) and checks whether it contains characters from more than one script block:

function detectMixedScriptHomograph(hostname) {
  const labels = hostname.toLowerCase().split('.');
  // Unicode block ranges for common script families
  const scripts = {
    cyrillic: /[\u0400-\u04FF]/,
    greek:    /[\u0370-\u03FF]/,
    armenian: /[\u0530-\u058F]/,
    latin:    /[a-z]/,
  };

  for (const label of labels) {
    const matchedScripts = Object.entries(scripts)
      .filter(([, regex]) => regex.test(label))
      .map(([name]) => name);

    if (matchedScripts.includes('latin') && matchedScripts.length > 1) {
      return {
        isHomograph: true,
        risk: 'CRITICAL',
        label,
        scripts: matchedScripts,
        reason: `Mixed scripts: ${matchedScripts.join(' + ')}`
      };
    }
  }
  return { isHomograph: false };
}

Step 3: Confusable Character Database Lookup

For all-Cyrillic or all-Greek domains (no mixed script), the scanner maps each character against Unicode's confusables.txt dataset (maintained by the Unicode Consortium). Each character in the domain is checked against its confusable mapping — if the mapped Latin equivalents form a known brand name, the domain is flagged as a potential homograph attack.

Example: gооgle.com (two Cyrillic 'о' characters) maps via confusables to google.com — a Levenshtein distance of 0 against a known brand, flagged as critical.

Step 4: Levenshtein Distance Against Brand Database

After converting confusable characters to their Latin equivalents, the scanner computes the Levenshtein edit distance between the normalized domain and every entry in a brand database. A distance of 0 (exact match after normalization) is a confirmed homograph. A distance of 1 indicates either typosquatting or a partial homograph substitution.

5. Enterprise Defense Strategies

A. Defensive Domain Registration

Register the most common homograph variants of your brand domain and point them to your primary domain via 301 redirect. This prevents attackers from acquiring them. For a brand like paypal.com, the highest-priority registration targets are domains where one or two high-frequency characters (a, p, o, e) are replaced with Cyrillic equivalents. A brand protection service can enumerate these programmatically and monitor for new registrations of confusable domains.

B. Certificate Transparency Monitoring

Homograph domains that use HTTPS (required to appear legitimate) must obtain TLS certificates. Every publicly trusted CA submits issued certificates to Certificate Transparency logs. Monitor CT logs for certificates issued to any domain that is a confusable of your brand — services like crt.sh, Facebook's CT monitor, or commercial brand protection platforms provide this feed. A new certificate for a homograph domain typically means a campaign is imminent.

C. Email Gateway Configuration

Configure your Secure Email Gateway to flag inbound messages where any URL in the message body decodes to a Punycode domain (xn-- prefix), or where a displayed domain in a hyperlink's text differs from the href domain. Both conditions are reliable homograph indicators.

D. User Training — The Right Mental Model

The most important training message: reading a URL is not the same as verifying it. Users who believe they can spot phishing by reading the URL are the most vulnerable to homograph attacks — precisely because the domain looks correct. Train users to paste suspicious links into a URL scanner rather than visually inspecting them, and to treat any credential prompt reached via a link in an email as suspicious until verified.

6. Detecting Homograph Attacks with IncogSay

IncogSay's URL Scanner and Safe Link Checker both run in-memory Punycode decoding and confusable character analysis on every submitted URL. The scanner:

  • Detects xn-- Punycode labels in submitted domains
  • Converts all characters to their Unicode confusable equivalents and checks against a major-brand database
  • Runs mixed-script analysis on each label and flags Cyrillic/Greek/Armenian characters mixed with Latin
  • Reports the raw Punycode form alongside the display form, so the difference is immediately visible
  • Combines homograph risk with domain age, SSL certificate, and entropy score into a consolidated threat score