Unicode Inspector
See every character, code point and hidden symbol
Inspect text character by character: code points, names, categories, invisible characters, look-alike letters and normalization.
- Unicode
- UTF-8
- Emoji
- Homoglyph
How to use Unicode Inspector
- Paste text into the Text box, drop a text file onto it, or use Open file. Try an example loads text with emoji, hidden characters and look-alike letters.
- Read the summary: counts of characters, code points, UTF-8 bytes and UTF-16 units, and a warning listing any invisible, bidi-control or look-alike characters found.
- Scan Characters (grapheme clusters) for highlighted characters, then click a character in the Code points table to see its name, category, script, block and HTML, JavaScript, CSS and URL escapes.
- Check Scripts and look-alikes for words that mix scripts, and Normalization to compare the NFC, NFD, NFKC and NFKD forms and copy one.
- Under Cleaned text, choose Turn unusual spaces into normal spaces and Replace look-alike letters, then Copy, Send to…, or Clean the input.
How it works
The text is split into user-perceived characters (grapheme clusters) with the browser’s Intl.Segmenter, so an emoji family or a letter with combining accents counts as one character, and each cluster is split into its code points. For every code point the tool shows the U+ value, its UTF-8 bytes and UTF-16 units, and its general category (such as Lu or Cf) and script, both taken from your browser’s Unicode data through regex property escapes. Block names and character names come from a built-in table that covers the common blocks, with algorithmic names for CJK ideographs, Hangul syllables, variation selectors, tags, regional indicators and many accented Latin letters.
Each code point is checked against lists of suspicious characters: bidirectional controls used in Trojan Source attacks, invisible characters such as zero-width spaces and the byte order mark, unusual spaces such as the no-break space, control characters, tag characters that can hide text, stray variation selectors, private-use, unassigned and lone surrogate code points, and the replacement character. Zero-width joiners, variation selectors and tags inside an emoji sequence are expected and only noted. Look-alikes are found with a small hand-picked list of Cyrillic, Greek and Latin letters that imitate Latin ones, plus fullwidth forms, and any word mixing letters from more than one script is listed with the Latin text it imitates.
Normalization runs String.prototype.normalize for all four forms and reports whether each differs from your text. Cleaned text removes invisible, bidi-control, tag, stray variation selector, control and lone surrogate characters while keeping emoji sequences intact, and can replace unusual spaces and look-alikes. Everything runs in your browser.
Limits
- Text over 1,000,000 characters isn’t inspected.
- The Code points table lists the first 2,000 code points and the character chips show up to 1,000 characters from the first 20,000; the counts and warnings cover the whole text.
- Character names come from a built-in table for common blocks; other characters show their block instead of a name.
- Look-alike detection uses a short hand-picked list, not Unicode’s full confusables data, so many look-alikes aren’t caught.
- Categories and scripts depend on your browser’s Unicode version, so very new characters may show as unassigned. Scripts outside a built-in list of 36 show as Other.
- Dropped or opened files are limited to 10 MB and must be text.
Privacy
Your text is analysed entirely in your browser; it is never uploaded or stored. The Share button copies a link with your text in its # fragment, which browsers don’t send to servers, but anyone you give the link to can read it. Send to… passes text between tools through this tab’s session storage, and it is removed as soon as the receiving tool reads it.
Frequently asked questions
Why is the character count different from the length in my code?
The tool counts what you see as one character (a grapheme cluster). JavaScript’s length counts UTF-16 units, so 👍🏽 is one character, two code points and four UTF-16 units. All three counts, plus the UTF-8 byte count, are shown at the top.
What is a Trojan Source attack?
Bidirectional control characters such as U+202E RIGHT-TO-LEFT OVERRIDE can make source code display in a different order than the compiler reads it, so code can look like a comment but still run. These characters are flagged in red.
How can I tell if a domain or word uses look-alike letters?
Paste it in. Letters such as Cyrillic а in pаypal.com are flagged as look-alikes, and Scripts and look-alikes lists words that mix scripts together with the Latin text they imitate.
Which normalization form should I use?
NFC is the usual choice for storing and comparing text: it composes e plus a combining accent into é. NFD splits them apart. NFKC and NFKD also fold compatibility characters, such as fi to fi and ① to 1, which helps with searching but changes how text looks.
More tools
- Clean Image: Inspect and remove hidden image metadata
- JWT Decoder: Decode and verify JSON Web Tokens
- Diff Checker: Compare two texts line by line
- JS Runner: Run JavaScript and TypeScript in your browser
- JSON Formatter: Format, validate and minify JSON
- Encode / Decode: Base64, URL, HTML entity and hex
- Hash Generator: MD5, SHA and HMAC of any text
- UUID Generator: Generate UUID v4 and v7 in bulk
- Timestamp Converter: Unix time ↔ human dates
- Regex Tester: Test regular expressions live
- URL Parser: Break a URL into its parts
- HTTP Status Codes: Look up any HTTP status code
- MIME Type Lookup: File extension ↔ MIME type
- Password Generator: Strong random passwords and passphrases
- Random String Generator: Random tokens, IDs and keys
- Slug Generator: Turn titles into URL slugs
- Case Converter: camelCase, snake_case, Title Case and more
- Word Counter: Count words, characters and reading time
- JSON to TypeScript: Generate TypeScript types from JSON
- JSON Diff: Compare two JSON documents structurally
- JSON to SQL: Turn JSON arrays into SQL inserts
- YAML ↔ JSON: Convert between YAML and JSON
- XML ↔ JSON: Convert between XML and JSON
- CSV ↔ JSON: Convert between CSV and JSON
- CSV Viewer: View, sort and filter CSV files
- SQL Formatter: Format and beautify SQL queries
- cURL ↔ Fetch: Convert cURL commands to fetch and back
- Markdown Editor: Write Markdown with a live preview
- Text Cleaner: Remove duplicate lines, empty lines and extra spaces
- Find & Replace: Find and replace in any text
- Cron Expression Builder: Build and explain cron schedules
- User-Agent Parser: Identify browser, OS and device from a user agent
- HTTP Headers Inspector: Paste response headers and get them explained
- JWT Generator: Create and sign test JSON Web Tokens
- Certificate Inspector: Decode PEM certificates and keys
- Meta Tag Inspector: Check a page's SEO and social tags
- UTM Builder: Build campaign URLs with UTM parameters
- URL Cleaner: Strip tracking parameters from links
- Robots.txt Generator: Create and test a robots.txt file
- Sitemap Generator: Create an XML sitemap from a list of URLs
- Image Compressor: Shrink JPEG, WebP and AVIF images in your browser
- Image Resizer: Resize images by pixels, percentage or to fit a box
- Image Converter: Convert between PNG, JPEG, WebP and AVIF
- Image to Base64: Encode images as Base64 data URIs and decode them back
- SVG Optimizer: Minify and sanitize SVG files
- Favicon Generator: Make favicon.ico, Apple and Android icons from an image or emoji
- Color Converter: HEX, RGB, HSL, OKLCH and contrast checks
- Number Base Converter: Binary, octal, decimal, hex and float bits
- IP / CIDR Calculator: Subnets, masks and IP ranges for IPv4 and IPv6
- JSONPath Query: Query JSON with JSONPath expressions
- JSON Schema Validator: Validate JSON against a schema, or generate one
- Semver Checker: Check versions against semver ranges
- chmod Calculator: Unix permissions: rwx ↔ octal
- .env Diff: Compare and validate .env files
- TOTP Generator: Generate and verify 2FA codes
- String Escaper: Escape and unescape strings for any language
- Mock Data Generator: Generate realistic fake data
- QR Code Generator: Create QR codes for links, Wi-Fi and contacts
- Lorem Ipsum Generator: Placeholder text in paragraphs, sentences or words
- Date Calculator: Date differences, business days and durations
- Unit Converter: Convert bytes, lengths, weights, temperatures and more
- Query CSV with SQL: Run SQL queries on CSV files
- PDF Merge & Split: Merge, split, reorder and rotate PDFs
- PDF Metadata Cleaner: See and remove hidden PDF metadata
- Office Metadata Cleaner: Remove author and revision data from Word, Excel and PowerPoint
- Images to PDF: Combine images into one PDF
- Image Editor: Crop, rotate, resize and adjust images
- Encrypt / Decrypt Text: Encrypt text with a passphrase (AES-GCM)
- SSH Key Generator: Generate Ed25519 and RSA SSH keys locally
- Email Header Analyzer: Trace an email's path and check SPF, DKIM and DMARC
- JSON to Code: Generate Go, Python, Rust, Java, C# and Kotlin models from JSON
- docker run ↔ Compose: Convert docker run commands to docker-compose and back
- Color Palette Extractor: Pull the dominant colours out of any image
- Password Strength Checker: How long would your password take to crack?
- SPF / DKIM / DMARC Checker: Validate and explain email DNS records
- Kubernetes YAML Checker: Validate and explain Kubernetes manifests
- .gitignore Generator: Build a .gitignore from presets
- CSP Builder: Build and check a Content-Security-Policy
- JSON-LD Generator: Create schema.org structured data
- Open Graph Image Generator: Make 1200×630 social preview images
- CSS Generator: Gradients, shadows, clamp() and more
- Time Zone Meeting Planner: Find meeting times across time zones