Back to blog

October 6, 2026

How to Validate UTF-8 Encoding

Learn how UTF-8 encoding works, what invalid sequences look like, and how to validate text before it breaks your pipeline.

UTF-8 is the default text encoding on the web and in most modern systems. Validating UTF-8 means checking that a byte sequence follows the encoding rules—so decoders, databases, and UIs do not choke on corrupt input.

What UTF-8 actually encodes

UTF-8 maps Unicode code points to one–four bytes:

Code points Bytes Pattern (binary)
U+0000–U+007F 1 0xxxxxxx
U+0080–U+07FF 2 110xxxxx 10xxxxxx
U+0800–U+FFFF 3 1110xxxx 10xxxxxx 10xxxxxx
U+10000–U+10FFFF 4 11110xxx 10xxxxxx 10xxxxxx 10xxxxxx

ASCII is valid UTF-8 by design. Multi-byte characters use a leading byte plus continuation bytes (10xxxxxx).

Common ways UTF-8 goes wrong

Invalid UTF-8 usually comes from:

  • Truncated multi-byte sequences — cutting a string mid-character
  • Unexpected continuation bytes — a 10xxxxxx byte with no valid lead
  • Overlong encodings — encoding ASCII with extra bytes (forbidden)
  • Surrogate halves — UTF-16 leftovers treated as Unicode scalars
  • Wrong encoding labeled as UTF-8 — Latin-1/Windows-1252 bytes mislabeled

Symptoms include replacement characters (�), mojibake (é instead of é), failed JSON parses, or database rejection.

How to validate UTF-8

1. Treat input as bytes

Validation is about bytes, not characters you already decoded. If a language has already turned bytes into a string with a lossy decoder, you may have lost the evidence.

2. Decode strictly

Use a strict UTF-8 decoder that errors on illegal sequences instead of inserting �. In pipelines, fail closed when encoding must be trusted (security-sensitive parsers, signatures, filenames).

3. Inspect suspects in hex

When validation fails, convert the bytes to hex to see the broken sequence. The Convert UTF-8 to Hex and Convert Hex to UTF-8 tools help you inspect and round-trip samples.

4. Validate in the browser

Paste suspicious text or decoded output into Validate UTF8 to check whether the content is well-formed and to surface encoding issues early.

Validation checklist

  1. Confirm the producer’s declared charset (headers, file metadata, DB collation).
  2. Decode with a strict UTF-8 checker.
  3. Reject or quarantine invalid byte sequences.
  4. Normalize only after validation (NFC/NFKC is a separate step).
  5. Re-encode explicitly when crossing system boundaries.

Related encoding tools

Summary

UTF-8 validation protects parsers and users from corrupt or mislabeled text. Decode strictly, inspect failures in hex, and fix the producer when possible. Use Validate UTF8 whenever you need a quick, local check before data enters a stricter system.

Try these browser tools connected to this article.