An HTML entity looks like unnecessary clutter โ why write & instead of just &? โ but it exists to solve a real ambiguity problem in how HTML gets parsed.
Why Certain Characters Break HTML
The ampersand is reserved in HTML as the character that begins an entity reference, and the angle brackets < and > are reserved as the start and end of a tag. A parser encountering a literal & or < not meant as part of an entity or a tag can misinterpret the surrounding markup, corrupting the page's structure โ escaping these characters as entities removes the ambiguity entirely.
Named vs. Numeric Entities
A named entity uses a memorable word โ & for an ampersand, © for a copyright symbol. A numeric entity instead references the character's Unicode code point directly, in decimal (&) or hexadecimal (&). Numeric entities work for virtually any character, including ones with no widely supported named entity, which is why they're the more universal fallback.
Escaping Non-ASCII Characters
On a page correctly declared with a modern encoding like UTF-8, accented letters, emoji, and other non-ASCII characters don't need entity escaping at all โ UTF-8 can represent them directly. Escaping them anyway is sometimes done defensively, for compatibility with older systems or specific data pipelines expecting pure ASCII, but it isn't a requirement of HTML itself on a properly encoded page.
Which Characters Actually Need It
The short, genuinely required list: & (starts an entity), < and > (start and end a tag), and inside attribute values, quote characters. Everything else is generally safe to include literally on a properly UTF-8 encoded page โ entity escaping matters specifically for this small set of structurally significant characters.
Encoding or Decoding Right Now
Use our free HTML Entity Encoder/Decoder to encode special characters into named, decimal, or hexadecimal entities, decode entities back into plain text, or optionally escape all non-ASCII characters too.
FAQ
Why does a raw ampersand break HTML if typed directly? The ampersand is reserved in HTML as the character that starts an entity reference โ the parser sees & and expects it to be followed by an entity name or number ending in a semicolon. A literal ampersand not meant as an entity confuses this parsing, which is why it needs to be escaped as & to be displayed as a plain ampersand character.
What's the difference between a named entity and a numeric entity? A named entity uses a memorable word, like & for an ampersand or © for a copyright symbol. A numeric entity instead references the character's Unicode code point directly, either in decimal (&) or hexadecimal (&) โ useful for characters that don't have a widely supported named entity, since virtually every character has a numeric code point.
Do I need to escape every non-ASCII character, like accented letters or emoji? Not strictly, as long as the page correctly declares a modern character encoding like UTF-8, which can represent virtually any character directly without needing an entity. Escaping non-ASCII characters as entities anyway is sometimes done defensively for maximum compatibility with older systems or specific data pipelines, but it isn't required by HTML itself on a properly UTF-8 encoded page.
Which characters actually need to be escaped in HTML? The characters with special meaning to an HTML parser: & (starts an entity), < and > (start and end a tag), and inside attribute values, quote characters (" or '). Everything else is generally safe to include literally on a properly encoded page, which is why entity escaping matters most for exactly this small set of structurally significant characters.