toolready. HTML Entity Encode / Decode

HTML Entity Encode / Decode

Escape and unescape HTML entities and character references.

What this does

Converts between raw text and HTML character references, both directions, live as you type. The page opens in Decode mode — paste entity-laden text pulled from a feed, an email or a database column and read it back as text; switch to Encode for the other direction. Nothing is sent anywhere, so customer content and internal templates are safe to paste.

Which characters does Encode escape?

Five, in this order — the ampersand first, so already-escaped output does not get mangled:

&   →  &
<   →  &lt;
>   →  &gt;
"   →  &quot;
'   →  &#39;

So <a href="x">&'</a> becomes &lt;a href=&quot;x&quot;&gt;&amp;&#39;&lt;/a&gt;. The apostrophe becomes numeric &#39; rather than &apos; on purpose: &apos; is an XML entity older HTML parsers and some email clients do not recognise, while the numeric form is understood everywhere.

What is the difference between named and numeric entities?

A named reference spells the character out — &copy;, &mdash;, &nbsp; — and only works if the parser knows the name. A numeric reference gives the Unicode codepoint directly, decimal (&#169;) or hex (&#x00A9;), and needs no lookup table. Both forms require the trailing semicolon here. Decoding handles all three shapes:

&copy; 2026 &mdash; &amp; &nbsp;end   →  © 2026 — & ␣end
&#169; &#x1F600;                →  © 😀
&notreal;                        →  &notreal;

Numeric decoding covers the whole Unicode range, emoji above U+FFFF included. Named decoding covers what turns up in the wild — the five reserved entities plus curly quotes, dashes, ellipsis, &nbsp;, currency and legal symbols — not the full HTML5 list of some two thousand names. Anything unrecognised is left exactly as it was, which is why &notreal; comes through untouched.

What does the "also escape non-ASCII" option do?

Off, only those five characters change and accented letters, CJK and emoji pass through as themselves — the right choice for any modern UTF-8 page. On, every character above codepoint 127 also becomes a decimal reference, so é turns into &#233;. Reach for it when the output has to survive an ASCII-only pipeline: legacy email templates, an XML transport with a mis-declared charset, a CMS field that mangles high bytes. It escapes per UTF-16 code unit, so characters above U+FFFF come out as a pair of references; leave it off for text containing emoji.

How do I safely put user text into an HTML page?

  1. Encode the text here, or with your framework's escaping, before it reaches the template.
  2. Put it in a text node or a quoted attribute value. Escaped text is only safe in those two places.
  3. Do not rely on it inside <script> or style blocks, or in href and src values — those contexts have their own rules, and javascript: in an href survives entity escaping intact.
  4. Prefer your template engine's automatic escaping. This tool is for one-off content and for inspecting what a system already produced.

Why is there a weird space after decoding?

Because &nbsp; decodes to U+00A0, a non-breaking space — a real character that looks identical to a normal space but does not wrap, does not match \s in some regex flavours, and breaks comparisons against hand-typed text. Word processors scatter them liberally, so they are a common cause of "these two strings look the same but aren't". Swap them out with find and replace if unwanted; the whitespace cleaner targets zero-width characters and line padding, not U+00A0. URL encode/decode solves the same escaping problem for links, and HTML to Markdown is the better route if you want prose rather than markup.