What this does
Converts between raw text and HTML character references, both directions, live as you type. The page opens in Decode mode — paste entity-laden text pulled from a feed, an email or a database column and read it back as text; switch to Encode for the other direction. Nothing is sent anywhere, so customer content and internal templates are safe to paste.
Which characters does Encode escape?
Five, in this order — the ampersand first, so already-escaped output does not get mangled:
& → &
< → <
> → >
" → "
' → '
So <a href="x">&'</a> becomes
<a href="x">&'</a>.
The apostrophe becomes numeric ' rather than
' on purpose: ' is an XML entity
older HTML parsers and some email clients do not recognise, while the numeric
form is understood everywhere.
What is the difference between named and numeric entities?
A named reference spells the character out — ©,
—, — and only works if the
parser knows the name. A numeric reference gives the Unicode codepoint
directly, decimal (©) or hex
(©), and needs no lookup table. Both forms require
the trailing semicolon here. Decoding handles all three shapes:
© 2026 — & end → © 2026 — & ␣end
© 😀 → © 😀
¬real; → ¬real;
Numeric decoding covers the whole Unicode range, emoji above U+FFFF
included. Named decoding covers what turns up in the wild — the five reserved
entities plus curly quotes, dashes, ellipsis, ,
currency and legal symbols — not the full HTML5 list of some two thousand
names. Anything unrecognised is left exactly as it was, which is why
¬real; comes through untouched.
What does the "also escape non-ASCII" option do?
Off, only those five characters change and accented letters, CJK and emoji
pass through as themselves — the right choice for any modern UTF-8 page. On,
every character above codepoint 127 also becomes a decimal reference, so
é turns into é. Reach for it when the
output has to survive an ASCII-only pipeline: legacy email templates, an XML
transport with a mis-declared charset, a CMS field that mangles high bytes.
It escapes per UTF-16 code unit, so characters above U+FFFF come out as a
pair of references; leave it off for text containing emoji.
How do I safely put user text into an HTML page?
- Encode the text here, or with your framework's escaping, before it reaches the template.
- Put it in a text node or a quoted attribute value. Escaped text is only safe in those two places.
- Do not rely on it inside
<script>orstyleblocks, or inhrefandsrcvalues — those contexts have their own rules, andjavascript:in an href survives entity escaping intact. - Prefer your template engine's automatic escaping. This tool is for one-off content and for inspecting what a system already produced.
Why is there a weird space after decoding?
Because decodes to U+00A0, a non-breaking space — a
real character that looks identical to a normal space but does not wrap, does
not match \s in some regex flavours, and breaks comparisons
against hand-typed text. Word processors scatter them liberally, so they are
a common cause of "these two strings look the same but aren't". Swap them out
with find and replace if unwanted; the
whitespace cleaner targets zero-width characters
and line padding, not U+00A0. URL encode/decode solves the
same escaping problem for links, and HTML to Markdown
is the better route if you want prose rather than markup.