What this does
Runs OpenAI's byte-pair encodings over your text and shows what the model sees: a token count, the chunk boundaries in colour, and the ID behind each chunk. Display them as text, IDs or both, and copy either as a JSON array. The stat row adds characters, words and characters per token. Encoding tables are megabytes each, so they load on demand once you type; your text is tokenized in the tab and never sent anywhere.
What is a token?
A fragment of text produced by a byte-pair encoding — a merge table built by repeatedly gluing the most frequent character pairs in a corpus. Common words survive whole, rarer ones shatter, and boundaries are unintuitive: a leading space belongs to the word after it, numbers split oddly.
Hello world! → "Hello" | " world" | "!" 3 tokens (13225·2375·0)
2026-09-08 → "202" | "6" | "-" | "09" | "-" | "08" 6 tokens Which is why models miscount letters, and why an ID column costs more than the prose around it.
Which encoding should I pick for my model?
- o200k_base — every current OpenAI model: GPT-4o, GPT-4.1, GPT-5, o1, o3, o4.
- cl100k_base — GPT-4, GPT-4-turbo, GPT-3.5-turbo, and
text-embedding-3/ada-002. - p50k_base — Codex and
text-davinci-002/003. - r50k_base — original GPT-3: davinci, curie, babbage, ada.
The number in each name is roughly its vocabulary size. When in doubt use o200k_base.
Why does the count change when I switch encoding?
A bigger vocabulary means longer merges, and the tables were built on different data. On English the gap is small — the pangram "The quick brown fox jumps over the lazy dog." is 10 tokens on both — but it widens:
你好世界 o200k_base 2 "你好" | "世界"
cl100k_base 5 "你" | "好" | (partial byte) | "世" | "界"
tokenizer o200k_base 2 "token" | "izer"
cl100k_base 1 "tokenizer" The second case runs the other way: a broader vocabulary does not guarantee fewer tokens. The blank chunk in the cl100k row is real — that encoding splits some CJK characters across several tokens, so a piece can be part of a UTF-8 sequence rather than a printable character.
How do I estimate what a prompt will cost?
- Paste the prompt as your code sends it — system message and retrieved context included, since those dominate the bill.
- Multiply by your provider's input price per million tokens, then add expected output, usually several times dearer.
- Sanity-check chars-per-token: English prose sits near 4; JSON, minified code and non-Latin scripts run lower.
Why is my API token count higher than this page shows?
A chat request is more than its text. Every message carries role and
delimiter overhead, tool schemas count as input, images and audio get their
own token budgets, and reasoning models bill thinking tokens you never see.
This page counts raw text — a floor, not a quote. One edge case: like the
reference tokenizer, it refuses text containing a literal marker such as
<|endoftext|> rather than counting it.
Does this work for Claude, Gemini, or Llama?
Not exactly. These are OpenAI's BPE tables; Anthropic and Google ship no open
in-browser tokenizer for their current models, and Llama and Mistral use
their own SentencePiece vocabularies. Other vendors land within roughly
10–20% of the o200k_base count on English prose — fine for capacity planning,
not billing. For exact figures use Anthropic's count_tokens
endpoint or the model's own library. For prose statistics the
word counter is simpler, and the
JSON formatter pairs well with structured prompts.