toolready. AI Tokenizer

AI Tokenizer

See how GPT-4o, GPT-4, and GPT-3.5 break your text into tokens.

Model
Tokens
0
Characters
0
Words
0
Chars / token
—
Token visualization
Type or paste text above to see tokens.

What this does

Runs OpenAI's byte-pair encodings over your text and shows what the model sees: a token count, the chunk boundaries in colour, and the ID behind each chunk. Display them as text, IDs or both, and copy either as a JSON array. The stat row adds characters, words and characters per token. Encoding tables are megabytes each, so they load on demand once you type; your text is tokenized in the tab and never sent anywhere.

What is a token?

A fragment of text produced by a byte-pair encoding — a merge table built by repeatedly gluing the most frequent character pairs in a corpus. Common words survive whole, rarer ones shatter, and boundaries are unintuitive: a leading space belongs to the word after it, numbers split oddly.

Hello world!  →  "Hello" | " world" | "!"                3 tokens (13225·2375·0)
2026-09-08    →  "202" | "6" | "-" | "09" | "-" | "08"   6 tokens

Which is why models miscount letters, and why an ID column costs more than the prose around it.

Which encoding should I pick for my model?

  • o200k_base — every current OpenAI model: GPT-4o, GPT-4.1, GPT-5, o1, o3, o4.
  • cl100k_base — GPT-4, GPT-4-turbo, GPT-3.5-turbo, and text-embedding-3 / ada-002.
  • p50k_base — Codex and text-davinci-002 / 003.
  • r50k_base — original GPT-3: davinci, curie, babbage, ada.

The number in each name is roughly its vocabulary size. When in doubt use o200k_base.

Why does the count change when I switch encoding?

A bigger vocabulary means longer merges, and the tables were built on different data. On English the gap is small — the pangram "The quick brown fox jumps over the lazy dog." is 10 tokens on both — but it widens:

你好世界    o200k_base   2   "你好" | "世界"
            cl100k_base  5   "你" | "好" | (partial byte) | "世" | "界"

tokenizer   o200k_base   2   "token" | "izer"
            cl100k_base  1   "tokenizer"

The second case runs the other way: a broader vocabulary does not guarantee fewer tokens. The blank chunk in the cl100k row is real — that encoding splits some CJK characters across several tokens, so a piece can be part of a UTF-8 sequence rather than a printable character.

How do I estimate what a prompt will cost?

  1. Paste the prompt as your code sends it — system message and retrieved context included, since those dominate the bill.
  2. Multiply by your provider's input price per million tokens, then add expected output, usually several times dearer.
  3. Sanity-check chars-per-token: English prose sits near 4; JSON, minified code and non-Latin scripts run lower.

Why is my API token count higher than this page shows?

A chat request is more than its text. Every message carries role and delimiter overhead, tool schemas count as input, images and audio get their own token budgets, and reasoning models bill thinking tokens you never see. This page counts raw text — a floor, not a quote. One edge case: like the reference tokenizer, it refuses text containing a literal marker such as <|endoftext|> rather than counting it.

Does this work for Claude, Gemini, or Llama?

Not exactly. These are OpenAI's BPE tables; Anthropic and Google ship no open in-browser tokenizer for their current models, and Llama and Mistral use their own SentencePiece vocabularies. Other vendors land within roughly 10–20% of the o200k_base count on English prose — fine for capacity planning, not billing. For exact figures use Anthropic's count_tokens endpoint or the model's own library. For prose statistics the word counter is simpler, and the JSON formatter pairs well with structured prompts.