LLM Token Counter

Free LLM token counter. Count GPT-5, GPT-4o and GPT-4 tokens exactly with OpenAI's tiktoken tables — in your browser, nothing uploaded.

Runs in your browserNothing uploadedFree · no signup
0tokens — exact
0Characters
0Words
0Chars per token
0.0%Of context window

128,000 tokens still free in a 128,000-token window. That budget covers the model's reply too, not just your prompt.

Counted in your browser with the same BPE tables OpenAI publishes for tiktoken — your prompt is never uploaded. Counts cover the text itself; a chat request also spends a few tokens per message on role and formatting overhead.

Frequently asked questions

How accurate is this token counter?
For the OpenAI encodings it is exact, not an estimate — it runs the same byte-pair encoding tables OpenAI publishes for tiktoken, so o200k_base, cl100k_base, p50k_base and r50k_base return the identical count their API bills you for. The models listed under "estimate only" are a different matter: those vendors do not publish a tokenizer that runs in a browser, so the number shown for them is a characters-divided-by-four rule of thumb and is labelled as such.
Which encoding does my model use?
o200k_base covers GPT-5, GPT-4.1, GPT-4o and the o-series reasoning models. cl100k_base covers GPT-4, GPT-4 Turbo, GPT-3.5 Turbo and the version-3 embedding models. p50k_base covers text-davinci-002/003 and Codex, and r50k_base covers the original GPT-3 base models and GPT-2. The tool is keyed on the encoding rather than the model name because model line-ups get renamed and retired, while the encoding a family uses does not change.
Can I count Claude or Gemini tokens exactly?
Not in a browser. Anthropic and Google do not publish tokenizer files you can run client-side, so any site claiming an exact Claude or Gemini count offline is guessing. For an exact figure, Anthropic offers a /v1/messages/count_tokens endpoint and the Gemini API offers countTokens, both of which return the real number for a specific model. The estimate here is useful for sizing a prompt, not for reconciling a bill.
Why is my token count higher than my word count?
Tokens are sub-word fragments, not words. Common English words are usually a single token, but rare words, names, long compounds and anything with unusual casing get split into pieces, and punctuation often costs its own token. English prose averages roughly four characters per token; code, JSON and non-Latin scripts such as Hindi, Japanese or Arabic pack far fewer characters into each token, so the same visible text can cost several times more.
Does the count include the model's reply?
No. The count covers only the text you paste. A real request also spends a few tokens per message on role and formatting overhead, and the model's reply is billed separately as output tokens — usually at a higher rate. The context window has to hold your prompt and the reply together, so leave headroom rather than filling it to the last token.
Is my text uploaded anywhere?
No. The encoding tables are downloaded to your browser and the tokenizing runs locally, so your prompt, document or source code never leaves your device.