Developer & encoding

What Is a Token in an LLM? How GPT and Claude Count Text

What is a token in an LLM? How tokens differ from words, why the same sentence costs more in Hindi than English, which encoding each GPT model uses, and how to count tokens exactly.

6 min readUpdated Aug 9, 2026

Every price list, rate limit and context window published by an AI company is denominated in tokens, and almost nobody can look at a paragraph and say how many tokens it is. The unit is precisely defined - it is just not one humans think in. This guide covers what a token is, why the count never matches your word count, why the same sentence costs three times more in one language than another, and how to get an exact number.

What a token actually is

A language model reads neither characters nor words. Your text is first chopped into tokens - fragments drawn from a fixed vocabulary the model was trained with - and each fragment is swapped for the integer that identifies it. The model only ever sees that list of integers.

The vocabulary is built by byte pair encoding: starting from raw bytes, the algorithm repeatedly merges the most frequent adjacent pair until it has as many entries as the designers wanted - roughly 100,000 for GPT-4, 200,000 for GPT-4o and GPT-5. Because the merges follow frequency, common sequences become single entries and rare ones stay broken apart, which explains nearly every surprise in a token count.

Why the token count never matches the word count

Common words are one token each, including the space in front of them: in GPT-4o's vocabulary the four characters of " the" are a single token, and so are " The" and " and". Rare words are not so lucky. "Antidisestablishmentarianism" is one word and 28 characters, but six tokens - " ant" + "idis" + "est" + "ablishment" + "arian" + "ism". Brand names fare badly too: " Zenoply" costs three, because the model has never seen it often enough to earn an entry of its own.

This is why a word counter cannot substitute for a token counter. For plain English prose the ratio is fairly stable at roughly 1.1 to 1.3 tokens per word, so you can estimate in your head. The moment your text contains code, JSON, identifiers or unusual names, that breaks down: a small JSON object of 84 characters comes to 25 tokens.

A worked example, token by token

The Token Counter ships with a sample that mixes the awkward cases on purpose: an English sentence, a paragraph containing a deliberately rare word, a line of JavaScript, and a sentence of Japanese. It runs to 266 characters and 39 words.

Under o200k_base - the encoding GPT-5, GPT-4.1 and GPT-4o use - that is 84 tokens, or 3.17 characters per token. Switch the dropdown to cl100k_base, which is what GPT-4 and GPT-3.5 Turbo use, and the same text becomes 88. Nothing changed except the vocabulary; the newer one has twice as many entries, so it finds longer matches and needs fewer of them.

The breakdown shades every token separately, so you can see where the splits fall - product names, UUIDs, timestamps and indented JSON are the usual offenders.

Non-English text costs more tokens

The training data behind these vocabularies is overwhelmingly English, so English gets the efficient merges and everything else is assembled from smaller pieces. One instruction - "Please summarise this document in three short bullet points." - and its translations, under o200k_base:

  • English, 60 characters: 11 tokens.
  • Hindi, 69 characters: 23 tokens - roughly twice the cost for a sentence of the same length.
  • Arabic, 42 characters: 13 tokens.
  • Japanese, 24 characters: 19 tokens - about 1.3 characters per token, compact on screen and expensive to send.

The gap has been closing. That same Hindi sentence costs 69 tokens under the older cl100k_base and 23 under o200k_base - a three-fold improvement, because the newer vocabulary spends some of its extra 100,000 entries on non-Latin scripts. If you work in a language other than English, the encoding matters far more to your bill than it does for English speakers.

Which encoding your model uses

OpenAI publishes its tokenizer, tiktoken, with the vocabulary files. Four encodings cover everything it has shipped:

  • o200k_base - GPT-5, GPT-4.1, GPT-4o and the o-series reasoning models.
  • cl100k_base - GPT-4, GPT-4 Turbo, GPT-3.5 Turbo and the version-3 embedding models.
  • p50k_base - text-davinci-002 and 003, and the Codex models.
  • r50k_base - the original GPT-3 base models and GPT-2.

The counter is keyed on the encoding rather than the model name, deliberately: line-ups get renamed and replaced every few months, but the encoding a family uses is fixed for its lifetime. Pick the encoding and the answer survives the next rename.

Why Claude and Gemini counts are estimates, not facts

Anthropic and Google do not publish tokenizer files you can run in a browser, so any tool claiming an exact offline count for Claude or Gemini is guessing. The honest fallback is the rule of thumb their own documentation uses - about four characters per token for English - and that is what the counter shows for those models, labelled as an estimate rather than dressed up as a measurement.

For English prose that rule is usually within 10 to 20 per cent; for code, JSON or any non-Latin script it can be badly wrong, for exactly the reasons above. When the number has to be right, both vendors expose a counting endpoint - Anthropic's /v1/messages/count_tokens and the Gemini API's countTokens - that returns the real figure for a specific model. Use the estimate to size a prompt; use the API to reconcile a bill.

Context windows have to hold the reply too

A context window is the total number of tokens a model can consider at once, and the common mistake is treating it as a budget for the prompt alone. It has to hold your system prompt, the conversation so far, any documents you pasted in, and the model's answer. Fill 127,000 tokens of a 128,000-token window and the model has room for about a sentence before it is cut off. The meter shows what fraction of a window your text fills and how many tokens remain, switching to a hatched bar once you go over; the presets run from 4,096 up to a million.

Turning tokens into money

Providers quote a price per million tokens, input and output at different rates. Because those prices change constantly, the counter takes the rate as an input rather than shipping a table that would go stale.

A worked example. A 4,000-word English document runs to roughly 4,400 tokens at 1.1 tokens per word. At $3 per million input tokens that is about $0.013 to send once - trivially cheap. Send it on every request in a pipeline running a hundred times a day and it becomes about $1.33 a day, or $485 a year, for one document you never cached. Tokens only matter in aggregate, which is why nobody notices them until the invoice.

Count your own text

Paste your prompt into the Token Counter for the exact figure under the OpenAI encodings, an honest estimate for the rest, the per-token breakdown, the share of a context window it fills and - if you enter a rate - what it costs. The encoding tables download to your browser and the tokenizing runs there, so a prompt you are still drafting never leaves your device.

Frequently asked questions

How many words is 1,000 tokens?
About 750 to 900 words of plain English prose, because English runs at roughly 1.1 to 1.3 tokens per word. Treat that as a planning figure and nothing more. The same 1,000 tokens might be only 400 words of dense technical writing full of product names and identifiers, around 250 lines of nothing but JSON punctuation, or a far shorter passage in Hindi or Japanese. Whenever the number actually matters - a context limit you are close to, or a cost you are forecasting - measure the real text rather than converting from a word count.
Why is my token count different from another site's?
Almost always because the two tools are using different encodings. The same sentence is 84 tokens under o200k_base and 88 under cl100k_base, so a site pinned to GPT-4's tokenizer will disagree with one set to GPT-4o. The other common cause is that the other site is estimating rather than tokenizing - a characters-divided-by-four figure presented without a caveat looks identical to a real count on screen. Check which encoding each tool names; if it does not name one, it is probably guessing.
Do tokens include spaces and punctuation?
Yes, and the way they are grouped surprises people. A leading space is usually absorbed into the word that follows, so " the" is one token rather than a space plus a word - which is why splitting text on spaces and counting the pieces gives the wrong answer. Punctuation is often merged into runs: in JSON, a sequence such as a quote, comma and brace together can be a single token. Line breaks cost tokens too, so heavily indented code and pretty-printed JSON are measurably more expensive to send than their minified equivalents.