Tutorials · · 1,072 words · 5 min read
Counting tokens before the bill: six tokenizers, six kinds of text
The same 20,000 characters came to 4,716 to 33,502 tokens depending on language and tokenizer. Measured counts, the divide-by-four rule tested, a forecast script.
tokenizers cost tiktoken forecasting
An API bill is tokens times price, and the price is the easy half: it sits on a pricing page. The token count is the part people estimate, usually by dividing the character count by four. This tutorial measures what that shortcut gets right and wrong. We counted the same six pieces of text with six public tokenizers, and the answer ranges from 4,716 tokens to 33,502 for 20,000 characters, depending on the language and the tokenizer. The listing code/token-counting-cost-forecast.py reproduces every count from public sources, and it turns a sample request into a monthly token volume and a bill once you supply your provider's prices.
What was measured
Six samples, chosen to cover what pipelines actually send:
- English prose: 20,000 characters of Pride and Prejudice (characters 5,000 to 25,000 of the Project Gutenberg body, line endings normalised to
\n). - German prose: the same slice of Kafka's Die Verwandlung.
- Chinese prose: the same slice of 西遊記 (Journey to the West).
- Python code: the first 20,000 characters of CPython's
json/decoder.pyandjson/encoder.py, from Python 3.13.12. - JSON: one real API response, ESPN's public NFL scoreboard for 24 September 2026, serialised compactly (17,894 characters) and with
indent=2(34,716 characters).
The six tokenizers are the two public OpenAI encodings in tiktoken 0.14.0 (cl100k_base and o200k_base) and the tokenizers shipped with four open-weight models: Qwen2.5, DeepSeek-V3, Mistral-7B-v0.3 and Phi-3.5-mini. Counts exclude special tokens and chat templates, so a real request adds a few more.
| Sample | Characters | cl100k_base | o200k_base | Qwen2.5 | DeepSeek-V3 | Mistral-7B-v0.3 | Phi-3.5-mini |
|---|---|---|---|---|---|---|---|
| English prose | 20,000 | 4,772 | 4,716 | 4,770 | 4,735 | 5,186 | 5,268 |
| German prose | 20,000 | 5,990 | 5,114 | 5,924 | 5,797 | 6,800 | 6,244 |
| Chinese prose | 20,000 | 29,829 | 20,576 | 18,811 | 17,073 | 28,221 | 33,502 |
| Python code | 20,000 | 4,726 | 4,780 | 4,758 | 5,029 | 5,887 | 5,882 |
| JSON, compact | 17,894 | 6,166 | 6,393 | 7,360 | 6,693 | 8,436 | 8,381 |
| JSON, indent=2 | 34,716 | 8,728 | 8,872 | 9,922 | 9,092 | 12,137 | 12,042 |
English prose and code: the rule of four mostly holds
For English prose, 20,000 characters divided by four is 5,000, and the six tokenizers counted between 4,716 and 5,268. The shortcut is within about 6% either way: 6.0% high against o200k_base, 5.1% low against Phi-3.5-mini. Python code behaves much the same, from 4,726 tokens with cl100k_base to 5,887 with Mistral-7B-v0.3.
Vocabulary size shows up even here. The two tokenizers with vocabularies of about 32,000 needed 8.7% to 11.7% more tokens for the same English than the four with 100,000 or more, and 17.0% to 24.6% more for the Python code. A bigger vocabulary stores more whole words and common code fragments as single tokens.
Other languages: the rule breaks
German took 5,114 tokens with o200k_base and 6,800 with Mistral-7B-v0.3, so the divide-by-four estimate of 5,000 undercounts by 2.2% to 26.5%. The newer OpenAI encoding needed 14.6% fewer tokens than cl100k_base for the same German.
Chinese is where the shortcut fails outright. Each character carries far more than an English letter, and the 20,000 characters became 17,073 tokens with DeepSeek-V3 and 33,502 with Phi-3.5-mini: 3.4 to 6.7 times the estimate. The choice of tokenizer alone is worth almost a factor of two, and o200k_base needed 31.0% fewer tokens than cl100k_base. For the same number of characters, a Chinese request bills 3.6 times as many input tokens as an English one under DeepSeek-V3 and 6.4 times as many under Phi-3.5-mini.
JSON is expensive, and indentation makes it worse
The compact scoreboard is mostly short keys, quotes, braces and numbers, and it ran at 2.12 to 2.90 characters per token. The divide-by-four estimate of 4,474 tokens undercounts by 27.4% to 47.0%. Pretty-printing the same response with two-space indentation nearly doubled the characters, a 94.0% increase, but raised the token count by 34.8% to 43.9%, because runs of spaces merge into a few tokens. Either way the lesson for a pipeline that puts tool results or records into a prompt is the same: serialise compactly, and drop the fields the model does not need, before worrying about anything else.
Forecasting a bill
The arithmetic is one line:
cost = requests x (input_tokens x input_price + output_tokens x output_price) / 1,000,000
with prices in dollars per million tokens, as providers publish them. The only uncertain term is input_tokens, so measure it on a real request instead of estimating it. The listing does this. Save one typical request's input text to a file and run:
python token-counting-cost-forecast.py forecast --prompt request.txt --output-tokens 300 --requests 100000 --price-in 1 --price-out 1
With both prices set to 1, the dollar column reads as millions of tokens, which makes the output easy to scale by your real prices. Using the 20,000-character English sample as the request, 100,000 requests a month came to 471.6 to 526.8 million input tokens across the six tokenizers, and $501.60 to $556.80 per dollar of price once the 30 million output tokens are added, a spread of 11%. The same run on the Chinese sample came to 1,707.3 to 3,350.2 million input tokens, and $1,737.30 to $3,380.20, a spread of 1.95 times. Without --tokenizer, the script prints all six and the range. That range is the honest forecast when your provider does not publish its tokenizer; if yours offers a token-counting endpoint, use it for the final number.
The dataset
datasets/token-counting-cost-forecast.csv holds all 36 counts, produced by python token-counting-cost-forecast.py reproduce (add --save-samples DIR to write the six sample texts for the other two commands).
_README:
- content: english, german, chinese, python code, json compact, or json indent=2
- tokenizer: cl100k_base, o200k_base, Qwen2.5, DeepSeek-V3, Mistral-7B-v0.3 or Phi-3.5-mini
- characters: length of the sample in Unicode characters
- tokens: tokens the tokenizer produced, special tokens excluded
- chars_per_token: characters / tokens
- chars4_error_pct: how far the divide-by-four estimate is from the real count, (characters / 4) / tokens - 1, in percent; negative means the shortcut undercounts
What the measurement does not cover
One slice of one book per language is a sample, not a language: technical German or classical versus modern Chinese will land somewhere else, so measure your own text. The JSON is a single response from a single API, and a live response for the same date can drift as the provider changes it. Special tokens, chat templates and tool schemas are left out, and for short requests they are not negligible. Most importantly, some hosted models do not publish their tokenizers at all, so none of these six is guaranteed to match a given bill. What the table does show is how wide the uncertainty is, and which content types deserve a real count before a budget is set.
Code and data
- token-counting-cost-forecast.py — the complete listing used in this article.
- token-counting-cost-forecast.csv — the data behind the numbers here.
Sources
- OpenAI, "tiktoken" BPE tokenizer library, version 0.14.0 (encodings cl100k_base and o200k_base), run 2026-09-25
- Qwen, "Qwen2.5-7B-Instruct" model repository, tokenizer.json (vocabulary 151,665), downloaded 2026-09-25
- DeepSeek, "DeepSeek-V3" model repository, tokenizer.json (vocabulary 128,815), downloaded 2026-09-25
- Mistral AI, "Mistral-7B-Instruct-v0.3" model repository, tokenizer.json (vocabulary 32,768), downloaded 2026-09-25
- Microsoft, "Phi-3.5-mini-instruct" model repository, tokenizer.json (vocabulary 32,011), downloaded 2026-09-25
- Project Gutenberg eBook 1342, Jane Austen, "Pride and Prejudice" (English)
- Project Gutenberg eBook 22367, Franz Kafka, "Die Verwandlung" (German)
- Project Gutenberg eBook 23962, Wu Cheng'en, "西遊記" (Journey to the West, Chinese)
- Python Software Foundation, CPython 3.13 json module (decoder.py and encoder.py), as installed with Python 3.13.12
- ESPN, public NFL scoreboard endpoint, response for 24 September 2026, pulled 2026-09-25