Skip to main content
BenchLM

LLM Token Counter & AI Token Calculator

Workspace

Paste any text and see the token count with the cost per model beside it, for every major LLM side by side.

Your Text

0 characters0 words~0 tokens

Select Models

Token counts & costs

Type or paste text in the input area to see token counts.

What Is a Token?

A token is the smallest unit of text that an LLM processes. Tokenizers split text into subword pieces using algorithms like Byte-Pair Encoding (BPE). Common words are usually a single token, while rare or long words get split into multiple tokens.

When tiktoken loads, OpenAI-labeled rows use one encoding, with a character-based fallback. Other families use character-based estimates. These text-only counts exclude request formatting and tool schemas. Cost estimates use current LLM API pricing. New to tokens? Read what an AI token is and what one costs.

"The cat sat" → ["The", " cat", " sat"] = 3 tokens

"Hamburger" → ["Ham", "bur", "ger"] = 3 tokens

"AI" → ["AI"] = 1 token

Rule of thumb: 1 token ≈ 4 characters ≈ 0.75 words in English. Code and non-Latin scripts may tokenize differently. For the full story, read our guide: How LLM Token Pricing Works.

Questions

What is a token in an LLM?

A token is the basic unit of text that large language models process. Tokens are subword units created by a tokenizer algorithm. In English, one token is roughly 4 characters or 0.75 words. Common words like "the" are single tokens, while longer or rarer words may be split into multiple tokens.

How many tokens are in 1,000 words?

Approximately 1,333 tokens for typical English text. The exact count depends on the model's tokenizer and the vocabulary used — technical jargon and code tend to produce more tokens per word than everyday English.

Do different AI models count tokens differently?

Tokenizers can split the same text differently. This tool uses one loaded tiktoken encoding for OpenAI-labeled rows and character-based estimates for other families. Check the serving API usage before treating a text estimate as a billable token count.

How do I count tokens before making an API call?

Use this free token counter tool to check counts across multiple models instantly. For programmatic use, OpenAI provides tiktoken, Anthropic offers a count_tokens endpoint, and Google provides a CountTokens API. All provider counting endpoints are free to use.

Why does token count matter for LLM costs?

LLM APIs charge per token processed. Both your input (prompt) and the model's output are billed separately, with output tokens typically costing 2-5x more than input tokens. Knowing your token count helps you estimate costs and choose the right model for your budget.

What is BPE tokenization?

Byte-Pair Encoding (BPE) is the tokenization algorithm used by most modern LLMs including GPT, Claude, and Llama. It breaks text into subword units based on frequency patterns learned during training, balancing vocabulary size with the ability to handle rare and novel words.

Stay on top of LLM changes

Get notified when new models launch, pricing changes, or benchmarks update.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.