The Long Run
← Back to blog

Data Science

Tokens 101: The Secret Language Your AI Actually Speaks

A beginner-friendly introduction to what tokens are, how they differ from words and characters, and why they matter when using LLMs like ChatGPT and Claude.

·8 min read

Part 1 of 4 in the "LLM & Tokenization" series

"Token" Why Should You Care About This Word?

If you have ever chatted with ChatGPT or Claude, you have probably bumped into the word token. Maybe you saw it in a pricing page ("$0.003 per 1K tokens") and thought, "Tokens? Are we buying gaming credits now?"

Here is the twist: you are not far off from thinking that way. Tokens really are the "currency" your AI spends every time it reads your message or writes a reply.

Understanding more about tokens, API costs, "context window" limits helps to understand why your chatbot sometimes cuts off mid-sentence, all tha you were struggling starts to make sense.

This is Part 1 of a 4-part series that takes you from "what even is a token" all the way to "how do context windows work." Let's dive in and unravel this concepts.

How You Actually Talk to an LLM

When you use an LLM, you type a normal sentence, like:

"Who moved my cheese?"

You read this the way as a normal humans do i.e. left to right, one word flowing into the next, no extra thought required. But a Large Language Model (LLM) doesn't "read" the way you do. It can't glance at a sentence and instantly grasp it as whole words the way your brain does.

Instead, before the model can understand anything, the sentence has to be chopped up into smaller pieces it can actually process. That chopping-up process of a large sentence into smaller pieces is called tokenization, and the pieces it produces are called tokens.

So, What Exactly Is a Token?

A token is the basic unit of text that a model processes. Depending on the tokenizer, a token might be:

  • A whole word e.g., cheese
  • A sub-word chunk e.g., mov + ed (from "moved")
  • A single character e.g., c, h, e
  • Punctuation or special symbols e.g., ?, !, or invisible markers like "end of sentence"

In short: tokenization breaks a sentence into small, manageable chunks so the model can "read" it. Our example sentence could be tokenized in more than one way:

Who | moved | my | cheese | ?

or, if the tokenizer is being extra thorough:

W | h | o | m | o | v | e | d | m | y | c | h | e | e | s | e | ?

Neither version is "wrong", it just depends on the model the tokenizer uses.

Token vs. Word vs. Character: They are no synonymous

This is where a lot of us get tripped up, so let's untangle it clearly.

ConceptWhat it meansExampleCarries full meaning?
CharacterA single letter, digit, or symbolc, 7, $No
WordA complete unit with standalone meaningcheeseYes
TokenWhatever chunk the tokenizer decides on could be a word, part of a word, or a characterche + ese, or cheese, or c+h+e...Sometimes

A token is not a fixed-size thing. It's flexible by design. Sometimes it lines up with a full word, sometimes it's just a fragment. That flexibility is exactly what makes tokenization powerful (more on that in Part 2).

Same sentence, wildly different token counts depending on the tokenization strategy The chart above shows how the same 5-word sentence, "Who moved my cheese?", produces a different number of tokens depending on whether you tokenize by word, by subword (closer to how real LLMs work), or by character.

Note that: Character-level tokenization creates far more tokens for the exact same sentence.

How to read this chart:

  • X-axis shows three tokenization strategies.
  • Y-axis shows the total number of tokens produced for the same sentence.
  • Taller bars mean more tokens are needed to represent identical text, which matters because more tokens usually means more processing cost and a bigger bite out of your token limit (we will get to that in Part 3).

How to build a similar chart with your own sentence?

Simple steps to replicate the recipe:

  1. Pick a sentence and tokenize it three ways: by word, by subword (you can use an online tokenizer visualizer for GPT/Claude-style models), and by character.
  2. Count the tokens produced by each method.
  3. Plot the three counts as bars using any spreadsheet or charting tool (Excel, Google Sheets, or Python's matplotlib).
  4. Compare the heights, you will see subword tokenization usually lands in a comfortable middle ground.

Tokens Aren't Meant for Humans

Here's a fun (and slightly humbling) fact: Tokens are not designed to be intuitive to you i.e. Humans. They are designed to be efficient for the machine. A human reading ["mov", "ed"] instead of "moved" would find it clunky and unnecessary. But for a model doing math on probabilities across millions of examples, breaking "moved" into reusable chunks like mov and ed is far more efficient because those same chunks can be reused in "moving," "movement," "removed," and so on.

Think of it like IKEA furniture. You and I want a finished bookshelf. IKEA ships you panels, screws, and an Allen key. Less intuitive to unbox? Sure. But wildly more efficient to manufacture, store, and ship at scale. Tokens are the Allen-key version of language.

Retail Example #1: Why a Chatbot "Sees" Your Query Differently Than You Do

Imagine a fictional online electronics retailer called SmallBazaar. A customer types into SmallBazaar's support chatbot:

"My noise-cancelling headphones aren't pairing with my phone."

You read that as one clear complaint. SmallBazaar's LLM-powered assistant, however, sees something closer to:

My | noise | - | cancel | ling | headphones | aren | ' | t | pairing | with | my | phone | .

That's roughly 14 tokens for a 10-word sentence, words like "noise-cancelling" and "aren't" get split into multiple pieces because they are less common as single units. This is exactly why SmallBazaar's engineering team monitors average tokens per customer query, not just word count, when estimating their monthly AI support costs.

Key Takeaways

  • A token is the basic unit of text an LLM processes.It can be a word, part of a word, a character, or punctuation.
  • Tokenization is the process of breaking text into these units so the model can read it.
  • Tokens, words, and characters are 3 different things, don't use them interchangeably.
  • Tokenization is optimized for machine efficiency, not human readability.
  • The number of tokens a sentence produces depends entirely on which tokenization strategy (algorithm or model) is used, and that number has real cost and performance implications.

Up next in Part 2: We crack open the hood and look at how tokenizers actually decide where to make the cuts including a walkthrough of BPE, WordPiece, SentencePiece, and Unigram, with a genuinely funny example thrown in.

In case of any queries or feedback feel free to drop comments or reach out using the Contact secion.

Comments