AI / A CONCEPT NOTE
Tokenization
how LLMs break text into pieces their neural networks can process
Overview · mechanism
pitfall · examples
01 / THE SHORT VERSION
The idea in a few sentences.
Tokenization is the first step of any LLM pipeline. Text is split into tokens — subwords, not whole words. 'unbelievable' becomes ['un', 'believe', 'able']. Each token maps to an integer ID in the model's vocabulary (typically 32K-200K tokens). Tokenization determines how the model 'sees' your text and directly affects cost, latency, and even comprehension.
02 / FOLLOW THE MECHANISM
How tokenization breaks down text
Raw text
"I'm running Kubernetes on AWS" enters the tokenizer.
BPE tokenizer
Byte-Pair Encoding scans the text for common subword patterns:
['I', "'", 'm', ' run', 'ning', ' Ku', 'ber', 'net', 'es', ' on', ' AWS'].Token IDs
each token is mapped to an integer:
[40, 466, 12, 1285, 596, 17283, 29768, 4793, 659, 1649, 21371].Model input
the token IDs are passed through the embedding layer — the model never sees raw text, only these integer sequences.
04 / COMMAND NOTES
Read the command, then the result.
Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.
count tokens in a string with tiktoken
python -c "import tiktoken; enc = tiktoken.get_encoding('cl100k_base'); print(len(enc.encode('Hello world')))"count tokens with js-tiktoken
node -e "const {getEncoding} = require('js-tiktoken'); const enc = getEncoding('cl100k_base'); console.log(enc.encode('Hello world').length)"05 / CHECK YOURSELF
Could you explain Tokenization to a teammate?
Try it out loud in two sentences: what it is, and the one detail that changes the picture. If you stall, the gap is the part to reread.
Up next in AI engineeringPrompt Cachingoptimizing LLM API latency and token reuse costs