AI / A CONCEPT NOTE

Tokenization

how LLMs break text into pieces their neural networks can process

~75 sec read

Overview · mechanism
pitfall · examples

01 / THE SHORT VERSION

The idea in a few sentences.

Tokenization is the first step of any LLM pipeline. Text is split into tokens — subwords, not whole words. 'unbelievable' becomes ['un', 'believe', 'able']. Each token maps to an integer ID in the model's vocabulary (typically 32K-200K tokens). Tokenization determines how the model 'sees' your text and directly affects cost, latency, and even comprehension.

02 / FOLLOW THE MECHANISM

How tokenization breaks down text

  1. Raw text

    "I'm running Kubernetes on AWS" enters the tokenizer.

  2. BPE tokenizer

    Byte-Pair Encoding scans the text for common subword patterns: ['I', "'", 'm', ' run', 'ning', ' Ku', 'ber', 'net', 'es', ' on', ' AWS'].

  3. Token IDs

    each token is mapped to an integer: [40, 466, 12, 1285, 596, 17283, 29768, 4793, 659, 1649, 21371].

  4. Model input

    the token IDs are passed through the embedding layer — the model never sees raw text, only these integer sequences.

04 / COMMAND NOTES

Read the command, then the result.

Inspect the flags and arguments before trying an example. Snippets can need local setup, replacement values, or resources in your own environment.

EXAMPLE 01 · REFERENCE

count tokens in a string with tiktoken

python -c "import tiktoken; enc = tiktoken.get_encoding('cl100k_base'); print(len(enc.encode('Hello world')))"

EXAMPLE 02 · REFERENCE

count tokens with js-tiktoken

node -e "const {getEncoding} = require('js-tiktoken'); const enc = getEncoding('cl100k_base'); console.log(enc.encode('Hello world').length)"

Explore command anatomy in the CLI lab