Subword Tokenization & Byte-Pair Encoding Mechanics
Large Language Models do not ingest raw characters or whole words directly. Instead, tokenizers segment strings into variable-length character chunks called tokens using algorithms such as Byte-Pair Encoding (BPE) or WordPiece.
1. Token Approximation Equation
Estimated Tokens=
Character Count4
≈Word Count × 1.332. Prompt Ingestion Cost Formula
Ingestion Cost ($)=
Total Tokens1,000,000
×Input Rate per Million ($)Step-by-Step Calculation Breakdown
Step 1: Estimate Tokens from 2,000 Characters
Estimated Tokens = 2,000 ÷ 4 = 500 Tokens
Step 2: Calculate Ingestion Cost on GPT-4o ($2.50 / 1M Tokens)
GPT-4o Cost = (500 ÷ 1,000,000) × $2.50 = $0.00125
Step 3: Compare to Gemini 1.5 Flash ($0.075 / 1M Tokens)
Gemini Flash Cost=(500 ÷ 1,000,000) × $0.075=$0.0000375 (33× Cheaper)
Tokenization Ratios Across Data Types
| Text Content Type | Avg Tokens / Word | Avg Chars / Token | Token Efficiency Note |
|---|---|---|---|
| Standard English Prose | 1.33 | 4.0 | Optimal compression ratio |
| JSON & YAML Payloads | 1.8 – 2.2 | 2.5 – 3.0 | Punctuation & quotes increase token count |
| Source Code (Python/TS) | 1.6 – 2.0 | 3.0 – 3.5 | Indentation & camelCase split into subwords |