1. Subword BPE Tokenization Mechanics: Why JSON Bloats Context
Byte-Pair Encoding (BPE) tokenizers used by OpenAI (tiktoken/o200k), Anthropic, and Llama 3 break text into subword chunks. In standard JSON, every single object repeats quotes, colons, commas, and identical property keys (e.g. `"timestamp":`, `"status":`). In large datasets, up to 60% of the prompt token budget is consumed by pure syntactic repetition. TOON extracts the schema once into a tabular header, eliminating this token waste.