What counts, what it costs, and what falls out of the window
Tokens are the unit of both billing and memory. Estimating them well is most of cost control and all of context management.
| Content | Rough tokens | Note |
|---|---|---|
| English prose | ~1 token per 4 characters | About 0.75 tokens per word |
| Code | ~1 token per 3 characters | Punctuation and indentation are dense |
| JSON | Higher still | Braces, quotes and keys all cost |
| Non-Latin scripts | 2–3× more per character | Budget generously for CJK and Arabic |
| An A4 page of text | ~500–700 tokens | Useful for sizing document pipelines |
| Strategy | Keeps | Loses |
|---|---|---|
| Sliding window | The last N turns | Anything agreed early on |
| Summarise-and-drop | A rolling summary plus recent turns | Detail, gradually |
| Pinned + recent | Explicitly marked turns, plus recent | Needs someone to choose the pins |
| Retrieval over history | Whatever matches the current turn | Continuity between turns |
No. Cost scales with what you send, latency grows with it, and retrieval accuracy over a very long context is not uniform. A big window buys headroom, not a free pass.
Tokens are sub-word pieces. English prose runs about 0.75 tokens per word; code, JSON and non-Latin scripts run considerably higher because punctuation and rare characters split into more pieces.