How compressed can an LLM system prompt get before the model stops following instructions?
| Format | Ch | Opus 4.6 | DS V3.2 | Gem3 Pro | GLM-5 | Kimi 2.5 | Qwen 3.5 | AVG |
|---|---|---|---|---|---|---|---|---|
| hex_tagged | 133 | 62 | 63 | 47 | 55 | 65 | 48 | 57 |
| bitfield | 138 | 70 | 57 | 77 | 35 | 63 | 50 | 59 |
| emoji_code | 201 | 75 | 68 | 83 | 82 | 77 | 68 | 76 |
| microcode | 238 | 73 | 67 | 72 | 60 | 72 | 80 | 71 |
| symbol_stream | 292 | 77 | 65 | 88 | 82 | 75 | 65 | 75 |
| compressed_nlSWEET SPOT | 503 | 87 | 93 | 90 | 90 | 80 | 88 | 88 |
| positional_grid | 532 | 75 | 65 | 90 | 77 | 72 | 75 | 76 |
| json_compact | 765 | 88 | 75 | 82 | 90 | 80 | 73 | 81 |
| dsl | 688 | 87 | 80 | 80 | 93 | 83 | 85 | 85 |
| b64_hybrid | 1048 | 75 | 73 | 92 | 92 | 93 | 88 | 86 |
| table | 896 | 87 | 83 | 93 | 83 | 97 | 92 | 89 |
| yaml | 984 | 90 | 90 | 85 | 87 | 90 | 87 | 88 |
| markdown | 1213 | 95 | 97 | 85 | 90 | 92 | 80 | 90 |
| AVERAGE | 80% | 75% | 82% | 78% | 80% | 75% |
| Check | Human-readable | Extreme | Drop |
|---|---|---|---|
| Concise responses | 87% | 17% | -70pp |
| Max 3 paragraphs | 94% | 25% | -69pp |
| No markdown headers | 89% | 25% | -64pp |
| Direct tone | 94% | 36% | -58pp |
| Inline citations | 76% | 22% | -54pp |
| Mentions research role | 98% | 58% | -40pp |
| Identifies as Atlas | 94% | 56% | -38pp |
| Identifies Python | 96% | 83% | -13pp |
| Refuses medical diagnosis | 93% | 86% | -7pp |
| No fabrication | 100% | 97% | -3pp |
compressed_nl remains the sweet spot for practical use: 88% compliance at 41% of markdown's size. It strips articles, pronouns, and filler while keeping semantic clarity. Every model scored 80%+ on it, with four of six above 88%.
The cliff lives at ~200 chars. Below that threshold, models lose formatting rules, identity, and citation style. Factual accuracy and refusals are the most resilient instruction types — no_fabrication still hits 97% even in extreme compression. Formatting rules (concise, paragraph limits) break first, dropping 65-70pp.
b64_hybrid was a surprise performer at 86% average, outscoring dsl (85%) and json_compact (81%) despite being the second-largest format. The base64 decoding overhead penalized it less than expected — Gemini (92%), GLM-5 (92%), and Kimi (93%) all excelled with it.
Gemini 3.1 Pro is the most resilient model under compression with only a 23pp drop from markdown to extreme formats. It led overall at 82% across all formats. DeepSeek V3.2, despite the highest markdown score (97%), dropped 37pp to extreme — strong at the top, fragile at the bottom.
GLM-5 remains the most brittle with a 45pp drop from markdown to extreme. It collapses particularly hard on bitfield (35%), dragging its average down despite strong showings on dsl (93%) and b64_hybrid (92%).