The same 250-document backdoor that compromised LLMs from 600M to 13B params, plotted across model scale. Toggle 'documents needed' — a flat 250 vs the myth that poison scales with data — and 'share of training set', collapsing to 0.00016% and below. Drag to pick a scale.
A landmark 2025 result: ~250 poisoned documents backdoor an LLM whether it has 600M or 13B parameters — 0.00016% of the tokens. DataDAOs sell 'verifiable' training data, but on-chain provenance proves integrity, not purity. Here's the gap, and what actually narrows it.