A keyword scanner searched 207,391 AI-agent posts for prompt injection and reported 9.87%. The real figure was 1.48%. Nothing changed but three lines of matching logic. Type into the box to watch the bug happen live — then read the 10,176 false positives it produced, every one of them a real item from the corpus.
Every keyword is an instruction-shaped phrase — something an attacker would write at a model. None of them are topic words. prompt injection, attack and exploit are deliberately absent, so writing about injection does not trip the scanner. Keywords marked Aa are matched in capitals only.
dan
Almost all of them inside redundant, dance, dangerous, redundancy. Not one was an attack.
Three quarters of every finding came from eight short single-word keywords — the ones a substring test mangles worst.
Reported in good faith from the broken scanner, and wrong. The dataset card now says so in the open.
Same corpus, same 56 keywords, same day. A 6.7× correction produced by matching logic alone.
A detector can manufacture the incident rate it claims to measure. Nothing about the 9.87% looked wrong. It came from a real corpus, a defensible keyword list, and code that ran without error. The number was plausible, reproducible, and false.
This is why the broken scanner's output ships alongside the fixed one in the published dataset rather than being quietly replaced. The pair is the most useful thing in it: a worked example of a measurement inflating 6.7× with no visible symptom.
Two honest numbers, both with stated denominators: 1.48% matched any keyword — an upper bound, because an agent discussing injection matches the same words as one performing it. 0.083% matched high-confidence phrases — a lower bound. The truth sits between them, and finding it needs a human to hand-label a sample. Nobody has yet.
On the handles shown throughout this page. Agent handles are published as they appear in the source corpora, matching the dataset cards. Some agent accounts are linked to real people — several platforms verify ownership through a social-media account. Do not use this page to target individual accounts. On the false-positive tab in particular, every handle belongs to an agent that was flagged in error by a broken scanner and did nothing wrong.