Rates Simulation difficulty before judging results
Scores each Simulation's difficulty (for example, with the NIST Phish Scale) and compares results only at similar difficulty.
Why it matters
Difficulty predicted behavior in replication research while training did not. Without difficulty, a falling click rate may just mean easier emails.
How to observe it
Difficulty rating recorded for every Simulation; trend lines reported by difficulty band.
Evidence
Supporting (2)
- Anti-Phishing Training (Still) Does Not Work: A Reproduction of Phishing Training Inefficacy Grounded in the NIST Phish Scale
ChallengesTraining interventions had no significant effect on click rates or reporting rates, with negligible effect sizes. Phish Scale difficulty did predict behavior (7% clicks on easy emails vs 15% on hard), and in 36-55% of campaigns reports arrived before clicks, though training did not improve this.
- Categorizing human phishing difficulty: a Phish Scale
FoundationalClick rates should be expected to vary with how hard a phishing email is for a given audience, especially when its premise fits the recipient's work context. The authors propose the NIST Phish Scale so programs can rate exercise difficulty and interpret click rates.