Book online or chat with usAnswered 24/7Licensed, insured & bondedSchedule a Consultation
Request serviceUrgentConsultation

Phishing Awareness Training That Reduces Clicks

Phishing awareness training reduces clicks when it is measured by reporting behavior rather than click rate alone, runs continuously instead of annually, and is paired with technical controls that assume some employees will always click. NIST research shows click rates rise mainly because a lure fits the target’s work context, not because training failed. Report rate, and how fast reports arrive, is the signal that matters.

Most organizations already run phishing training. Far fewer can say whether it works. The usual evidence is a click-rate chart that trends down for two quarters, then jumps after a harder simulation, at which point somebody asks whether the training was worth the money. The chart is not the problem. The metric is. Click rate, read on its own, tells you more about the difficulty of the email you sent than about the judgment of the people who received it.

This guide is written for the executive, general counsel, or IT leader who has to defend a security awareness budget and wants a program that changes behavior rather than one that generates a compliance certificate. It covers what the click rate hides, why reporting is the metric worth managing, how to design simulations that teach instead of humiliate, and which technical controls have to sit behind any human-layer program.

Why does phishing still work after two decades of training?

Because it remains the cheapest reliable way into an organization, and because the volume is enormous. In the FBI Internet Crime Complaint Center’s 2025 annual report, phishing and spoofing was the most-reported crime type by complaint volume, with 191,561 complaints out of 1,008,597 total complaints and $20.877 billion in reported losses across all categories. Business email compromise, which almost always begins with a phishing or credential-theft step, accounted for 24,768 complaints and $3,046,598,558 in reported losses in the same year.

Those figures describe reported incidents only, and they describe outcomes rather than attempts. The operational point is narrower: attackers do not need a clever technical exploit when a convincing message and a plausible pretext will do. That is why the human layer keeps getting attention, and why training that only teaches people to look for bad spelling is obsolete. A well-built lure aimed at a specific department contains no spelling errors at all.

For a fuller treatment of the attack techniques themselves, including vishing, pretexting, and physical tailgating, see our guide on how social engineering bypasses your firewall. This article is about the program you build in response.

What does a click rate actually measure?

The most useful public work on this question is the NIST Phish Scale, published as NIST Technical Note 2276 in November 2023. It was built for exactly the problem security teams keep running into: two simulations produce wildly different click rates, and nobody can say whether the workforce got better, got worse, or simply received a harder email.

The Phish Scale rates how hard a given message is for a human to detect, using two inputs. The first is a count of observable cues, 21 of them, grouped into five categories: errors such as spelling and grammar mistakes; technical indicators such as spoofed sender addresses, hidden URLs, and suspicious attachment types; visual presentation problems such as missing or imitated branding; language and content signals such as manufactured urgency, generic greetings, and requests for sensitive data; and common tactics such as humanitarian appeals, unrealistic offers, time limits, and impersonation of authority. More cues make a message easier to spot.

The second input is premise alignment: how well the story in the email matches the recipient’s actual work context. High alignment means the premise maps onto the person’s real responsibilities or a genuine organizational event. Low alignment means the premise is irrelevant to them. NIST combines cue count, binned as few, some, or many, with premise alignment in a three-by-three matrix to produce a detection difficulty rating of least difficult, moderately difficult, or very difficult.

The finding that matters most for program design is that premise alignment carries more weight than cue count. A message with several visible flaws will still catch people if its story lands squarely in their job. An invoice-approval lure sent to accounts payable during month-end close is hard to detect even when it is technically sloppy. The same email sent to the warehouse is easy. Two teams, one email, two very different click rates, and no difference at all in training quality.

How should each simulation result be interpreted?

Once difficulty is scored before the send, results become readable. The table below shows how to interpret an outcome against the difficulty of the lure, and what each combination should trigger.

Lure difficultyHigh clicks, low reportsHigh clicks, high reportsLow clicks, low reports
Least difficult
many cues, low premise alignment
Genuine gap. Basic recognition is not there. Return to fundamentals for that group.People are deceived briefly but recover and escalate. Reinforce, do not alarm.Expected. This tells you almost nothing. Raise difficulty next cycle.
Moderately difficult
mixed cues and alignment
Reporting habit is missing. Fix the reporting path before adding more content.Healthy. The workforce is behaving as designed.Either strong performance or an email that never reached inboxes. Verify delivery.
Very difficult
few cues, high premise alignment
Normal clicks, abnormal silence. Reporting friction is the problem to solve.Best realistic outcome. Detection will not reach zero here.Confirm the message was actually delivered and rendered before celebrating.

Read the columns, not just the click number. NIST’s own guidance is that click rates have to be considered alongside reporting behavior. A rising click rate on a very difficult lure alongside a rising report rate is a program working correctly. A flat click rate with no reports at all is a program that has taught people to stay quiet.

Why is report rate the metric worth managing?

Because clicks cannot be driven to zero, and the response window can be driven to minutes. Assume a determined attacker will eventually write a message that fits someone’s work context closely enough to get a click. What determines whether that click becomes an incident is how quickly somebody tells the security team, and whether the team can act on it.

That reframes the goal. The program is not trying to produce a workforce that never errs. It is trying to produce a workforce that reports fast, including the people who clicked. The second group is the harder one to win, because reporting your own mistake requires believing nothing bad will happen to you for it.

Three measures are worth tracking every cycle, and they beat click rate individually and together:

  • Report rate: the share of recipients who escalated the message through the official channel.
  • Time to first report: minutes between delivery and the first escalation. This is the number that determines containment.
  • Self-report rate after a click: the share of people who clicked and then said so. This is the clearest available proxy for whether your security culture is punitive.

Track all three by department and role rather than as a single company-wide figure. Finance, executive assistants, HR, and anyone with vendor or payment authority face different lures than the rest of the organization and deserve to be measured separately.

What does NIST say a training program should look like?

In September 2024 NIST published Special Publication 800-50 Revision 1, Building a Cybersecurity and Privacy Learning Program. It supersedes the 2003 edition of SP 800-50 and SP 800-16 from 1998, and the shift in framing is the substance of the update. The revision takes a lifecycle approach, is explicitly written to serve organizations of any size including those starting from nothing, and treats behavior change as part of risk management rather than treating training as an annual compliance event. It also asks organizations to build in metrics and evaluation so the program improves over time.

CISA’s small-business guidance lands in the same place in plainer language: threats change constantly, so once-a-year training is not enough, and employees need to know exactly to whom and how to report a suspicious message. CISA also recommends designating someone, an internal lead or an IT provider, to track emerging threats and brief the team between formal training sessions, and it points organizations to no-cost tabletop exercises they can adapt.

One specific habit from that guidance is worth teaching directly because it defeats most pretexting: verify out of band. If a message asks for money, credentials, or a change to payment details, confirm it using a phone number or address you already have, never a number or link contained in the message itself.

A nine-step program that changes behavior

The following sequence reflects how we build human-layer programs for clients and what tends to separate programs that hold up from programs that produce charts.

  1. Define the behavior, not the score. Write down the two or three actions you want: report suspicious mail through one channel, verify payment and credential requests out of band, never reuse a password across systems. Everything else is in service of those.
  2. Make reporting effortless. One button in the mail client, or one address that everyone knows. If reporting takes more than a few seconds, your report rate is measuring friction, not culture.
  3. Acknowledge every report. An automatic confirmation plus a short human reply on the ones that matter. Reports that vanish into silence stop arriving.
  4. Baseline honestly. Run one moderately difficult simulation before any training, score its difficulty first, and record report rate and time to first report alongside clicks.
  5. Score difficulty before every send. Use the Phish Scale inputs: count the cues, judge premise alignment for the specific audience. Without this, your trend line is not comparable across cycles.
  6. Segment by exposure. Build role-specific lures for finance, executive support, HR, and IT administrators. Generic company-wide sends teach generic lessons.
  7. Train at the moment of the click. A short, specific, non-punitive explanation delivered immediately, naming the cues that were present in that message. Annual slide decks do not survive contact with a real lure.
  8. Remove punishment from the design. No public lists, no manager escalation for a first click, no disciplinary language. The moment clicking becomes dangerous to admit, self-reporting collapses and your detection window widens.
  9. Review quarterly with the difficulty score attached. Present report rate, time to first report, and self-report rate by department, each labeled with the difficulty of the lure that produced it.

Steps two, three, and eight are the ones organizations skip, and they are the ones that determine whether the rest works. A program can have excellent content and still fail because reporting is inconvenient and admitting a mistake feels risky.

Which technical controls have to sit behind the training?

Training is a control, not a perimeter. The joint guidance issued by CISA, the NSA, the FBI, and the MS-ISAC in October 2023, Phishing Guidance: Stopping the Attack Cycle at Phase One, is built on that premise: it addresses network defenders and software manufacturers, with a section written specifically for small and medium-sized businesses, because the human layer cannot be the only layer.

The table below sets out what each layer is realistically responsible for.

LayerWhat it preventsWhat it cannot prevent
Awareness training and simulationsRecognition failures on moderate lures; slow escalation; unverified payment and credential requestsClicks on well-aligned, low-cue lures aimed at a specific role
Phishing-resistant multi-factor authenticationAccount takeover from stolen or relayed credentials, including real-time proxy pagesThe click itself, malware delivered by attachment, or authorized fraudulent payments
Mail authentication and filteringLarge volumes of spoofed and commodity phishing before deliveryTargeted mail from a compromised but legitimate partner account
Endpoint detection and monitoringExecution and lateral movement after a successful clickCredential theft that produces no malware at all
Tested response plan and retainerSlow, improvised containment during the first hoursThe initial compromise

Two of those rows do most of the work. Phishing-resistant multi-factor authentication removes the value of a stolen password, which is what a large share of phishing is actually after. A tested response plan decides what a click costs once it happens. If you are weighing how to resource the response side, our comparison of an incident response retainer versus on-demand engagement covers the trade-offs, and ongoing monitoring is covered in our guide to managed security services.

Payment fraud deserves its own control, separate from email entirely. Business email compromise succeeds at the approval step, not the inbox. A written rule that any new or changed payment instruction is verified by callback to a previously known number, with no exceptions for urgency or seniority, prevents more loss than any amount of inbox vigilance. Our guide to business email compromise sets out the full control set.

How do you run simulations without damaging trust?

Badly run simulations cost more than they teach. The classic failures are a fake bonus announcement, a fake layoff notice, or a fake benefits change. These produce excellent click rates and lasting resentment, and they train employees to distrust internal communications, which is the opposite of the goal.

Three rules keep a program credible. Do not use emotionally exploitative premises that touch pay, employment status, or family. Tell the organization in advance that simulations happen, without saying when; the deterrent value comes from the program existing, not from ambush. And publish what you do with the data, specifically that individual results are not used for discipline.

There is a business reason as well as an ethical one. A workforce that expects blame reports later, and every minute of delay is dwell time an attacker gets to use.

What should improvement look like after twelve months?

Expect the click rate to move less than you hoped and the reporting numbers to move more. A realistic first-year pattern looks like this: click rate on comparable difficulty drifts down modestly; report rate rises substantially; time to first report falls from hours to minutes; and self-reporting after a click becomes routine rather than rare.

The strongest single indicator that the program has taken hold is unprompted reports of real phishing arriving from across the organization, not just from IT. That means people have a channel, they trust it, and they use it without being tested. At that point the human layer is doing the job a technical control cannot: catching the targeted message that got through the filters because it came from a real, compromised partner account.

Honeybadger Solutions builds and runs these programs alongside the technical controls behind them, and handles the forensic work when a lure succeeds. Cybersecurity and digital forensics are delivered in-house, nationally and internationally. You can review our full cyber services, and if an investigation is needed after an incident, our digital forensics team produces court-ready findings under a documented chain of custody.

Program scope varies with headcount, regulatory exposure, and how many high-risk roles you have. Pricing depends on the case. We quote after a short scoping call, in writing, before any work starts.

Frequently asked questions

How often should phishing simulations run?

Monthly or every six weeks for most organizations, with role-specific sends layered on top for finance, executive support, and administrators. CISA is explicit that once-a-year training is not enough because threats change continuously. What matters more than raw frequency is that each send has its difficulty scored beforehand, so results are comparable, and that a short, specific, non-punitive lesson is delivered at the moment of the click.

Is a rising click rate a sign the training is failing?

Not by itself. NIST’s Phish Scale work found that detection difficulty, and particularly how closely a lure’s premise aligns with the recipient’s work context, drives click rates more than cue count does. A harder, better-targeted simulation will produce more clicks from the same workforce. Read the click rate against the scored difficulty of that specific email and against report rate. Clicks up and reports up is a program working.

Should employees face consequences for clicking a simulated phishing email?

No, and building discipline into the program actively damages security. The measurable cost is that self-reporting after a click collapses, which widens the window between compromise and containment. Repeated clicks by someone in a high-risk role is a coaching and access-review question, handled privately. The behavior to reward is reporting, including reporting your own mistake.

Can training alone prevent business email compromise?

No. BEC succeeds at the payment-approval step, and the request often arrives from a genuine, compromised account belonging to a real vendor or colleague, so there may be nothing suspicious in the message at all. Training helps, but the control that stops the loss is procedural: verify every new or changed payment instruction by callback to a number you already hold, with no exception for urgency or seniority.

About Honeybadger Solutions

Honeybadger Solutions is an Arizona-licensed security, investigations, and cyber firm serving executives, general counsel, and organizations nationally and internationally. We are a Service-Disabled Veteran-Owned business, veteran-led, with a team that includes former military and law-enforcement professionals. Cybersecurity, digital forensics, financial investigations, and background intelligence are delivered in-house, so a phishing program, the controls behind it, and the forensic investigation if a lure succeeds all sit under one accountable chain of command.

Offices: Casa Grande (HQ), Phoenix, and Oro Valley, Arizona.
Phone: Book a consultation online
Confidential consultation: use the AI agent on this site to scope a phishing awareness program for your organization.

Browse by topic

Security guard services  ·  Private investigations  ·  Cybersecurity  ·  Digital forensics  ·  Financial fraud investigation  ·  Executive protection  ·  All articles