What each number in a phishing simulation report measures
A phishing simulation returns a small set of numbers, and each supports a narrow claim. What click rate, credential submission, time-to-click, report rate and repeat exposure each measure.
A phishing simulation returns a small set of numbers, each supporting a narrow claim. Most of the trouble a programme meets later comes from a number being asked to carry a claim it cannot.
The obligations these numbers discharge are specific about measurement and silent about method. Commercial banks shall "evaluate the awareness level of employees periodically" (RBI Directions, 2026, Commercial Banks, ¶202, 31 July 2026; the same sentence is ¶201 for Payments Banks and Small Finance Banks, ¶197 for Credit Information Companies). NBFCs shall "deploy a formal mechanism to measure and track the effectiveness of such training through periodic assessments or testing" (NBFC Direction, ¶36). SEBI requires that "REs shall periodically assess level of employee cybersecurity awareness, for e.g., through phishing test success rate, etc." (CSCRF v1.0, 20 August 2024, GV.RM Guidelines item 1(e), p. 87).
One metric, in one instrument, offered as an example. Everything else in the report is there because somebody chose it. Every figure below is a property of a population under one pretext on one day.
Click rate
The proportion of the delivered population that followed the link. Delivered, not sent: bounces, quarantined mail and mailboxes nobody reads belong in the reconciliation, not in the denominator. A report that does not state its denominator has not stated a click rate.
It is evidence of how much of the workforce a single pretext could move, once. Two variables sit inside it.
The first is pretext difficulty. A bespoke pretext built on real organisational detail and a generic one produce different numbers from the same people; the difference is a fact about the design, not the population. Rates from campaigns never banded to a common difficulty are a collection of numbers, not a series.
The second is the state of the mail path. If the sending infrastructure was allowlisted at the gateway, the rate measures human behaviour with the technical control deliberately removed. If it was not, the rate measures the gateway and the people together, and a movement between campaigns may be a movement in filtering. Both designs are legitimate; a series that mixes them silently is not.
Credential submission
A materially different event. Following a link is one decision, taken in a mail client in seconds. Submitting credentials is a second, taken on a page the person has now looked at. A click is exposure; a submission is compromise.
The gap between the two carries most of the information. A high click rate with few submissions describes a population that inspects the destination on arrival: the pretext survived the mail client and not the browser. The two close together describe a lure and a destination that both held. The metric counts submission events against the same delivered denominator; it is not a measure of credential validity.
Time-to-click
A distribution across the execution window, not an average. This is the metric most often reduced to a single figure it cannot support: a workforce that responds in the first minutes and one that responds evenly across three days can share a mean and nothing else. Report the shape against a stated window: when the first click landed, where the mass sits, how long the tail runs.
What it is evidence of is speed: how fast a message moves through this organisation once it arrives. It sets the size of the problem the defensive side is asked to solve, and means something only against the two figures below.
Report rate
The proportion of the delivered population who told somebody. Of everything in the set, this figure best predicts what the organisation will do during a real campaign, because it is the only one counting an action that helps rather than harms.
It is also the figure most dependent on something other than the people measured. A population with nowhere to send a suspicion can only fail to report one: where there is no mailbox, no named owner and no acknowledgement, a rate near zero is a finding about the route before it is one about the workforce.
That places the metric inside the purpose CERT-In states for such testing: to evaluate "overall awareness and the effectiveness of security processes" (CISG-2025-02, Version 1.0, 25 July 2025, §15.2.2(iii), pp. 55–56). The reporting route is a security process; the report rate measures it.
Time-to-first-report
Elapsed time from first delivery to the first report reaching somebody who could act on it. This is the number a defender actually cares about, because it bounds how early the organisation could have responded.
Read it against the time-to-click distribution. If the first report lands before the bulk of the clicks, there was a window in which the campaign could have been blocked and the workforce warned. If it lands after, the organisation would have learned about the campaign from its consequences.
It measures a chain rather than a behaviour: somebody recognised, a route existed, and a person at the other end registered it inside the window. A programme that lifts report rate without moving time-to-first-report has improved recognition and left response where it was.
Repeat exposure
Whether the same individuals respond across campaigns, or the responding set is resampled each time. Two organisations can hold an identical click rate across four quarters, one where the same people click every time and one where a different part of the workforce does. Repeat exposure is the only thing in the set that separates them.
It carries a data-protection precondition, settled before the first campaign rather than after the third. The analysis needs the identified layer of two or more campaigns to stay linkable, and the DPDP Act, 2023 requires erasure "as soon as it is reasonable to assume that the specified purpose is no longer being served" (s.8(7)(a)), with s.8(5) safeguards over that data while it is held. If repeat-exposure analysis is a purpose of the programme, it is named as one at commissioning, with the retention window it needs. If it is not named, the identified layer goes when the campaign's purpose is served and cannot be reconstructed.
The output is a cohort statistic either way. That the responding population is largely the same across four campaigns is a finding about a programme; a list of who they are is not something this measurement produces.
What each instrument asks to be measured
| Instrument | What it asks for | Where |
|---|---|---|
| RBI Directions, 2026 (31 July 2026) | "evaluate the awareness level of employees periodically" | Commercial Banks ¶202; Payments Banks ¶201; Small Finance Banks ¶201; Credit Information Companies ¶197 |
| RBI Directions, 2026 — NBFCs | A formal mechanism to measure training effectiveness through periodic assessments or testing, and an "up-to-date repository of the training and awareness status of all users" | ¶35–36 |
| SEBI CSCRF v1.0 (20 August 2024) | Periodic assessment of employee awareness, phishing test success rate named as an example. Applies to "All REs except small-size, self-certification REs (Mandatory)" | GV.RM Guidelines item 1(e), p. 87 |
| SEBI CSCRF v1.0 — Cyber Capability Index | Percentage of information system security personnel trained in the past year; target 100%, weighting 5%; MIIs and Qualified REs | Annexure-K, p. 166, Measure 3 [PR.AT.S1] |
| CERT-In CISG-2025-02 (25 July 2025) | Anonymised or statistical technique, no individual personally identified or penalised | §15.2.2(iii), pp. 55–56 |
The Cyber Capability Index measure is a coverage metric — who received training, inside what period — scoped to information system security personnel rather than the whole workforce. The rest is behaviour over a defined employee population, and a programme is asked for both.
What none of these numbers is evidence of
None of them supports a statement about an individual's competence.
CISG-2025-02 applies to CERT-In empanelled auditing organisations (§4), and on general staff it is explicit about the form the result takes:
"...must utilize anonymized or statistical techniques—ensuring no individual is personally identified or penalized. The purpose is to evaluate overall awareness and the effectiveness of security processes, not to single out individuals."
A rate is a property of a population under one pretext on one day. One person's response to one email, under a lure built to work, is a single observation, and no inference about their judgement survives being written as a number.
That has a consequence for cohort reporting. Department and role cuts show where a process rather than a person produced the result, but a cohort small enough to name somebody is a person with extra steps. The minimum cohort size is set before the campaign runs, not after somebody asks who.
Two further claims the set does not support:
- How the same population would respond to a different pretext. The rate belongs to the lure as much as to the people.
- Improvement across quarters, unless pretext difficulty banding and population drift were controlled by design. Comparability is decided when the series is designed, not when it is charted.
What to hold
Keep four things beside every rate: the denominator, the pretext difficulty band, whether the mail path was allowlisted, and the date. A number without them cannot be re-used the following year, which is when it is wanted.
That is what makes the set defensible as evidence: the NBFC Direction asks for an "up-to-date repository of the training and awareness status of all users" (¶36), and a repository of rates whose conditions were never recorded is one nobody can stand behind.
Settle the purposes at commissioning. Written authorisation precedes the testing — CISG-2025-02 §13.2.7(ii), p. 50, requires that "Specific written permissions must be obtained" from the auditee organisation before process testing or social engineering — and it is the place to fix which metrics the programme produces, which cohort cuts it reports, the minimum cohort size, and how long the identified layer lives.
About the author
Siddarth G
Practice Director — Cybersecurity
Leads Security Brigade's offensive security practice with deep expertise in vulnerability research, penetration testing, and red team operations. Ranked Top 80 globally on Bugcrowd.
Continue reading
All articles →Spear phishing versus bulk simulation: what changes in the test and in the numbers
A spear campaign and a bulk campaign measure different things over different populations. Why their rates cannot share a trend line, and how a six-person cohort is reported when a percentage would identify people.
Scoping a phishing simulation programme in India
What has to be decided before the first send: the population and its cohorts, the scope boundary, the cadence, the channels, the exclusions, the escalation contacts and the written authorisation CERT-In requires.
Measuring the executive population
RBI ¶204 addresses the Board and Senior Management separately, and CERT-In's reporting rule makes a percentage over twelve people an individual result. What an executive exercise reports instead.