How a phishing simulation runs, and what you receive
A phishing simulation is a measurement exercise over a defined employee population: what is agreed before it runs, what the campaign does, and what the report and the debrief actually contain.
A phishing simulation is a controlled measurement exercise. A defined employee population receives messages built to be plausible, what those people do next is recorded, and the result is reported as population figures rather than as a list of names. What the client receives is that record, a reading of which controls saw the traffic and which did not, and a session walking through both with the people who have to act on them.
The unit of interest is the workforce, measured on a cadence, and the output is meant to hold up in front of an auditor a year later.
Before anything is sent
Two things are settled in writing before design begins: the authorisation, and the population.
CERT-In's Comprehensive Cyber Security Audit Policy Guidelines, CISG-2025-02, Version 1.0, 25 July 2025, addresses empanelled auditing organisations directly, and it is explicit about the first at §13.2.7(ii), p. 50:
"Specific written permissions must be obtained from the auditee organization before conducting tests that involve survivability failures, denial-of-service (DoS), process testing, or social engineering."
It is equally explicit about the second at §15.2.2(iii), pp. 55–56:
"Social engineering and process testing must only target group of employees explicitly included within the agreed audit scope. These tests must not involve external entities such as customers, business partners, vendors, or other third parties, unless specific written consent is obtained from the target organization."
In practice that turns the opening conversation into a document. The population is named as cohorts, by department, role, location and seniority, with a count for each. Exclusions are recorded, and so is the reason for them. Escalation contacts on the client side are named. Contractors, outsourced desks and partner staff are included only where a specific written consent covers them, and where it does not yet exist, that is a scoping decision taken before the calendar is set.
§6 of the same guidelines lists engagement types "including, but not limited to", and this work is scoped under (x) Process Security Testing and (xvi) Red Team Assessment. §15.2.2(iii)'s own phrase, "Social engineering and process testing", is the link.
Campaign design and pretext development
Pretexts are written for the organisation, not chosen from a library. That means reading how it actually communicates: the internal systems people log into, the notices they expect, and the seasonal traffic that makes a message unremarkable, such as appraisal cycles, payroll changes and policy attestations. For spear phishing variants the same work is done per target group from public reconnaissance, which is why those campaigns run against smaller populations and produce figures not comparable with a bulk round.
Two design decisions belong to the client, settled before drafting. The first is the pretext boundary: themes that are off limits are named in advance rather than adjudicated after a complaint. The standard in the clause is that this testing "should be conducted in a controlled and ethical manner", which is easier to satisfy at the design table than at the debrief.
The second is difficulty. Every pretext is recorded against a difficulty band, because a round of obvious lures followed by a round of well-built ones moves the click rate for reasons that have nothing to do with awareness. If the programme is to be read as a trend, the band has to travel with the number.
Infrastructure and the execution window
The campaign runs on infrastructure the firm operates: the sending domains and mail path used for the exercise, the landing pages, and any credential-capture page the design calls for. Delivery is external to the client estate, so a round needs no deployment and no internal change request.
One infrastructure decision changes what the round measures. If the client's mail gateway is configured to allow the campaign through, the round measures people. If it is not, the round measures the whole chain, filter and gateway and mail client and then people, and a low click rate may be evidence of a good filter rather than an aware workforce. Both are legitimate; they answer different questions. A programme that switches between them without recording which was used has broken its own series.
The execution window is a fixed period with agreed dates. Sends are staggered so the first cohort to notice does not pre-warn the rest. A named client-side contact is reachable throughout, and a stop condition is agreed in advance in case the exercise collides with a real incident.
Where vishing is within the authorisation, it is scoped as a process test rather than a test of a person: calls placed against a defined procedure, such as a service desk reset or a payment detail change, measuring whether the procedure holds when it is pushed.
The report
Reporting is on population behaviour, by cohort.
| Measure | What it records | How it appears |
|---|---|---|
| Click rate | The proportion of the delivered population that followed the link | Percentage, per cohort and overall |
| Credential submission | The proportion that went on to submit credentials at the landing page | Percentage, per cohort |
| Time-to-click | The interval between delivery and first interaction | Distribution across the execution window |
| Department and role heatmaps | Where behaviour concentrates across the organisation | Cohort grid mapped to the reporting structure |
The shape of that report is set by the same guideline that set the authorisation. §15.2.2(iii), pp. 55–56:
"When targeting general staff (e.g., untrained or non-security personnel), such testing must utilize anonymized or statistical techniques—ensuring no individual is personally identified or penalized. The purpose is to evaluate overall awareness and the effectiveness of security processes, not to single out individuals."
So the report describes a population, and the cohort design has to mean it. A four-person team given its own row is an individual with extra steps. Cohorts are sized so a figure cannot be resolved back to a person, and the heatmap is a question about where to look next, not an answer about who was careless.
Across rounds the report becomes a series: a baseline, then each round carrying its population, its pretext difficulty band and its figures. The series is what an auditor asks for. A single round is a data point.
The debrief
The report is handed over in a session rather than emailed, because half of what the round produced is not in the numbers. The purple-team walkthrough covers the other half: what did and did not detect the campaign. Which controls saw the mail and at which hop, what was logged, what raised an alert and what did not, and what a defender could have acted on and when.
That gives two readings of one exercise. The human-layer reading is the population figures. The control-layer reading is the detection timeline, and it is often the more immediately actionable: a filter rule or a logging gap can be closed this quarter.
Who is in the room matters. Security owns the control-layer findings. Where the programme exists because a regulator requires the organisation to evaluate awareness, compliance and internal audit need the figures and the method behind them. HR and the department heads whose cohorts appear on the heatmap should hear what the numbers mean before they hear them second-hand.
What the round is evidence of
| Instrument | The obligation | Citation |
|---|---|---|
| RBI Directions, 2026, Commercial Banks | "The bank shall evaluate the awareness level of employees periodically" | ¶202, 31 July 2026. The same sentence sits at PB ¶201, SFB ¶201 and CIC ¶197 |
| RBI Directions, 2026, NBFCs | A "formal mechanism to measure and track the effectiveness of such training through periodic assessments or testing", and an "up-to-date repository of the training and awareness status of all users" | ¶36, 31 July 2026 |
| SEBI CSCRF v1.0, 20 August 2024 | "REs shall periodically assess level of employee cybersecurity awareness, for e.g., through phishing test success rate, etc." | GV.RM Guidelines item 1(e), p. 87, standards column GV.RM.S3; applicability "All REs except small-size, self-certification REs (Mandatory)" |
| CERT-In CISG-2025-02, 25 July 2025 | Specific written permission before the test; anonymised or statistical reporting after it; the target population limited to the agreed employee groups | §13.2.7(ii), p. 50 and §15.2.2(iii), pp. 55–56 |
Read the SEBI wording exactly. The obligation is to assess employee awareness periodically, and phishing test success rate is named by the regulator as an example of how. "For e.g." is the regulator's own phrase and belongs in every summary of the clause. The RBI construction reaches the same place from the other side: a standing duty to evaluate, drawn as its own paragraph and separate from the duty to train, with the method left to the entity. A simulation is a defensible way to discharge either.
For an NBFC the report is also the thing ¶36 asks to be kept: population, dates, method, coverage and result, held current.
What to hold before commissioning
Four items, all obtainable before anything is signed:
- The written authorisation §13.2.7(ii) requires, naming who signs it on the client side and what it covers.
- The population as cohorts, with counts, exclusions, and the consent position for any third-party staff.
- A sample of the report, so its cohort granularity and its anonymisation can be checked against §15.2.2(iii) before the first round, not after.
- The record format for pretext difficulty, because that is what makes round two comparable with round one.
A supplier who can produce all four before the engagement starts is describing a programme. One who can produce none of them is describing a send.
About the author
Abhed Indulkar
Lead — ShadowMap
Senior full-stack engineer and ShadowMap product lead. 5+ years building security platforms across Vue.js, Laravel, React, Python, Django, AWS, and Azure. Architects the technology that powers continuous attack surface monitoring.
Continue reading
All articles →Spear phishing versus bulk simulation: what changes in the test and in the numbers
A spear campaign and a bulk campaign measure different things over different populations. Why their rates cannot share a trend line, and how a six-person cohort is reported when a percentage would identify people.
Scoping a phishing simulation programme in India
What has to be decided before the first send: the population and its cohorts, the scope boundary, the cadence, the channels, the exclusions, the escalation contacts and the written authorisation CERT-In requires.
Measuring the executive population
RBI ¶204 addresses the Board and Senior Management separately, and CERT-In's reporting rule makes a percentage over twelve people an individual result. What an executive exercise reports instead.