Designing an anonymised recruiting study
Design an anonymised recruiting study with a clear question, data-minimisation plan, re-identification review, denominators, access controls and stop rules.

An anonymised recruiting study should answer one bounded question while reducing recognition risk. Define the population, collect only what the question needs, control access, test re-identification and set stop conditions. This research-design and arithmetic guide is not legal, privacy, employment or statistical advice; anonymisation is not a product setting or guarantee. Obtain privacy, security and statistical-owner approval for the relevant context.
What question should an anonymised recruiting study answer?
Start with a question answerable without identifying a person. For example: Among the fictional team's completed technical screens in one quarter, what share reached a hiring-manager review under two versions of the same screening rubric? It defines an event, period, population, exposure and outcome, not a people list or prediction that the rubric improves hiring.
Write the protocol before extracting records:
Study question and decision owner:
Population and inclusion/exclusion rule:
Unit of analysis and one qualifying event:
Observation period and data freeze date:
Exposure or comparison definition:
Outcome definition and denominator:
Minimum fields and permitted derived fields:
Anonymisation or pseudonymisation method and residual-risk review:
Permitted viewers, access expiry and disclosure reviewer:
Analysis plan, missing-data rule and version:
Stop conditions, incident route and deletion/review date:The study permits bounded counts, shares or differences with stated limits. It does not support causal, quality, fairness, speed, conversion, performance or hiring-success claims. Label each table observed, calculated, assumed, unknown or not measured.
What is the difference between anonymisation and pseudonymisation?
Use the terms precisely. Anonymisation aims to irreversibly transform information so a person is not identifiable in the release context. Pseudonymisation replaces direct identifiers with a code, but a separate key can restore the link. A pseudonymised file remains sensitive to the team holding that key; calling it anonymised does not change the boundary.
The Article 29 Working Party's Opinion 05/2014 on Anonymisation Techniques, adopted 10 April 2014, identifies singling out, linkability and inference as risks. NIST's IR 8053, published October 2015, notes that de-identified data can be re-identified. NIST SP 800-188, published 14 September 2023, treats de-identification as risk reduction matched to goal and governance.
Record viewers, auxiliary data, repeated-query access, key custody and unusual combinations that could single someone out. Test the released table and access route, not only the raw file after names are removed.
How to minimise recruiting data before anonymising it
The ICO's data-minimisation guidance, checked 5 September 2026, says data should be adequate, relevant and limited to what is necessary. Apply it before pseudonymising or aggregating:
- Keep a study ID rather than name, email, profile URL or CV text when a row-level link is genuinely needed.
- Drop free text, exact timestamps, rare titles, tiny locations and unusual details unless required; coarsen dates, geography, seniority or categories when precision does not change the analysis.
- Exclude protected or sensitive traits unless an approved design requires a strictly necessary, controlled field; never infer them from names, photos, schools or location.
- Separate key, source records and analysis extract; record who can join them and when access expires.
Do not collect a larger candidate history "just in case." A smaller table can still be risky when remaining combinations are rare. Minimise collection, then assess residual risk in the actual output.
How to design the study sample and denominator
Describe how records entered the study: frame, dates, census or sample, selection, strata/weights, duplicates and missing-data rule. A convenience sample answers only a narrow question, not all candidates or roles. Do not replace an unavailable frame with easy-to-find records and call it coverage.
Define one unit and denominator before counting. Candidate, application, completed screen and requisition differ. Document whether two screens by one candidate count once or as two events. Do not change the unit after results; preserve approved weights and never invent one to make groups comparable.
Keep a selection record:
| Sampling field | Record | Limitation to state |
|---|---|---|
| Frame | System, queue or list from which eligible records could be selected | People or events absent from the frame were not observed |
| Selection | Census, random, systematic, stratified, quota or convenience rule | The rule affects coverage and comparability |
| Unit and deduplication | Candidate, application, screen or event; duplicate and repeat-event treatment | Results change if one person can contribute multiple rows |
| Nonresponse and missingness | Declined, incomplete, unavailable or suppressed fields and the analysis rule | Missing records can change the observed composition |
| Weighting | Approved weight, calibration or no weight, with owner and version | An unapproved weight is an assumption, not a correction |
Use descriptive arithmetic only inside the frozen population:
stage share = qualifying events at the stage / eligible events entering that stage
absolute difference = comparison share - reference share
relative difference = (comparison share - reference share) / reference shareShow n beside every percentage. Withhold or aggregate when a cell is too small, a total permits deduction or repeated tables enable differencing. Do not invent a universal minimum cell size; sensitivity, auxiliary information, release channel and policy matter.
The ONS policy for social survey microdata, last updated 5 May 2017 and checked 5 September 2026, warns that sparse combinations, especially a count of one, can support identification. Its statistical disclosure control policy, checked the same date, covers output review and record handling. Review counts, margins, linked tables and release channel; this is not a recruiting threshold.
Fictional worked example: two screening rubrics
The rubric example below is fictional. It is not a benchmark, fairness result or Talent Summoner output. A fictional company freezes 120 screens: 60 under rubric A and 60 under B. Same-period hiring-manager review is 18 for A (18 / 60 = 30%) and 15 for B (15 / 60 = 25%), a 5 percentage point difference.
The analyst does not call this causal improvement; role mix, timing and reviewer assignment were not measured. The release table contains group counts and the approved outcome. A restricted pseudonymous-ID file supports an authorised check but is not distributed. Aggregate, suppress or hold any group that could expose a rare candidate or permit deduction.
Who can access the study data, and under which controls?
NIST SP 800-53 Rev. 5, published 23 September 2020, includes least-privilege, access-enforcement, audit and separation-of-duty controls. Translate them into named study roles:
| Role | Permitted action | Boundary |
|---|---|---|
| Study owner | Approves question, protocol, population and release | Cannot silently change the outcome or denominator after review |
| Data custodian | Creates the minimum extract and protects the key, if any | Does not publish the join key or raw candidate material |
| Analyst | Runs the versioned calculation on the approved extract | Cannot add fields or query new slices without approval |
| Disclosure reviewer | Tests small cells, linkability, inference and differencing | Can require aggregation, suppression, restricted access or hold |
| Recruiting lead | Interprets the bounded result for the stated decision | Cannot treat it as individual evidence or an automated hiring rule |
Use named accounts, least privilege, strong authentication, expiry, logs and revocation. Keep the protocol, extract version, transformations, missing-data decisions, output and review decision together; record corrections. A privacy incident, identity match, unapproved field, denominator change or unexplained result is a stop.
When should an anonymised study stop or go on HOLD?
Move the study to HOLD when:
- the question, population, event or denominator is unsettled, or a sample is self-selected, duplicated, incomplete or changed without a rule;
- a pseudonymous key is called anonymisation without residual-risk assessment, or unnecessary identifiers/free text/rare combinations remain;
- counts or differences are presented as causation, fairness, quality, performance or hiring probability;
- small cells, margins, linked outputs or repeated queries could reveal a person or sensitive attribute; or
- a viewer, expiry, deletion/review date or incident owner is missing, or the recruiting lead wants individual candidate evidence.
The release decision can be restricted analysis, aggregate release, revise protocol, hold for review or close and delete/anonymise under the approved policy. If it cannot support the question safely, do not release it.
Where Talent Summoner fits
Talent Summoner is our product for candidate sourcing and candidate ranking, not study governance. Its candidate-ranking workflow, checked 28 September 2026, compares supplied CVs with a job description and produces a report for human review. It does not design a threat model, draw a sample, validate a denominator, run disclosure review or establish causal or fairness evidence.
Do not upload a study extract merely to calculate a statistic. Keep research records under the organisation's approved data, access and retention process. Candidate ranking can organise supplied CVs; it is not an ATS, research database or automated hiring decision-maker.
Is removing names enough to anonymise a recruiting study?
No. Rare combinations, dates, locations, free text and external information can make someone identifiable. Distinguish anonymisation from pseudonymisation, state the release context and test singling out, linkability, inference and differencing.
What denominator should a recruiting study use?
A recruiting study should use the eligible population entering the defined stage, or the one qualifying event rule named in the protocol. Put the count beside each percentage, keep candidates, applications and screens distinct, and withhold the result when the denominator is missing, unstable or too disclosive.
Can a study compare two recruiting rubrics and prove one is better?
Not from descriptive counts alone. A share or difference describes the frozen sample. Role mix, timing, reviewer assignment, selection and missing data may limit interpretation; do not make causal, fairness, performance or hiring-success claims without an approved design and analysis.
Should a pseudonymised file be shared with the hiring panel?
Usually not: the panel needs an aggregate result, if any, not row-level data. Keep the key and source records restricted, apply least privilege, review re-identification risk and document the disclosure decision. Follow the organisation's approved policy and qualified advice.
Can Talent Summoner design or validate this study?
No. Talent Summoner can organise supplied CVs against a role description for human review. It does not design the protocol, anonymise a dataset, validate a sample or denominator, assess disclosure risk or make a hiring decision.
Next step: hand the study protocol to its owners
Freeze the question, event, denominator, minimum fields, threat model, role owners and stop conditions. Have the privacy, security and statistical owners review the protocol before any record-level extract is made. For a separate review of CVs already in hand, read the candidate-ranking methodology and keep its output outside the study's aggregate evidence.


