Marketing says "anonymised." Security says "hashed." The Personal Information Protection Law (PIPL) asks a different question: whether re-identification is reasonably possible, and who still holds the means of re-identification. In China, hashing a name or applying a pseudonym does not automatically make a health dataset anonymised — and calling it anonymised does not make it so. This article explains the legal definition of anonymisation in PIPL, the difference between anonymisation and pseudonymisation under the national security standard, and the audit that tells a data controller whether a health dataset is actually outside the scope of the personal-information rules or merely pretending to be.
Why this matters: the anonymisation claim is the load-bearing wall of every health-data project
For healthcare companies, anonymisation is the legal basis that lets a dataset escape the full weight of the personal-information framework — the consent requirements, the purpose restrictions, the cross-border mechanisms, and the individual-rights obligations. If a dataset is genuinely anonymised under PIPL's definition, the personal-information rules do not apply to it, and the data can be used for research, analytics, and secondary purposes with far fewer restrictions. If it is only pseudonymised or de-identified, the full framework still applies — and the company has been building its compliance architecture on a claim that does not hold.
The stakes are operational, not academic. A research team that processes "anonymised" health data that can in fact be re-identified has, in the regulator's eyes, been processing personal information — and likely sensitive personal information — without the separate consent, security measures, and transfer mechanisms that the law requires. The same dataset that the deck calls anonymised can become the enforcement file that cites unapproved processing of sensitive health data. The first job of the compliance review is therefore to test the anonymisation claim, not to admire it.
- Marketing “anonymised” and security “hashed” are not the legal standard. PIPL question: is re-identification reasonably possible, and who still holds the key?.
- Irreversibility in practice
- True anonymisation means re-ID is not reasonably
Governing legal and statutory framework
Personal Information Protection Law of the People's Republic of China (2021) — Article 73
PIPL Article 73 defines the key terms, and the definition of anonymisation is deliberately strict. Anonymisation means the process by which personal information is processed so that a specific natural person cannot be identified, and cannot be restored. The two limbs are cumulative: not only must the current dataset fail to identify an individual, but the processing must make re-identification impossible — including restoration. A dataset that can be re-identified by joining, by re-engineering a hash, or by a key held elsewhere in the organisation is not anonymised; it is pseudonymised at best.
Personal Information Protection Law of the People's Republic of China, Article 73: "Anonymisation" means the processing of personal information such that a specific natural person cannot be identified, and the processed information cannot be restored to identify that natural person.
GB/T 35273-2020 Personal Information Security Specification
The national standard GB/T 35273-2020 provides the operational distinction between anonymisation and pseudonymisation. Under the standard, anonymisation removes the personal-information character of the data entirely — no individual can be identified from the result, alone or in combination. Pseudonymisation, by contrast, replaces identifiers with aliases but keeps the data identifiable through the alias-key relationship; it reduces risk but does not remove the data from the scope of the personal-information framework. The standard's guidance is the practical bridge between the law's definition and what the data team actually does.
Data Security Law and health-data protection
Where the dataset is health data, the Data Security Law's classification and grading regime adds a second layer. Even a genuinely anonymised dataset may remain subject to data-security duties under the classification system, and the governance expectations for health data are higher than for ordinary data. Anonymisation removes the personal-information obligations; it does not remove the data-security obligations.
Key legal analysis and enforcement precedents
Chinese courts and regulators have examined the boundary between pseudonymisation and anonymisation in the health-data and platform-data context. The judicial and administrative treatment of hashed or coded identifiers is consistent: where the data holder retains the key, the mapping, or the means of re-identification, the dataset remains personal information, and claims of anonymisation are rejected. Decisions have recognised that a hash of a name is trivially reversible against common dictionaries, that a code held in the same organisation is a pseudonym rather than an anonymised value, and that the question is answered by re-identification risk, not by the label.
For healthcare data specifically, the enforcement posture is stricter than for ordinary consumer data. Health information is sensitive personal information under Article 28 of PIPL, and regulators apply the sensitive-data framework — separate consent, specific purpose, sufficient necessity — to datasets that have been merely pseudonymised. A company that has built a research programme on the assumption that hashed names are anonymised has, in effect, been processing sensitive personal information without the required safeguards, and the discovery of the hash key in the same organisation is the fact that converts the marketing claim into a regulatory finding.
The cross-border angle makes the distinction decisive. PIPL imposes specific mechanisms for transferring personal information abroad — CAC security assessment for important data or large volumes, standard contract clauses, or certification. A "de-identified" dataset that is still personal information cannot be sent to an overseas research partner on the strength of an anonymisation claim. Where the dataset contains genetic or related data, the human genetic resources (HGR) rules add their own approval and filing requirements. The anonymisation analysis is therefore the gate to the cross-border file as well.
The practical test for anonymisation in a health-data context is not whether the data team can re-identify with effort, but whether the data controller and its processors have designed the pipeline so that re-identification is impossible in the ordinary course. The national security standard GB/T 35273-2020 points in this direction by requiring that the risk of identification be assessed against the data's intended use and the other data available to the processor. For a research dataset shared with an overseas partner, the assessment must include the partner's ability to join the data with its own datasets; a "de-identified" file that becomes identifiable inside the partner's environment was never anonymised from the partner's perspective, and the transfer that carried it was a transfer of personal information.
Operational vulnerabilities and transactional pitfalls
- The hash-key illusion: the data team hashes names and patient IDs, but the hash key, the mapping table, or the original data remains in the same system. Re-identification is one query away, and the dataset is pseudonymised, not anonymised.
- The token-and-key split: identifiers are replaced with tokens, but the token-key mapping is retained for "future linkage." The retained mapping is the re-identification mechanism, and its existence defeats the anonymisation claim.
- The external-linkage blind spot: the dataset is de-identified in isolation, but it can be joined with publicly available or licensed data to identify individuals. Anonymisation analysis must test the dataset in combination, not alone.
- The small-sample risk: a dataset with rare-disease records, unusual treatment combinations, or small geographic cells can be re-identified even without identifiers. The anonymisation standard requires assessing re-identification risk in context, and small-cell health data fails the test.
- Uncontrolled residual access: the "anonymised" dataset is shared broadly, while the original identifiers remain accessible to the same analysts, vendors, or overseas partners. The access path is the re-identification path.
- The board-pack claim: the compliance deck states "all data anonymised" in one line, and nobody audits the actual pipeline. The claim becomes the exhibit when the pipeline contradicts it.
The Singapore lens: anonymisation claims travel badly
In my life-sciences and clinical-research practice in Singapore, I review health-data projects across the Asia-Pacific region, and the anonymisation claim is the wall that every project leans on — and the wall that fails most often when it crosses a border. The claim that works in one jurisdiction’s marketing deck does not survive another regulator’s test, and China’s standard under the PIPL and GB/T 35273-2020 is deliberately strict: anonymisation requires that re-identification be impossible, not merely difficult, and the retained hash key, mapping table or token-key relationship is the mechanism that turns the claim into pseudonymisation. The cross-border angle is where the Asia-Pacific practice bites: a dataset that is “anonymised” for an internal analytics project may still be personal information for the purposes of the cross-border transfer rules, and a Singapore-based CRO that receives a de-identified China dataset with the mapping retained has received personal information that the transfer gate still applies to. The teams that get this right audit the pipeline before they rely on the claim: what fields were removed, what was transformed, what was retained, and whether re-identification is possible with the resources the organisation already holds. The anonymisation claim is a legal conclusion, not a data-engineering description; the pipeline audit is the evidence, and the claim without the evidence is the project that the regulator unpicks.
- Define release purpose
- Who receives what fields, for what
- Inventory direct + quasi IDs
- Names, IDs, free text, rare
Strategic compliance roadmap and action plan
Audit the anonymisation claim with a five-step test before relying on it:
- Map the pipeline: document how the raw data becomes the "anonymised" dataset — what fields are removed, what is transformed, and what is retained. The map shows the difference between what was intended and what was built.
- Test reversibility: determine whether re-identification is possible using the same organisation's resources — the hash key, the mapping table, the original data, or joinable datasets. Any realistic path back to an individual means the data is not anonymised.
- Assess the combination: test the dataset against external data — public registries, published research, licensed databases — for the small-cell and quasi-identifier risks that make de-identified health data re-identifiable.
- Identify residual access: list every person, system, and vendor that can still see the original identifiers or the key. Residual access to raw identifiers is the audit's most direct finding.
- Fix the classification: where the dataset is pseudonymised, stop calling it anonymised; apply the full PIPL framework — separate consent for sensitive health data, purpose restrictions, security measures, and the correct cross-border mechanism — or invest in true anonymisation that removes the re-identification path permanently.
Document the audit result in the compliance file, and re-run the test whenever the dataset, the pipeline, or the access community changes. The anonymisation claim is only as strong as the last test that verified it.
The audit should also address the human and organisational layer, because re-identification risk is not only a technical question. An "anonymised" dataset loses its protection the moment an analyst who worked on the original raw data carries the linkage knowledge into the new project, or when a vendor retained for the de-identification work keeps a copy of the original file for "testing." The compliance review therefore includes access review, personnel movements, and vendor contracts: who can still see the raw identifiers, who changed roles, and what the service agreement says about retained copies. In the enforcement record, the decisive fact is often not the algorithm but the leftover copy — the original file in a backup, the key in a spreadsheet, the analyst who remembers the mapping.
What not to do
Do not let marketing define anonymisation. Do not retain the hash key and call the data clean. Do not join de-identified health data with other datasets and call the result anonymous. Do not send a "de-identified" dataset offshore on an anonymisation claim that the cross-border regulator will test. The law asks one question — can the person be identified — and the answer is found in the pipeline, the keys, and the access, not in the deck.
The anonymisation analysis should be documented in the compliance file and re-run whenever the dataset, the pipeline, or the access community changes. The claim is only as strong as the last test that verified it, and the test is only as good as the record that proves it was run.
Read next: HGR compliance · Clinical data · Privacy
Discussion
Share experience or questions about this topic. This is a public discussion — not legal advice. Do not post confidential case details.
Have a question after reading? Leave it here, or Ask a Lawyer for a free initial consultation.
Comments are moderated. China Legal Portal is a directory and information resource; no attorney–client relationship is formed by posting here.