Sharing and publishing data with personal information
Research data containing information that can be linked directly, or indirectly by some kind of complementary means (e.g. a code key), to individual persons containing personal data.
The Data Protection Regulation (GDPR) states that personal data must be protected with organisational and technical measures to avoid unauthorised access to the material. This type of information is therefore normally not allowed to be published or shared with unauthorised persons.
Only if certain conditions are met may personal data be disclosed or shared with others. It requires a legal basis, such as the use of the material in research or that the data do not fall under confidentiality under the provisions of The Public Access to Information and Secrecy Act (2009:400). In the case of sensitive personal data, an approved ethical review is normally required by the party requesting the documents.
Please note that third country transfers of personal data (outside the EU/EEA area) may only take place if the recipient country ensures an adequate protection level.
If the risk of re-identifying individuals is deemed sufficiently low, the data set can be considered anonymised. According to the data protection regulation, it is no longer considered personal data, and the information can normally be shared and published. But formal anonymisation exists under the Data Protection Regulation only if there is no longer any code key or other reasonable possibility to attribute information in the data to individuals.
Irreversible anonymisation of all personal data in a research project should be avoided, for various reasons:
- In many studies there is a need to follow up participants in a longer perspective (e.g. in longitudinal epidemiological studies).
- In cases where published results need to be reviewed, all material in a study, including complete raw data, should be available.
- Code keys or other documentation linking data to a person may belong to the type of documents that according to the Archives Act cannot be deleted until a certain time has elapsed.
As indicated in recital 26 below, the Data Protection Regulation imposes high requirements on whether data based on personal data can be considered to be anonymised. Even if data does not containt direct identifiers, one must take into account the risk that combining information in an otherwise anonymised material may lead to the re-identification of individuals. By combining values of different variables, e.g. occupation, diagnosis, municipality and age, there is always a risk that individual persons can be identified even if direct identifiers are missing.
To reduce the risk of identifying individuals, you can limit the number of variables and generalise values, for example by specifying a region instead of a city or age group instead of an exact year. Data then gets a certain degree of so-called k-anonymisation, which means that the same value is shared by at least k persons and where k is an integer. Each combination of properties therefore occurs several times in a data set and fits into k-number of persons. K-anonymisation can then be supplemented with measures such as l-diversity and t-nearness to further limit the risk of re-identification of individuals. The Disciplinary Domain of Medicine and Pharmacy states that data should have at least 10 individuals per combination of variables when all variables are combined with each other to be considered anonymous (K-value=10).
See: Guideline on the distinction between personal data and anonymised data. Disciplinary Domain of Medicine and Pharmacy (MEDFARM 2023/3967). Uppsala University
There are various tools that facilitate anonymisation of data in data, for example ARX, sdcMicro and Amnesia. However, some types of data are difficult or impossible to anonymise, such as films, biometric data and genetic code.
The Statistics Department can provide guidance on anonymisation of personal data. Counselling is free of charge for researchers and doctoral students in the Faculty of Social Sciences. Other researchers within the university can get help at a cost and in terms of time.
Given that there is always a risk of re-identification of individuals, caution should be exercised when sharing and publishing anonymised data. It is important to make an assessment of the risk of re-identification in each individual case, as well as to justify and document on what grounds the chosen method and level of anonymisation can be considered sufficient.
More about anonymisation and publishing personal data:
- Research data management – Anonymisation, UK Data Service
- Guide to basic data anonymisation techniques, Personal Data Protection Commission Singapore (2018)
- Making Qualitative Data Reusable, DANS (2024)
- Anonymisation tools and techniques, Vrije Universiteit Brussel (2020)
When are personal data considered to be anonymised under the Data Protection Regulation (GDPR)?
Recital 26, which supports the interpretation of the Data Protection Regulation, states that:
The principles of data protection should apply to any information concerning an identified or identifiable natural person. Personal data which have undergone pseudonymisation, which could be attributed to a natural person by the use of additional information should be considered to be information on an identifiable natural person.
To determine whether a natural person is identifiable, account should be taken of all the means reasonably likely to be used, such as singling out, either by the controller or by another person to identify the natural person directly or indirectly.
To ascertain whether means are reasonably likely to be used to identify the natural person, account should be taken of all objective factors, such as the costs of and the amount of time required for identification, taking into consideration the available technology at the time of the processing and technological developments.
The principles of data protection should therefore not apply to anonymous information, namely information which does not relate to an identified or identifiable natural person or to personal data rendered anonymous in such a manner that the data subject is not or no longer identifiable. This Regulation does not therefore concern the processing of such anonymous information, including for statistical or research purposes.