Guidelines for processing personal data within Data Foundry

Introduction

Data Foundry (DF) platform serves as a Design Research Data Management system that allows researchers to collect, store, process, and export data. DF enables data collection from various sources (devices, e.g., IoT, wearables; and persons, e.g., survey), storage of data in a unified format in one place, and provides tools to con-nect, curate, and share data. If research data that you are going to process within the platform contains personal data, there are several measures you need to put in place to ensure data safety. Necessary measures are described in this document.

Are you a researcher (student or staff member) using the DF platform and do you work with personal data? Then please read this document carefully and adhere to the guidelines described below. This evolving document will be updated depending on future DF functionalities and TU/e policies._

Scope of the guidelines

The term 'data' in this document refers to different type of data and covers:

  • Raw data: the original data that you have collected but have not yet processed or analyzed. These can in-clude data from experiments and observations, including audio/video recordings, images, field notes, sensor data.
  • Processed data: the data that you have modified, digitized, translated, transcribed, cleaned, validated, checked and/or anonymized
  • Analyzed data: the models, graphs, tables, texts, etc. that you have created based on the raw and the pro-cessed data.

User data

When you register or log in to Data Foundry and ID Projects, your first name and last name and email address are processed to provide access to and use of the respective platform. These details are essential for authentication and account management. Additionally, both platforms allow users the possibility to add a link to a personal website. ID Projects also provides the option to include links to social media accounts and a Github page. However, adding these links is entirely voluntarily and not required for platform use.

Within ID Projects, your first and last name and email address are retained indefinitely for attribution purposes. This ensures that project contributions remain properly credited over time while maintaining the integrity of the platform’s records. Within Data Foundry, your first and last name and email address are retained indefinitely to associate your account with the TU/e single sign-on system.

Collection of personal data

This section provides explanation on what is considered as personal data and provides general measures to con-sider while working with this type of data.

What is personal data?

The General Data Protection Regulations (GDPR) defines 'personal data' as any information relat-ing to an identified or identifiable natural person ('data subject'). Personal data also includes data that indirectly reveals something about a natural person. Personal data can lead to the physical, physiological, genetic, mental, economic, cultural or social identity of a natural person. However, data of deceased persons or organizations are not considered to be personal data be-cause they are not natural persons. There are different types of personal data: normal personal data and special category personal data (Figure 1).

CATEGORIES OF PERSONAL DATA

Figure 1: CATEGORIES OF PERSONAL DATA

The usage of special category personal data in research requires extra security measurements in order to safeguard the privacy of data subjects and to comply with the GDPR.

Informed consent: how to

Processing personal data requires a legal basis. For research this typically is TU/e's public task (as mentioned in the FAQ).

Nevertheless, participants in research still need to provide (ethical) permission to participate in a research project, and they must be informed about how their personal data will be used. TU/e has designed templates that address both personal data processing and ethical permission re-quirements. We highly recommend using these templates in your research. The forms can be integrated into online surveys. Make sure you use the latest version of the templates since they are updated on a regular basis. Please note that the participants of the research should be able to understand the content of consent forms. This means that while choosing a correct template you need to select one that is compatible with participants' language comprehension (Dutch and/or English).

Typically, the forms include an information sheet which explains the overall project and purpose for personal data collection. It is important to describe the goal of the data collection and en-visaged use of the personal data, also in future research. The use of data is limited to the specif-ic processing and purpose described in the form and no use outside those areas is permitted. Do you have multiple purposes for processing data? If so, you must inform the data subject about this. The purpose may not change along the way. Permission can be given by signature, digital signature, check mark or even verbally recorded (the latter for example in online interviews that are recorded). For children under the age of 16, it is necessary that the parent/guardian gives consent (this may differ per EU country).

If you need help with completing the informed consent template or have any questions, please contact a data steward.

To process personal data, you need to have a legal basis. For research this is often consent. Please use TU/e informed consent templates that are available here.

TU/e’s template informed consent forms fulfill the following legal requirements (see also Figure 2):

  • Be easy to understand, prominent, and concise
  • Include the name of your organization and other third parties involved in the project
  • Explain which data will be collected
  • Explain why you want to collect the data (i.e., purpose of data processing)
  • Explain what you are planning to do with the data
  • Remind data subjects (i.e., participants) that they can withdraw consent at any time
  • Be as specific as possible
  • Explain how long the data will be stored

WHAT A CONSENT MUST AND MUST NOT BE

Figure 2: WHAT A CONSENT MUST AND MUST NOT BE

Personal data processing: measures

There are various ways you can process personal data within your project, and they include: collection, recording, organization, structuring, storage, adaptation or alteration, sharing, dis-semination or otherwise making available, alignment or combination, restriction, erasure or destruction.

Processing personal data in research should be done with caution and should include measures that are important to safeguard the privacy and rights of your data subjects. Consider the following:

  • Work safely. With personal data, it is particularly important to work safely. Do not use public Wi-Fi, do not work where others can easily watch your screen or hear you talk, do not leave your laptop logged in when you are away, etc.
  • Apply data minimization. Collect as little personal information as possible and necessary for your research.
  • Work with de-identified data. Personal data can either be anonymized (data cannot (under any circumstances) be traced back to a person) or pseudonymized (no immediate identification, but it remains possible to identify a person from the data with additional information). Up until the anonymization takes place, the data is still personal data (not anonymized). Consequently, in these cases data subjects should not be informed that the survey is anonymous (this erroneously happens often).

The anonymization process aims at irreversibly preventing the identification of an individual study participant. It is quite a challenging process, which involves the use of complicated and ever-changing techniques, such as randomization, generalization or masking. Which type of techniques works best depends on the situation. For the data to be considered truly (i.e., effec-tively and sufficiently) anonymized, all direct identifiers (such as name, address, telephone number, email address) and all indirect identifiers (such as age, place of birth, occupation, edu-cation, income) need to be irreversibly altered in such a way that a data subject can no longer be directly or indirectly identifiable either by the data controller alone or in collaboration with any other party. If you are not sure whether your data is completely anonymized or you have questions about the anonymization process contact the data steward of your department for help.

Pseudonymization: replacing any identifying characteristic(s) of a data subject with an artificial pseudonym, or in other words, a value which does not allow the data subject to be directly identified. This means that identification is still possible with the identification key. The identi-fication key must be stored securely and separately from the pseudonymized data. If the data subject can be identified by combining data with additional information, the data is also called pseudonymous. Privacy laws do apply to pseudonymized data. Pseudonymization typically in-volves techniques like tokenization, data masking, and encryption.

  • Ensure storage limitation. Store personal data for no longer than is necessary for your research purposes. Do not store or distribute unnecessary copies of the data.
  • Restrict access. In case of collaboration, define who has access to collected data, how you will restrict this, how you will enable access to those authorized, and where you will describe who gets access to the data. Do not share or keep the data in a form which permits identification of data subjects for any longer than is necessary for the purposes for which the personal data is processed.
  • Be prepared for data subject (i.e., study participants) requests. Respondents should be able to successfully invoke their data subject rights within four weeks if the study includes their personal data.

If you receive formal data subject requests that you cannot handle within four weeks, please consult your data steward and the privacy team via privacy@tue.nl.

Using DF functionalities/features

When you decide to use Fitbit/Google Fit and Apple Health or Telegram in your project, make sure to include information in the informed consent form that Fitbit/Google Fit/Apple Health or Telegram will be an independent controller of data provided within these platforms. Additional-ly, include a link to the privacy statements of these companies in the informed consent form.

Telegram

You can use the functionality of Telegram bot (chat bot) to send notifications to and prompt participants to perform actions. It allows you as a researcher to collect diary, images or location data from participants in the study. Using Telegram is optional (deactivated by default), depending on a research objective. It will require your action as a researcher (sign-in with an e-mail address and a digital token) to set up a connection.

Do not use Telegram for collecting sensitive and confidential information. Furthermore, do not use it to collect special category of personal data.

OpenAI

DF provides access to the OpenAI GPT API for you as a researcher (and owner of the DF project) with the follow-ing features: 1) text completions (text generation), when you submit text "prompts" and receives "completions", 2) moderation, when you submit text "prompts" and receive scoring of the text in terms of appropriateness, safe-ty, and security according to the OpenAI general usage policies, 3) chat completions, when you can use the series of messages (e.g., existing chat with a list of prior messages) to generate a new response. The use of this func-tionality (and the associated tokens) within the DF platform is controlled by the admin of DF.

Do not use OpenAI for collecting sensitive and confidential information. Furthermore, do not use it to collect personal data including special category of personal data.

Uploading media files

Within DF platform different types of media files are processed:

  • Media files within research projects: you can upload images or audio/video recordings that are part of your research project into your DF account. Make sure that these types of media are available only to you and are not made public. Furthermore, make sure you describe handling of such files in the ERB ap-plication of your study.
  • (Applicable to ID projects and ID Demoday platform) Pictures of platform users: as a student you may wish to upload your picture next to your project description.

Data selection and storage

Selecting data to be stored on DF

You can collect and select data directly within the DF platform. You can also upload datasets, which were created outside the platform, to your account on DF. It is highly recommended to pseudonymize or even anonymize these datasets before you upload them to DF. Uploaded datasets cannot be edited within the DF platform in a straight-forward way. Such editing needs to be done outside the platform and a new version of a dataset can then be up-loaded to DF. Important: remove an old version from the platform once a new version is uploaded. However, make sure that the old version is kept for further storage on TU/e approved storage.

Data sharing and making datasets public

Data that you collect within or upload to your DF account is set as “private” by default. This means that only you, as the account owner, have access to the data. Nevertheless, during the course of your project you may need or want to share collected data and make your data publicly available to other DF platform users. Before you make a dataset publicly available, you need to make sure that data has been fully anonymized (i.e., the anonymization process is irreversible, and data cannot (under any circumstances) be traced back to a person!). Do not make a dataset public if such a dataset still contains directly or non-directly identifying information about study partici-pants.

Remember that it may not always be possible to share your data publicly. This includes situations like:

  • Your research data includes personal data which cannot be anonymized. Personal data can only be shared publicly if informed consent for data sharing has been given or when data is truly anonymized.
  • Your research data are confidential due to arrangements made with, for example, a third (commercial) party sponsoring your research or because of the confidential nature of the data.
  • You intend to make a patent application and want to avoid prior disclosure. Your data is already publicly available and/or you legally do not own the data, and consequently, you do not need to and/or cannot share your data (again) publicly.

Sharing data outside the DF platform

DF gives you (as a researcher) the possibility to store html and other website-related files (e.g., CSS, JavaScript, images) of a digital/interactive prototype or customized survey, and create a public link to it, so it becomes a pub-lic website. You can then share such a link with people (“participants” or “users”) outside of the DF platform and TU/e community. As described in the above section, you need to make sure that information presented on such a public page do not contain any information that could lead to identification of an individual (i.e, your research par-ticipants). This sharing option should only be used to present the final results/outcomes of you final study project.

Data breach

A data breach involves access to or destruction, loss or alteration of personal data. Note that it is a broad defini-tion: a data breach therefore includes not only the leakage of, but also the unlawful processing of personal data in different formats and in different ways. Examples of a data breach:

  • the loss of a USB drive containing unencrypted personal data
  • the loss of an unencrypted device (such as a laptop or tablet) that makes personal data accessible to the wrong people
  • a cyber-attack in which personal data has been captured
  • an email containing personal data of students or employees that is forwarded to the wrong recipient, or to a too large group of recipients
  • incorrect handling of personal data on paper (e.g., leaving lists of personal data at a printer, or throwing documents containing personal data into a container without first shredding them).
  • The making public of a dataset which is not (yet) anonymized whilst the data subjects have not given their consent for this.

Have you discovered a data breach? Report this immediately via the Selfserviceportal or call: 040-2475678.