Data types
In research, various types of data are handled – from detailed information about individuals to compiled statistics and metadata about the data sets. Different data types have different areas of application and handling requirements.
Microdata and individual data
Microdata is detailed information about individual objects within a data set. Examples of such units include natural persons, companies, workplaces or schools. In research and statistics, the term individual data is often used, which refers to information about individuals that is so detailed that it is possible to identify a specific person. Individual data is normally also personal data.
Individual and microdata are collected by public authorities or the health and medical care sector. They may contain information about individuals' age, gender, finances, education, medical interventions or other personal attributes. Individual data may also have been collected in research projects, for example through interviews and surveys. Data from several different sources relating to an individual can be linked via their personal identity number.
Health data
The meaning of the term health data varies depending on the context. The European Health Data Space (EHDS) Regulation defines electronic health data. In general, it covers information about health and genetics, such as data produced in healthcare (diagnoses, care measures, use of medicines), data on factors affecting health, genetic data, data from medical registers and biobanks.
One approach is to categorise data by source, meaning where health data is documented or in connection with which organisation the information is generated. This includes data generated in Swedish healthcare and social care or research studies conducted in connection with these activities. Health data is also generated outside healthcare, but from a research perspective, this data is rarely available to order. There is also a wealth of lifestyle data that is generated and stored completely independently of healthcare, but which for certain diagnoses can be of great importance to the course of the disease and is therefore recorded in medical records when a patient is treated by healthcare services.
The National Board of Health and Welfare's report Kartläggning av datamängder av nationellt intresse på hälsodataområdet describes different types of health data. These include for example:
- socioeconomic data (occupation, education, income)
- organisational data linked to the healthcare system (geography, economy, staff composition)
- environmental data (living environment, air and water quality).
Metadata
Metadata is usually described as ‘data about data’ because it provides characteristics and information about data rather than its content. Metadata can provide information about various aspects of data and is important for both facilitating work with a particular data set and for tracking and documenting data. Unlike microdata, metadata does not contain any individual data points (measurement values).
Aggregated data
Aggregated data is information that has been combined and summarised at a group level to provide an overview or summary of the detailed information. This means that individual data points (measurement values) are combined into a total sum or summary.
Aggregated data is often used to report the results of research or surveys in a way that is easy to understand and communicate to a wider audience. When conducting research with sensitive data, it is also necessary to publish results in aggregated form in tables and similar formats. For example, aggregated data can show patterns or trends within the group studied, making it possible to draw conclusions and make informed decisions without the data being linked to an individual. Aggregated data can be ordered or downloaded directly from several authorities.
Open data
Many Swedish authorities provide open data, meaning digitised, public information that anyone can access and use for any purpose without restrictions or fees. Open data is usually published according to standards or in formats that make it easily accessible and usable for a broad audience.
Public microdata
Public microdata (also known as Public Use Files) are files or datasets created for use as teaching materials, for example. Public microdata are always anonymised or pseudonymised.
Synthetic data
Synthetic data is artificial data that mimics the structure and statistical properties of the original data, but does not contain actual information about individuals or objects. This enables researchers to work with data that resembles real data sets, but without using personal data.
Examples of when synthetic data can be useful:
- Protecting privacy: Since synthetic data does not contain real personal data, it can be used to create test environments, develop algorithms and train AI models without the risk of revealing identities.
- Developing and testing: New software, analysis methods and models can first be tested on synthetic data before being applied to real data sets.
Synthetic data can therefore be a valuable complement to real data, but it is never an exact copy of the original and should therefore be used with an awareness of its limitations.