Inputs

In this document we describe the inputs to Sprout at a technical level. For a more general description of what domain we expect data to come from and how it may look like conceptually, see the input data document.

The objects we primarily incorporate in the interface are properties and data, as described in the naming document.

Properties

These are Python dataclasses we’ve created to represent the properties (metadata) of a data package or resource. They always use the *Properties suffix, such as SproutProperties and ResourceProperties. While SproutProperties holds general Data Package properties, we intentionally avoided a name such as DataPackageProperties to communicate that Sprout uses stricter criteria for some properties than the full flexibility that is available in the Data Package specification.

Data

These data input objects are always Polars DataFrame objects.

Sprout expects the data objects to be in a tidy format. Tidy data is a conceptual framework around how data should be structured to be optimally usable for analysis. This is similar to what is called third normal form in relational database development, excluding the emphasis on creating additional tables to reduce redundancy.

Tidy data has the following properties:

  1. Each variable forms a column.
  2. Each observation forms a row.
  3. Each cell has a single value that represents both the column and row entity.

This tidy structure makes it easier to process and use the data for later analysis. If it is already in this form, there is less need to transform it later.

We have set a requirement on using tidy data frames because it simplifies the workflow and allows users to use their own, potentially highly customised, workflows for processing their data. This also allows for a more flexible and user-friendly interface to Sprout, since it removes the need to support or consider all the wide variety of formats data can be stored or structured in. Because data processing and organizing can be so strongly dependent on the domain, we leave this to the user to handle.

Staging directory

This directory contains data for data resources that are tidied and processed from the raw data. These are bulk collections of data that are processed together. This staging/ folder is where the data is stored before it is processed and saved in the data.parquet file. This folder makes the history between raw data and the final data resource clearer. This can help with troubleshooting issues that may arise during data entry or processing or for auditing purposes. The default location is ./staging/<resource-name>/.

Staging files

These are Parquet files that store data in a structured, columnar format (see our decision post for why we use Parquet files). These files are created when timestamped raw data has been tidied and processed in preparation for being added to a data resource. These files are usually kept in a file path in the format staging/<resource-name>/<timestamp>.parquet. Files in staging/ are not linked within the datapackage.json file. Every time a user adds or updates the raw data, these staging/ files should also be updated. Before being joined into a data resource, they are checked against the properties.

Observational units (for deletion)

An observational unit is the level of detail on the entity (e.g. human, animal, event) that the data was collected on at a given point in time. Based on legal regulations on the “right to be forgotten” in many countries (most prominent is the EU GDPR), a person can request that any personally identifiable and sensitive data of theirs is deleted. In order to achieve this legal requirement, the data must contain adequate identifiers that actually delete the correct data.

Note

An example of a more granular observational unit might be a person in a research study who came to the clinic at a specific time in May 2024 to have their blood collected and filled out a survey. This person may only want the data from that May 2024 visit to be removed.

A “coarser” observational unit might be a person who participated in a research study and had data collected for many different measurements at multiple points. This person may want all their data removed, from all periods of time and all types of data.

Technically, it is near-impossible to completely remove all traces of data on a specific observational unit (e.g. a person). That’s because data can be saved within any backups done by the data owner or automatically in the server, or when stored within source management systems (e.g. like Git). The only legal requirement is that the data that is requested to “be forgotten” is no longer used for any other purposes aside from what it has already been used for. Because of this, Sprout takes as input a list of observational units to exclude from the final data package that is intended for distribution and use.