flowchart
in_properties[/properties/]
in_path[/path/]
function("write_properties()")
out[("Path: ./datapackage.json")]
in_properties --> function
in_path --> function
function --> out
Functions and classes
We created this document mainly as a way to help us as a team all understand and agree on what we’re making and what needs to be worked on. Which means that the descriptions and explanations of these functions will likely change quite a bit and may even be deleted later when they are no longer needed.
Based on the naming scheme and the Frictionless Data Package standard, these are the external-facing functions in Sprout. See the Outputs section for an overview and explanation of the different outputs provided by Sprout.
Nearly all functions have a path argument. Depending on what the function does, the path object will represent a different location. It’s designed this way to make it more flexible to where individual packages and resources are stored and to make it a bit easier to write tests for the functions. For similar reasons, most of the functions output either a dict Python object, a custom BaseProperties dataclass, or a path object.
Several of the functions have an argument called properties. The properties argument is a custom BaseProperties dataclass (a collection of key-value pairs, much like a JSON-style dict object) that describes the package and its resource(s). This metadata is stored in the datapackage.json file and follows the Frictionless Data specification.
write_*() functions always overwrite their target file and always create the parent folders of the file if they don’t exist.
Functions shown with a icon are not yet implemented while those with a icon are implemented.
For some reason, the diagrams below don’t display well on some browsers like Firefox. To see them, try using a different browser like Chrome or Edge.
Data package functions
write_properties(properties, path)
See the help documentation with help(write_properties) for more details.
Data resource functions
read_staging(resource_properties, paths)
See the help documentation with help(read_staging) for more details.
flowchart
in_path[/paths/]
in_properties[/resource_properties/]
function("read_staging()")
out[("List[DataFrame]")]
in_path --> function
in_properties --> function
function --> out
join_staging(data_list, resource_properties)
See the help documentation with help(join_staging) for more details.
flowchart
in_data[/data_list/]
in_properties[/resource_properties/]
function("join_staging()")
out[("DataFrame")]
in_data --> function
in_properties --> function
function --> out
write_resource_data(data, resource_properties)
See the help documentation with help(write_resource_data) for more details.
flowchart
in_data[/data/]
in_properties[/resource_properties/]
function("write_resource_data()")
out[("./resources/{name}/data.parquet")]
in_data --> function
in_properties --> function
function --> out
extract_field_properties(data)
See the help documentation with help(extract_field_properties) for more details.
flowchart
in_data_path[/data/]
function("extract_field_properties()")
out[("list[FieldProperties]")]
in_data_path --> function
function --> out
Properties dataclasses
These dataclasses contain an explicit, structured set of official properties defined within a data package. The main purpose of these is to allow us to pass structured properties objects between functions. They also enable users to create valid properties objects more easily and get an overview of optional and required class fields.
SproutProperties
See the help documentation with help(SproutProperties()) for more details on the properties.
Properties functions
read_properties(path)
See the help documentation with help(read_properties) for more details.
flowchart
in_path[/path/]
function("read_properties()")
out[("SproutProperties")]
in_path --> function
function --> out
Reads the datapackage.json file, checks that it is correct, and then outputs a SproutProperties object.
Data functions
check_data(data, resource_properties)
See the help documentation with help(check_data) for more details. This function checks the data against the properties in the datapackage.json file. It checks column names in the data against field names (field.name) in the properties, column data types against field types (field.type), and the data itself against any constraints on column values (field.constraints). See the function flow page for more details on the internal flow of this function.
flowchart
in_data[/data/]
in_properties[/resource_properties/]
function("check_data()")
out[("DataFrame<br>or Error")]
in_data --> function
in_properties --> function
function --> out
Base functions
write_file(text, path)
See the help documentation with help(write_file) for more details.