Skip to content

Commit c019901

Browse files
committed
docs: add entity relationships detail to docs
1 parent b63c2ac commit c019901

3 files changed

Lines changed: 39 additions & 0 deletions

File tree

Lines changed: 5 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,5 @@
1+
::: dve.core_engine.configuration.v1.hierarchy
2+
handler: python
3+
options:
4+
show_root_heading: true
5+
heading_level: 2
Lines changed: 32 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,32 @@
1+
---
2+
title: Entity Relationships
3+
tags:
4+
- Linkage
5+
- Relationships
6+
- Missing
7+
- Parent
8+
- Group
9+
- Rejections
10+
---
11+
12+
Sometimes a user may choose to use the file transformation stage to `normalise` a heavily nested dataset into separate entities during the initial reading of data. This would be done by specifying different entities in the dataset section of the contract configuration in the `dischema` file. This allows for easier interaction when customising errors in the data contract or writing transformations in the business rules. However, if the dataset being processed requires more complex validation, for example removing orphaned records or implementing group rejections, then how to link normalised entities needs to be provided. This can be provided in the `entity_relationships` section of the `dischema`
13+
14+
## Entity Relationships Content
15+
16+
To allow the DVE to link between normalised assets, the following information should be provided (per linkable entity):
17+
18+
- parent_entity: the immediate parent of the entity
19+
- join_fields: how to join the entity with its parent in dictionary form (parent_field_name: child_field_name)
20+
- mandatory: whether the child entity is a mandatory field in the immediate parent
21+
22+
There is also the functionality to customise errors related to either missing parent or group rejections:
23+
24+
- missing_parent_id_error_code: the error code to display if a record is rejected as it hs no valid parent record
25+
- missing_parent_id_error_message: the error message to display if a record is rejected as it hs no valid parent record
26+
- no_valid_records_error_code: the error code to display if parent records are removed due to no valid children in a mandatory field
27+
- no_valid_records_error_message: the error message to display if parent records are removed due to no valid children in a mandatory field
28+
29+
## Entity Hierarchy Object
30+
31+
The details provided in the entity_relationships section of the dischema are used to create an EntityHierarchy object.
32+
Please refer to [Advanced User Guidance: Entity Hierarchy](../advanced_guidance/package_documentation/entity_hierarchy.md).

‎docs/user_guidance/getting_started.md‎

Lines changed: 2 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -68,6 +68,8 @@ Within the example above, there are two parent keys - `schemas` and `datasets`.
6868
!!! note
6969
The "splitting" of entities is considerably more useful in situtations where you want to normalise/de-normalise your data. If you're unfamiliar with this concept, you can read more about it [here](https://en.wikipedia.org/wiki/Database_normalization). However, you should keep in mind potential performance impacts of doing this. If you have rules that requires fields from different entities, you will have to perform a `join` between the split entities to be able to perform the rule.
7070

71+
To support with the application of more complex validation relating to parent and child records within normalised data, the [entity_relationships](entity_relationships.md) section of the `dischema` enables users to specify parent-child relationships and to customise error codes related to missing parent and group level validation issues.
72+
7173
For each dataset definition, you will need to provide a `reader_config` which describes how to load the data during the [File Transformation](file_transformation.md) stage. So, in the example above, we expect `movies` to come in as a `JSON` file. However, you can add more readers if you have the same data in different data formats (e.g. `csv`, `xml`, `json`). Regardless of what file format, the [File Transformation](file_transformation.md) stage will convert the submitted data into a "stringified" parquet format which is a requirement for the subsequent stages.
7274

7375
To learn more about how you can construct your Data Contract please read [here](data_contract.md).

0 commit comments

Comments
 (0)