diff --git a/docs/data-tests/dbt/dbt-package.mdx b/docs/data-tests/dbt/dbt-package.mdx index dd868994b..009f32fd7 100644 --- a/docs/data-tests/dbt/dbt-package.mdx +++ b/docs/data-tests/dbt/dbt-package.mdx @@ -56,3 +56,85 @@ the following var: vars: elementary_full_refresh: true ``` + +## BigQuery settings + +The following settings only affect BigQuery. Other warehouses ignore them. + +### Test table expiration + +Elementary stores test results in temporary tables during a dbt invocation, and the `on-run-end` hooks read them when +the invocation ends. On BigQuery, these tables expire after 1 hour. If your dbt invocation runs longer than that, the +tables can expire before the hooks read them, and the hooks fail. + +To keep the tables longer, set their expiration in hours: + +```yaml +vars: + temp_table_expiration_hours: 6 +``` + +This applies to all temporary tables Elementary creates on BigQuery, not only test tables. + +### Nanosecond timestamps + +dbt Fusion and dbt Core 1.11 record run timings (such as `execute_started_at`) with nanosecond precision, which +BigQuery timestamps don't support. Starting with version 0.26.1, Elementary truncates these values to microseconds +when it uploads them. Rows uploaded by earlier versions may still contain nanosecond values, so by default Elementary +also truncates timestamps whenever it casts them on BigQuery. + +The truncation converts each timestamp to a string and back, so BigQuery can't prune partitions on it. As a result, +anomaly tests on partitioned tables scan the whole table, and tables with `require_partition_filter` reject the query. + +To use plain timestamp casts again: + +1. Fix the nanosecond values that earlier versions already uploaded: + + ```shell + dbt run-operation elementary.fix_nanosecond_timing_values + ``` + + This updates the timing columns of `dbt_run_results` and `dbt_source_freshness_results` in place, truncating them + to microseconds. It scans both tables, and only needs to run once. + +2. Disable the truncation: + + ```yaml + vars: + bigquery_truncate_nanosecond_timestamps: false + ``` + + + Run `fix_nanosecond_timing_values` before you set `bigquery_truncate_nanosecond_timestamps` to `false`. Otherwise, + queries that read rows with nanosecond values fail. + + +### Partitioning of test result tables + +Starting with version 0.26.1, `elementary_test_results` and `test_result_rows` are partitioned by day on +`created_at`, like `dbt_run_results` and `dbt_invocations`. This reduces the data scanned when Elementary reads +recent test results. + +dbt only applies partitioning when it creates a table, so tables created by earlier versions keep their current +layout and keep working as before. To partition an existing table and keep its rows, rebuild it while no dbt +invocation is running: + +```sql +create table `..elementary_test_results__partitioned` +partition by timestamp_trunc(created_at, day) +as select * from `..elementary_test_results`; +``` + +Then check that the row counts match, drop the original table, and rename the new table to `elementary_test_results`. +Repeat for `test_result_rows`. + + + Don't use a full refresh to partition these tables. They hold your test history, and a full refresh deletes it. + + +To create Elementary tables without partitioning, set: + +```yaml +vars: + bigquery_disable_partitioning: true +```