MergeCfg
class MergeCfg(on: MergeColumnsSpec, op_column: MergeColumnSpec)
Bases: DescriptorBase
Categories: destination
How if_table_exists="merge" matches incoming rows onto a target table.
The Databricks spelling of the shared MergeCfg. Columns are named as
Databricks resolves them: case-insensitively, so on=["order_id"] hits
ORDER_ID and OrderId alike. A backtick-quoted name is for a
special-char or reserved name -- unlike Snowflake, quoting does not make it
case-sensitive.
Every incoming row carries a marker in op_column saying what to do with it,
and the marker is what decides which clause claims the row:
'i'-- insert the row, if it matches no key.'u'-- update the row it matches.'d'-- delete the row it matches.
The idempotency watermark is named on the destination rather than here --
see DatabricksDest.watermark_column. It belongs to the write, not to the
way one table's rows are matched, so a DatabricksDest naming a list of
configs still stamps every one of its targets with the same column.
Examples
A CDC feed whose op column says what to do with each row:
MergeCfg(on=["order_id", "customer_id"], op_column="op")
Parameters
onMergeColumnsSpec (list[str])The key columns rows are matched by, joined with AND. Must be
present on both sides. Each is a Databricks column identifier --
bare, or backtick-quoted for a special-char name.
op_columnMergeColumnSpec (str)The column holding each row's i / u / d marker. Read
to decide each row's fate and never written, so the target does not
need the column and will not gain it.
Methods
validatedef validate()
Field validation, delegated to the shared config: on names at
least one key and op_column does not collide with one of them.
Called by the framework; raises SqlCommonValidateException.