BigQuery
The BigQuery connector lets Tabsdata write tables into BigQuery. BigQuery is destination-only: there is no publisher-side connector for reading from BigQuery.
The BigQueryDest connector can be used by a subscriber function to write data from a Tabsdata table into BigQuery. See the full subscriber walkthrough
Connection
BigQuery subscribers use BigQueryDestConn to define the GCS staging location and credentials used by the subscriber. The connector stages each write as parquet in GCS, then runs a load job against BigQuery referencing that object.
kind: connectionDef
apiVersion: '1.0'
type: tabsdatak.conn.bigquery:BigQueryDestConn
spec:
staging:
bucket: 'acme-stage'
credentials:
kind: gcpJsonCredentials
apiVersion: '1.0'
type: tabsdatak.conn.common.types:GCPJson
spec:
json: 'secret:{"type": "service_account", ...}'
base_path: '/bigquery'
credentials:
kind: gcpJsonCredentials
apiVersion: '1.0'
type: tabsdatak.conn.common.types:GCPJson
spec:
json: 'secret:{"type": "service_account", ...}'
- required
- required
- optional
- optional
- optional
- optional
Connection Config Parameters
staging
The GCSLocation used as the parquet staging bucket for load jobs: a bucket, credentials (GCP service-account JSON), and base_path (defaults to /, same rules as other connectors' base_path: starts with /, cannot end with /, cannot contain empty path segments).
credentials
The GCP service-account JSON credentials (GCPJson) used to run the BigQuery load job, separate from the credentials on staging used to upload to GCS.
project
The default GCP project id. Required when dataset is set.
dataset
The default BigQuery dataset id. Requires project to be set.
entity_whitelist
An optional list of regex patterns narrowing which tables the connector considers. Unset by default, which filters nothing.
conn_cfg
Optional connector configuration overrides.
Subscriber
Configure BigQueryDest as the destination of a subscriber function to define which tables Tabsdata writes in BigQuery.
from tabsdatak.api import subscriber
from tabsdatak.conn.bigquery import BigQueryDest
@subscriber(
destination=BigQueryDest(
tables=[
"orders",
],
),
input_tables=["orders"],
)
def write_orders(orders):
return orders
- required
- optional
- optional
- optional
Function Config Parameters
tables
Defines the destination tables the subscriber writes to in BigQuery.
Each element in tables represents one destination slot and maps positionally to a value returned by the subscriber function. Each name is [[project.]dataset.]table; a bare table or dataset.table is completed from the connection's project/dataset defaults. BigQuery wraps the whole name in one backtick pair, e.g. `my-project.ds.t`.
if_table_exists
Controls what happens when the destination table already exists. "append" (the default) adds rows. "replace" drops and recreates the table.
schema_evolution
Controls how new columns in the returned frame are handled. "update" (the default) allows new columns. "strict" rejects schema changes.
dest_cfg
Sets additional destination configuration. The only supported key is bigquery.logging_level.