Skip to main content
GuideServerConfigure and deploy Tabsdata servers on your machine.TutorialsConfigure data integration workflows within a running Tabsdata server.Advanced TutorialsBuild end-to-end workflows between two specific systems.API ReferenceCLI ReferenceRelease Notes
Version: 2.0.2

Google Cloud Storage

The Google Cloud Storage connector lets Tabsdata read files from and write files to Google Cloud Storage.

The GCSSrc connector can be used by a publisher function to read data from Google Cloud Storage into a Tabsdata table.

The GCSDest connector can be used by a subscriber function to write data from a Tabsdata table into Google Cloud Storage.

Supported formats​

FormatPublishSubscribe
AUTOYesYes
CSVYesYes
JSONYesYes
AVROYesYes
PARQUETYesYes
LOGYesNo

Publisher​

The GCSSrc connector can be used by a publisher function to read data from Google Cloud Storage into a Tabsdata table. See the full publisher walkthrough

Connection​

Google Cloud Storage publishers use GCSSrcConn to define the bucket and credentials used by the publisher.

conn-gcs.yaml
kind: connectionDef
apiVersion: '1.0'
type: tabsdatak.conn.gcs:GCSSrcConn
spec:
bucket: 'acme-ingest'
credentials:
kind: gcpJsonCredentials
apiVersion: '1.0'
type: tabsdatak.conn.common.types:GCPJson
spec:
json: 'secret:{"type":"service_account",...}'
  • required
  • required
  • optional
  • optional
  • optional

Connection Config Parameters​

bucket​

The name of the GCS bucket that contains the files to publish.

credentials​

The service-account JSON credential used to access the bucket. GCPJson requires a json key holding the full service-account key document.

project​

Optional GCP project id associated with the bucket.

base_path​

Sets the base path within the bucket. All paths configured in GCSSrc are relative to this path.

base_path starts with /, cannot end with /, and cannot contain empty path segments. The default is /.

conn_cfg​

Optional connector configuration overrides.

Publisher​

Configure GCSSrc as the source of a publisher function to select the files that Tabsdata reads from Google Cloud Storage.

publish_orders.py
from tabsdatak.api import publisher
from tabsdatak.conn.gcs import GCSSrc, SrcFileFormat

@publisher(
source=GCSSrc(
paths=[
"incoming/orders.parquet",
],
),
output_tables=["orders"],
)
def publish_orders(orders):
return orders
  • required
  • optional
  • optional
  • optional
  • optional

Function Config Parameters​

paths​

Defines the files that the publisher reads from Google Cloud Storage.

Each path is relative to the base_path configured in GCSSrcConn. Source paths:

  • cannot begin or end with /
  • cannot contain empty path segments
  • can contain a * glob in the final path segment

For example:

GCSSrc(
paths=[
"orders/2026/*.parquet",
"customers/customers.csv",
]
)

Each element in paths represents one source slot and maps positionally to an argument in the publisher function.

When a path uses a glob to match multiple files, the matching files are provided through the corresponding publisher input.

format​

Sets the format used to read the source files.

Supported formats are AUTO, CSV, JSON, AVRO, PARQUET, and LOG. The default is AUTO, which determines the format from the file extension.

format_cfg​

Sets format-specific reader options. Configuration is keyed by SrcFileFormat.

initial_last_modified​

Sets the initial cutoff for incremental ingestion.

On the first run, Tabsdata imports files modified at or after the specified timestamp. The cutoff then advances so subsequent runs only import newer files.

src_cfg​

Sets additional source configuration.

The following options are supported:

OptionDescription
tabsdata.file.chunk_sizeControls the size of chunks used when reading files.
tabsdata.src_metadata.dropWhen True, excludes the @td.file.path metadata column.

By default, rows read from Google Cloud Storage include a @td.file.path metadata column containing the source file's URI.

Subscriber​

The GCSDest connector can be used by a subscriber function to write data from a Tabsdata table into Google Cloud Storage. See the full subscriber walkthrough

Connection​

Google Cloud Storage subscribers use GCSDestConn to define the bucket and credentials used by the subscriber.

conn-gcs.yaml
kind: connectionDef
apiVersion: '1.0'
type: tabsdatak.conn.gcs:GCSDestConn
spec:
bucket: 'acme-exports'
credentials:
kind: gcpJsonCredentials
apiVersion: '1.0'
type: tabsdatak.conn.common.types:GCPJson
spec:
json: 'secret:{"type":"service_account",...}'
  • required
  • required
  • optional
  • optional
  • optional

Connection Config Parameters​

bucket​

The name of the GCS bucket where the subscriber writes files.

credentials​

The service-account JSON credential used to access the bucket. GCPJson requires a json key holding the full service-account key document.

project​

Optional GCP project id associated with the bucket.

base_path​

Sets the base path within the bucket. All paths configured in GCSDest are relative to this path.

base_path starts with /, cannot end with /, and cannot contain empty path segments. The default is /.

conn_cfg​

Optional connector configuration overrides.

Subscriber​

Configure GCSDest as the destination of a subscriber function to define where Tabsdata writes files in Google Cloud Storage.

write_orders.py
from tabsdatak.api import subscriber
from tabsdatak.conn.gcs import GCSDest, DestFileFormat

@subscriber(
destination=GCSDest(
paths=[
"exports/orders.parquet",
],
),
input_tables=["orders"],
)
def write_orders(orders):
return orders
  • required
  • optional
  • optional
  • optional

Function Config Parameters​

paths​

Defines where the subscriber writes files in Google Cloud Storage.

Each path is relative to the base_path configured in GCSDestConn. Destination paths:

  • cannot begin or end with /
  • cannot contain empty path segments
  • cannot contain globs

Each element in paths represents one destination slot and maps positionally to a value returned by the subscriber function.

Destination paths can contain variables that Tabsdata substitutes when the subscriber runs:

VariableValue
${EXEC_PLAN_ID}Current execution plan ID.
${TRX_ID}Current transaction ID.
${EXEC_PLAN_TS}Timestamp when the execution plan started.

For example:

GCSDest(
paths=[
"exports/${EXEC_PLAN_ID}/orders.parquet",
]
)

Dynamic paths can be used to create a new object for each execution instead of overwriting an existing object.

format​

Sets the format used to write the destination file.

Supported formats are AUTO, CSV, JSON, AVRO, and PARQUET. The default is AUTO, which determines the format from the destination file extension.

format_cfg​

Sets format-specific writer options. Configuration is keyed by DestFileFormat.

dest_cfg​

Sets additional destination configuration.

The tabsdata.file.chunk_size option controls the size of chunks used when writing files.