AzureBlobSrc
class AzureBlobSrc(
paths: SrcPathsSpec,
format: SrcFileFormat = SrcFileFormat.AUTO,
format_cfg: SrcFormatCfgSpec | None = None,
initial_last_modified: datetime | None = None,
src_cfg: SrcCfgSpec | None = None,
purge: bool = False,
)
Bases: Src
Categories: source
Azure Blob source for @publisher -- one output slot per path.
Every table this source publishes automatically carries a
@td.file.path metadata column: the az:// object URI the row was read
from. A path that globs several files reports the specific file
each row came from, and a worksheet read from a workbook keeps its
#<n> fragment, so the value can be pasted straight back into
paths. Setting the src_cfg key tabsdata.src_metadata.drop to
True leaves the column off.
Parameters
Source paths relative to the connection's base_path,
one output slot per entry. Each path is relative (no
leading /), must not end with /, and must not contain
empty segments (//). A path may use a * glob in its
final segment only; backslashes, NUL and ${...} tokens
are not allowed. An .xlsx path may end with #<n> to name
the worksheet -- 1-based over every sheet the workbook
declares, hidden ones included. Two paths may name different
sheets of one workbook, giving a slot each.
formatSrcFileFormatFile format -- one of AUTO, CSV, JSON,
AVRO, PARQUET, LOG, EXCEL. AUTO (the default)
infers it from the file extension. EXCEL reads one
worksheet of an .xlsx workbook, named by the path's #<n>
fragment.
format_cfgSrcFormatCfgSpec | None (dict[SrcFileFormat, dict[Literal['separator', 'quote_char', 'eol_char', 'encoding', 'null_values', 'missing_is_null', 'truncate_ragged_lines', 'comment_prefix', 'try_parse_dates', 'decimal_comma', 'has_header', 'skip_rows', 'skip_rows_after_header', 'raise_if_empty', 'ignore_errors'], Any] | dict[ExcelSrcKey, Any] | dict[EbcdicSrcKey, Any] | dict] | None)Optional per-format reader options (source-side
keys only), keyed by the SrcFileFormat. EXCEL takes
has_header (default True); with False the first row is
data and the columns are named column_1, column_2, ...
initial_last_modifieddatetime | NoneOptional cutoff for incremental ingestion. The first run imports only files modified at or after this timestamp; the cutoff then advances each run so only newer files are re-imported.
src_cfgSrcCfgSpec | None (Mapping[Literal['tabsdata.file.chunk_size', 'tabsdata.src_metadata.drop'], Any] | None)Optional connector config; supported keys are
tabsdata.file.chunk_size and
tabsdata.src_metadata.drop.
purgeboolDelete each source file that is older than the incremental
initial_last_modified cursor. Files
older than the original initial_last_modified timestamp set at
registration time are never deleted -- they
were never read. Only valid when initial_last_modified is set; the default
is False.
Methods
validatedef validate()
Cross-field validation of format, paths and format_cfg
(source side).
Called by the framework; raises AzureBlobValidateException
on failure.