This product is currently in Tech Preview and is not available for General Availability (GA). It should not be used in production environments, as features and functionality may change before the final GA release.

Understanding the Architecture

Understand the Architecture to install the Python Iceberg Protector.

The architecture of the Iceberg Protector using Python is depicted in the following diagram:

  1. Client Applications Layer: Two entry points access the data.

    • Python / PySpark / Databricks / Trino / Snowflake: Query engines and compute frameworks that read/write via Python Iceberg.
    • Python App / Pandas / DuckDB, etc.: Lightweight Python-based applications that access data directly through PyArrow.
  2. Python Iceberg: The table-format layer that sits between the query engines and storage. It handles Iceberg table semantics like snapshots, schema, partitions. It also communicates with the Catalog or metadata store to resolve table locations and metadata.

  3. PyArrow: The in-memory columnar data layer used by both Python Iceberg and direct Python apps. It hosts the Parquet Modular Encryption (PME) component, which manages encryption/decryption of Parquet column data in-flight.

  4. Parquet Modular Encryption (PME): Embedded inside PyArrow, it contains:

    • Int (Internal crypto): The built-in Parquet encryption path uses a KMS directly for key material.
    • External Crypto Hook: A pluggable interface that delegates cryptographic operations to an external provider instead of the internal implementation.
  5. DBPS Crypto (External Crypto / PTY Crypto): The external cryptographic service invoked via the External Crypto Hook. It performs the actual encrypt/decrypt of column blocks plus metadata and retrieves encryption keys from its own KMS.

  6. Crypto / Column Config Infra: A cross-cutting configuration channel that supplies crypto and per-column policy settings to the client apps, PyArrow/PME, and DBPS Crypto, ensuring consistent column-level protection rules across the stack.

  7. Parquet PME Encrypted Files: The physical storage output. Files are written and read as Parquet with PME-encrypted column blocks like data and metadata. This ensures the data remains protected at rest regardless of which client path reads it.

End-to-end flow: The query engines call Python Iceberg → Python Iceberg resolves metadata via the Catalog → data I/O flows through PyArrow → PME intercepts column reads/writes → for external protection, the External Crypto Hook routes column blocks to DBPS Crypto, which uses its KMS → encrypted bytes are written to Parquet PME Encrypted Files. Direct Python apps use the same PyArrow plus PME path, bypassing Python Iceberg/Catalog.

Last modified : September 15, 2026