This product is currently in Tech Preview and is not available for General Availability (GA). It should not be used in production environments, as features and functionality may change before the final GA release.

Python Iceberg Protector Architecture on Databricks

Understand the Architecture to install the Python Iceberg Protector on Databricks.

The architecture of the Iceberg Protector using Databricks is depicted in the following diagram:

Python Iceberg and Parquet Modular Encryption (PME) Architecture

Write Path in Databricks environment

  1. Warehouse: The data platform like Databricks initiates the data write and interacts with the Unified Catalog to register/manage table metadata.

  2. Unified Catalog: Serves as the central metadata registry. It integrates with catalog providers such as HMS, Delta, Unity, Polaris, Horizon/Open/REST, and Glue, and receives encryption instructions from the Column Encryption Config.

  3. Column Encryption Config: Supplies the policy which columns to encrypt, key references, etc. to the Unified Catalog so encryption is applied consistently at write time.

  4. Iceberg: Consumes data from the Warehouse and coordinates with the Unified Catalog to produce Iceberg-formatted table data with encryption metadata attached.

  5. Arrow (Parquet PME): The Iceberg layer hands data to the Arrow/Parquet PME engine, which performs Parquet Modular Encryption on the specified columns.

  6. Parquet files with Encrypted Columns: The PME engine outputs Parquet files where sensitive columns are encrypted at the column level rather than encrypting the whole file.

  7. Storage (S3, Ozone, BLOB, …): The encrypted Parquet files are persisted to object storage, which is the shared source of truth for readers.

Read Path for external or independent analytics

  1. Storage → Parquet files with Encrypted Columns: Any external consumer reads the same encrypted Parquet files directly from storage.

  2. Arrow (Parquet PME): An independent Arrow/Parquet PME reader decrypts the column data, driven by its own Column Encryption Config (key references and access policy).

  3. Any other Analytical Program: After PME decryption, the analytical tool (outside the Snowflake/Databricks/Trino/Cloudera boundary) can process the plaintext columns it is authorized to see.

Key Design Points

  1. Encryption is column-level, not file-level: enabled by Parquet Modular Encryption, so different consumers can decrypt different subsets of columns based on their key access.

  2. Storage is the interoperability point: both the internal warehouse stack and external analytical programs share the same encrypted Parquet files; access control is enforced by whoever holds the keys defined in the Column Encryption Config.

  3. Catalog-agnostic: the Unified Catalog abstraction lets the same encrypted-Iceberg pattern work across HMS, Delta, Unity, Polaris, Horizon/Open/REST, and Glue.

Last modified : September 15, 2026