This product is currently in Tech Preview and is not available for General Availability (GA). It should not be used in production environments, as features and functionality may change before the final GA release.
Python Iceberg Protector Architecture on Snowflake
The architecture of the Python Iceberg Protector using Snowflake is depicted in the following diagram:

User writes code: A developer/data engineer authors application logic like Python in a Notebook that runs inside Snowflake.
Notebook runs inside a Custom Runtime on Snowflake SPCS: The notebook is hosted in a Custom Runtime Environment, which is also referred to as CRE. The CRE is deployed to Snowflake Snowpark Container Services, which is abbreviated as SPCS. SPCS provides the compute sandbox for the whole stack.
Build pipeline delivers the runtime image: A separate Build Pipeline uses Docker and Snow CLI to build the CRE image and pushes it to an Image Registry. The image is then deployed as CRE into Snowflake SPCS, which is how the Notebook, PTYPyIceberg/PTYPyArrow, AP-C, and DevOps Policy/RPAgent components get installed together.
Notebook reads/writes tables via PTYPyIceberg and PTYPyArrow: When the notebook issues table reads or writes, it calls into the PTYPyIceberg and PTYPyArrow layer, which is the Iceberg/Arrow data-access library used inside the runtime.
PTYPyIceberg and PTYPyArrow encrypts/decrypts columns via AP-C: Sensitive columns are passed to Application Protector – C before the data leaves or after it arrives. Application Protector – C performs the actual field-level encryption on write and decryption on read.
AP-C is driven by DevOps Policy or RPAgent: AP-C uses the DevOps Policy / RPAgent component for its security policy and key material. The policy defines which fields to protect, with which method or key, and for which users. This ensures that protection is consistent and centrally governed.
Data is stored as Parquet Iceberg tables in AWS S3: After encryption, PTYPyIceberg and PTYPyArrow reads/writes Parquet files that make up the Iceberg Tables in AWS S3. Therefore, the data at-rest in S3 is already column-level protected.
Snowflake REST Catalog manages the Iceberg metadata: The Snowflake REST Catalog is also known as Polaris. The runtime talks to Polaris over a REST API authenticated with a Personal Access Token, which is abbreviated as PAT. The API call resolves Iceberg table metadata, such as namespaces, table locations, and snapshots.
Polaris vends S3 credentials for data access: Polaris then vends short-lived S3 credentials to the runtime, which PTYPyIceberg and PTYPyArrow uses to actually read/write the Parquet files in the S3 Iceberg tables. Therefore, S3 access is brokered by the catalog rather than using long-lived static keys.
- ←Previous: Downloading the DevOps Policy
- Next: System Requirements for the Python Iceberg Protector on Snowflake→
Feedback
Was this page helpful?