Preparing the Environment
Prepare the Environment to Install the Python Iceberg Protector on Snowflake.
This product is currently in Tech Preview and is not available for General Availability (GA). It should not be used in production environments, as features and functionality may change before the final GA release.
The Protegrity Python Iceberg Protector on Snowflake delivers column-level data protection for Apache Iceberg tables. These tables are managed by the Snowflake REST Catalog (Polaris) and stored as Parquet in cloud object storage, such as AWS S3. It enables data engineers and analysts to read from and write to Iceberg tables from Python workloads while sensitive fields are transparently protected. Protection uses the same Protegrity policy that governs the rest of the enterprise data estate.
The protector is delivered as a Custom Runtime Environment (CRE) that runs inside Snowflake Snowpark Container Services (SPCS). The runtime image is built and published through a standard container pipeline like Docker and the Snowflake CLI and deployed to SPCS as a managed container. Inside the CRE, a Snowflake Notebook hosts user code that calls the Protegrity-instrumented Iceberg and Arrow libraries, PTYPyIceberg and PTYPyArrow. These libraries are drop-in replacements for the standard Python Iceberg and PyArrow APIs, so existing Iceberg workloads can adopt protection with minimal code changes.
Protection is enforced by the Application Protector for C (AP-C), which is co-located in the runtime and invoked by PTYPyIceberg and PTYPyArrow on the columns identified by policy. On write, protected column values are encrypted, tokenized, or masked before the Parquet files are persisted to S3. On read, the same operations are reversed in memory based on the caller’s entitlements. Because protection is applied in the client runtime, the Parquet objects that land in the Iceberg table are already protected at rest, independently of the storage layer’s own encryption.
AP-C obtains its policy and key material from the DevOps Policy and Remote Protection Agent (RPAgent) components that ship inside the CRE. Policy is authored and managed centrally on the Protegrity Data Security Platform (ESA) and distributed to the runtime. Data element definitions, protection methods, and role-based access rules remain consistent with the customer’s existing Protegrity deployment.
Access to Iceberg metadata and data is brokered by Snowflake. The runtime authenticates to the Snowflake REST Catalog (Polaris) using a Personal Access Token (PAT) to resolve namespaces, table locations, and snapshots. Polaris then vends short-lived, scoped credentials that PTYPyIceberg and PTYPyArrow use to read and write the underlying Parquet files in S3. This removes the need for long-lived storage credentials in the runtime.
The following sections describe how to prepare the environment, build and deploy the CRE, and configure the Snowflake and Polaris resources. They also show how to use the Python Iceberg Protector from a notebook to read and write protected Iceberg tables.
Prepare the Environment to Install the Python Iceberg Protector on Snowflake.
Understand the Python Iceberg Protector Architecture on Snowflake.
Understand the System Requirements for the Python Iceberg Protector on Snowflake.
Install the Python Iceberg Protector on Snowflake.
Was this page helpful?