Working with Apache Iceberg V3
Apache Iceberg Version 3 (V3) is the latest version of the Apache Iceberg table format specification, introducing advanced capabilities for building petabyte-scale data lakes with improved performance and reduced operational overhead. V3 addresses common performance bottlenecks encountered with Version 2 (V2), particularly around batch updates and compliance deletes.
AWS provides support for the capabilities defined in the Apache Iceberg Version 3 (V3) specification. These capabilities include deletion vectors, row lineage, and column default values. They also include the following data types:
variant
geometry
geography
unknown
nanosecond-precision timestamps without a time zone
nanosecond-precision timestamps with a time zone
S3 Tables support all of the data types that V3 introduces.
You can use Apache Iceberg V3 with Apache Spark on Amazon EMR, AWS Glue
ETL, Amazon
SageMaker Unified Studio Notebooks, and tables in AWS Glue Data
Catalog, including Amazon S3
Tables
Key Features in V3
- Deletion Vectors
-
Replaces V2's positional delete files with an efficient binary format stored as Puffin files. This eliminates write amplification from random batch updates and GDPR compliance deletes, significantly reducing the overhead of maintaining fresh data. Organizations processing high-frequency updates will see immediate improvements in write performance and reduced storage costs from fewer small files.
- Row-lineage
-
Enables precise change tracking at the row level. Your downstream systems can process changes incrementally, speeding up data pipelines and reducing compute costs for change data capture (CDC) workflows. This built-in capability eliminates the need for custom change tracking implementations.
- Variant data type
-
With the variant data type, you can write semi-structured data like JSON directly in Iceberg tables without defining a fixed schema in advance. V3 compatible engines shred your semi-structured data into hidden columns as you write it, generating Parquet column statistics that query engines use for optimizations like file pruning. This reduces the data your analytical queries scan. S3 Tables provides ongoing table maintenance for variant columns, including compaction, so you can consolidate data from semi-structured sources into larger files that Iceberg engines can read efficiently.
- Geospatial data types
-
The
geometryandgeographytypes store points, lines, and polygons as native columns instead of encoded strings or paired latitude and longitude values.geometryoperates on a Cartesian plane.geographytreats coordinates as spherical coordinates on a spheroid. Both types recognize spatial reference system identifier (SRID) 0 and SRID 4326. Workloads such as fleet tracking and asset mapping can filter on location at query time. For examples, see Using the geospatial data types. - Nanosecond-precision timestamps
-
timestamp(9)records event times to nanosecond precision without a time zone, andtimestamptz(9)records them with a time zone. Telemetry, sensor fusion, and financial workloads can store timestamps at source precision instead of encoding them as integers and converting them on read. For examples, see Using nanosecond-precision timestamps. - Column default values
-
A field can carry an
initial-defaultvalue, used to populate the field for records written before the field was added to the schema, and awrite-defaultvalue, used when a writer omits the column. Column default values populate a newly added column for existing rows, with no backfill. For examples, see Using column default values.Important
The V3 specification requires that all columns of
unknown,variant,geometry, andgeographytype default to null. Non-nullinitial-defaultorwrite-defaultvalues are invalid for those four types. - Unknown type
-
The
unknowntype represents a column whose type has not yet been resolved. It must be optional with null defaults and is not stored in data files, which makes it useful for placeholder columns during schema evolution and when migrating from other table formats. For examples, see Using the unknown type.
Version Compatibility
V3 maintains backward compatibility with V2 tables. AWS services support both V2 and V3 tables simultaneously, allowing you to:
-
Run queries across both V2 and V3 tables
-
Upgrade existing V2 tables to V3 without data rewrites
-
Execute time travel queries that span V2 and V3 snapshots
-
Use schema evolution and hidden partitioning across table versions
Important
V3 is a one-way upgrade. Once a table is upgraded from V2 to V3, it cannot be downgraded back to V2 through standard operations.
Getting Started with V3
Prerequisites
Before working with V3 tables, ensure you have:
-
An AWS account with appropriate IAM permissions
-
Access to one or more AWS analytics services (EMR, Glue, Amazon SageMaker Unified Studio Notebooks, or S3 Tables)
-
To use the
geometryandgeographytypes in Apache Spark, setspark.sql.geospatial.enabled=truein your Spark configuration. On AWS Glue, set it through the job's--confargument. The variant, unknown, nanosecond timestamp, and column default value features need no additional Spark configuration. -
An S3 bucket for storing table data and metadata
-
A table bucket to get started with S3 Tables or a general purpose S3 bucket if you are building your own Iceberg infrastructure
-
AWS Glue catalog configured
Creating V3 Tables
Creating New V3 Tables
To create a new Iceberg V3 table, set the format-version table property to 3.
Using Spark SQL:
CREATE TABLE IF NOT EXISTS myns.orders_v3 ( order_id bigint, customer_id string, order_date date, total_amount decimal(10,2), status string, created_at timestamp ) USING iceberg TBLPROPERTIES ( 'format-version' = '3' )
Upgrading V2 Tables to V3
You can upgrade existing V2 tables to V3 atomically without rewriting data.
Using Spark SQL:
ALTER TABLE myns.existing_table SET TBLPROPERTIES ('format-version' = '3')
Important
V3 is a one-way upgrade. Once a table is upgraded from V2 to V3, it cannot be downgraded back to V2 through standard operations. Before you upgrade, verify that every engine that reads or writes the table supports V3. Amazon Athena can't read V3 tables. For engine support, see Troubleshooting.
What happens during upgrade:
-
A new metadata snapshot is created atomically
-
Existing Parquet data files are reused
-
Row-lineage fields are added to the table metadata
-
The next compaction will remove old V2 delete files
-
New modifications will use V3's Deletion Vector files
-
The upgrade does not perform a historical backfill of row-lineage change tracking records
Enabling Deletion Vectors
To take advantage of Deletion Vectors for updates, deletes, and merges, configure your write mode.
Using Spark SQL:
ALTER TABLE myns.orders_v3 SET TBLPROPERTIES ('format-version' = '3', 'write.delete.mode' = 'merge-on-read', 'write.update.mode' = 'merge-on-read', 'write.merge.mode' = 'merge-on-read' )
These settings ensure that update, delete, and merge operations create Deletion Vector files instead of rewriting entire data files.
Leveraging Row-lineage for Change Tracking
V3 automatically adds row-lineage metadata fields to track changes.
Using Spark SQL:
# Query with parameter value provided last_processed_sequence = 47 SELECT id, data, _row_id, _last_updated_sequence_number FROM myns.orders_v3 WHERE _last_updated_sequence_number > :last_processed_sequence
The _row_id field uniquely identifies each row, while _last_updated_sequence_number tracks when the row was last modified. Use these fields to:
-
Identify changed rows for incremental processing
-
Track data lineage for compliance
-
Optimize CDC pipelines
-
Reduce compute costs by processing only changes
Using the variant data type
Important
The variant data type is available only in specific AWS Regions. For the full list of supported Regions, see Availability.
With the variant data type, you can write semi-structured data like JSON directly in your Iceberg tables without defining a fixed schema in advance. You can write data faster while still getting efficient analytical query performance. Iceberg V3 compatible engines shred your semi-structured data into hidden columns as you write it. These hidden columns generate Parquet column statistics that query engines use for optimizations like file pruning.
Creating a table with a variant column using Spark SQL:
CREATE TABLE IF NOT EXISTS myns.events ( event_id bigint, event_timestamp timestamp, source string, event_data VARIANT ) USING iceberg TBLPROPERTIES ( 'format-version' = '3' )
Inserting semi-structured data into the variant column:
INSERT INTO myns.events VALUES ( 1, current_timestamp(), 'web-app', PARSE_JSON('{"user_id": "u-1234", "action": "page_view", "page": "/products", "duration_ms": 350}') ); INSERT INTO myns.events VALUES ( 2, current_timestamp(), 'mobile-app', PARSE_JSON('{"user_id": "u-5678", "action": "purchase", "items": [{"sku": "A100", "qty": 2}], "total": 49.99}') );
With S3 Tables, table maintenance for variant columns, including compaction, runs automatically. Compaction consolidates data from semi-structured sources from small files into larger files that Iceberg engines can read more efficiently, improving query performance over time.
Using the geospatial data types
Creating a table with geospatial columns using Spark SQL:
CREATE TABLE IF NOT EXISTS myns.vehicle_telemetry ( vehicle_id string, event_time timestamp, position geometry(4326), service_area geography(4326), firmware string ) USING iceberg TBLPROPERTIES ('format-version' = '3')
Inserting geospatial values:
INSERT INTO myns.vehicle_telemetry VALUES ( 'v-1024', TIMESTAMP '2026-09-18 14:22:31.123456', ST_SetSrid(ST_GeomFromWKT('POINT (1 2)'), 4326), ST_SetSrid(ST_GeogFromWKT('POLYGON ((0 0, 10 0, 10 10, 0 10, 0 0))'), 4326), 'fw-3.2.1' )
Filtering on location:
SELECT vehicle_id, event_time FROM myns.vehicle_telemetry WHERE ST_Intersects( position, ST_SetSrid(ST_GeomFromWKT('POLYGON ((0 0, 10 0, 10 10, 0 10, 0 0))'), 4326) )
Note
Existing columns that hold encoded coordinates, such as paired
double values or WKT strings, are not converted when you upgrade a
table to V3. To adopt the geospatial types for existing data, add a column of
the new type and populate it.
Note
Geodesic spatial predicates over geography are engine-dependent.
Spatial predicates in Apache Spark on Amazon EMR and AWS Glue operate on
geometry.
Using nanosecond-precision timestamps
Creating a table with a nanosecond timestamp column using Spark SQL:
CREATE TABLE IF NOT EXISTS myns.trades ( trade_id bigint, executed_at timestamptz(9), symbol string, price decimal(18,8) ) USING iceberg TBLPROPERTIES ('format-version' = '3')
timestamp(9) stores a timestamp without a time zone.
timestamptz(9) stores one with a time zone.
Note
Reading or writing a nanosecond column with an engine that supports only microsecond precision truncates the value. Confirm that your engine supports nanosecond precision before you rely on it.
Using column default values
Adding a column with a default value using Spark SQL:
ALTER TABLE myns.orders ADD COLUMN currency string DEFAULT 'USD'
Rows written before the currency column was added return
USD rather than null, and no data files are rewritten. The value is
stored once in the table's schema metadata as the field's
initial-default and substituted at read time.
Note
Columns of unknown, variant,
geometry, and geography type must default to null, so
a non-null default value is invalid for them. For examples of those types, see
Using the unknown type, Using the variant data type,
and Using the geospatial data types.
Using the unknown type
Creating a table with an unknown column using Spark SQL:
CREATE TABLE IF NOT EXISTS myns.staging_events ( event_id bigint, payload string, reserved unknown ) USING iceberg TBLPROPERTIES ('format-version' = '3')
An unknown column must be optional, always reads as null, and is not
written to data files. You can later evolve it to a concrete type.
Best Practices for V3
When to Use V3
Consider upgrading to or starting with V3 when:
-
You perform frequent batch updates or deletes
-
You need to meet GDPR or compliance delete requirements
-
Your workloads involve high-frequency upserts
-
You require efficient CDC workflows
-
You want to reduce storage costs from small files
-
You need better change tracking capabilities
Optimizing Write Performance
-
Enable Deletion Vectors for update-heavy workloads:
SET TBLPROPERTIES ( 'write.delete.mode' = 'merge-on-read', 'write.update.mode' = 'merge-on-read', 'write.merge.mode' = 'merge-on-read' ) -
Configure appropriate file sizes:
SET TBLPROPERTIES ( 'write.target-file-size-bytes' = '536870912' — 512 MB )
Optimizing Read Performance
-
Leverage row-lineage for incremental processing
-
Use time travel to access historical data without copying
-
Enable statistics collection for better query planning
Migration Strategy
When migrating from V2 to V3:
-
Test in non-production first - Validate upgrade process and performance
-
Upgrade during low-activity periods - Minimize impact on concurrent operations
-
Monitor initial performance - Track metrics after upgrade
-
Run compaction - Consolidate delete files after upgrade
-
Update documentation - Reflect V3 features in team documentation
Compatibility Considerations
-
Engine versions - Ensure all engines accessing the table support V3
-
Third-party tools - Verify V3 compatibility before upgrading
-
Backup strategy - Test snapshot-based recovery procedures
-
Monitoring - Update monitoring dashboards for V3-specific metrics
The V3 data types are supported only for tables that use the Parquet data file format. They are not supported for tables that use the ORC or Avro data file formats.
For V3 tables, the sort and Z-order compaction strategies don't support the variant, geometry, geography, and nanosecond-precision timestamp data types.
Considerations for compaction
Compaction writes shredded variant Parquet files by default. Older readers that
do not support shredding might fail to read compacted files. You can disable
shredding by setting the table property
write.variant.shredding.enabled=false.
Troubleshooting
Common Issues
- Error: "format-version 3 is not supported"
-
-
Check your query engine catalog for compatibility with Iceberg V3.
-
Ensure that you are using the latest AWS service versions.
-
Verify your engine version supports V3
V3 support for Amazon AWS services is as follows:
Service V3 Support V3 variant support EMR Spark Release 7.12+ Release 8.0+ AWS Glue ETL Version 5.1+ Version 6.0+ Amazon SageMaker Unified Studio Notebooks Yes No AWS Glue: Iceberg REST API, Table Maintenance Yes No Amazon S3 Tables: Iceberg REST API, Table Maintenance Yes Yes* Amazon Athena (Trino) No No Amazon Redshift Patch 204+ No *Partial Region availability
S3 Tables support all V3 data types.
Column default values are supported in S3 Tables. They are also supported in AWS Glue version 5.1 and later, and in Amazon Redshift.
-
- Performance degradation after upgrade
-
-
Verify there are no compaction failures. See Logging and monitoring for S3 Tables for more details.
-
Check if Deletion Vectors are enabled. Ensure the following properties are set:
SET TBLPROPERTIES ( 'write.delete.mode' = 'merge-on-read', 'write.update.mode' = 'merge-on-read', 'write.merge.mode' = 'merge-on-read' ) -
You can verify table properties with the following code:
DESCRIBE FORMATTED myns.orders_v3 -
Review partition strategy. Over partitioning can lead to small files. Run the below query to get the average file size for your table:
SELECT avg(file_size_in_bytes) as avg_file_size_bytes FROM myns.orders_v3.files
-
- Incompatibility with third-party tools
-
-
Verify tool supports V3 specification
-
Consider maintaining V2 tables for unsupported tools
-
Contact tool vendor for V3 support timeline
-
Getting Help
-
AWS Support: Contact AWS Support for service-specific issues
-
Apache Iceberg Community: Iceberg Slack
-
AWS Documentation: AWS Analytics Documentation
Pricing
-
Amazon EMR: Compute and storage pricing
-
AWS Glue: Job run and Data Catalog pricing
-
S3 Tables: Storage and request pricing
Availability
Apache Iceberg V3 support for deletion vectors and row lineage is available across all AWS Regions where Amazon EMR, AWS Glue Data Catalog, AWS Glue ETL, and S3 Tables operate.
Column default values and the geometry, geography, unknown, and nanosecond-precision timestamp data types are available in all AWS Regions where S3 Tables are available.
The variant data type in S3 Tables is available in the following AWS Regions: US East (N. Virginia), US East (Ohio), US West (Oregon), Asia Pacific (Mumbai), Asia Pacific (Seoul), Asia Pacific (Singapore), Asia Pacific (Sydney), Asia Pacific (Tokyo), Canada (Central), Europe (Frankfurt), Europe (Ireland), Europe (London), Europe (Paris), Europe (Stockholm), and South America (São Paulo).