View a markdown version of this page

Working with Apache Iceberg V3 - Amazon Simple Storage Service

Working with Apache Iceberg V3

Apache Iceberg Version 3 (V3) is the latest version of the Apache Iceberg table format specification, introducing advanced capabilities for building petabyte-scale data lakes with improved performance and reduced operational overhead. V3 addresses common performance bottlenecks encountered with Version 2 (V2), particularly around batch updates and compliance deletes.

AWS provides support for the capabilities defined in the Apache Iceberg Version 3 (V3) specification. These capabilities include deletion vectors, row lineage, and column default values. They also include the following data types:

  • variant

  • geometry

  • geography

  • unknown

  • nanosecond-precision timestamps without a time zone

  • nanosecond-precision timestamps with a time zone

S3 Tables support all of the data types that V3 introduces.

You can use Apache Iceberg V3 with Apache Spark on Amazon EMR, AWS Glue ETL, Amazon SageMaker Unified Studio Notebooks, and tables in AWS Glue Data Catalog, including Amazon S3 Tables.

Key Features in V3

Deletion Vectors

Replaces V2's positional delete files with an efficient binary format stored as Puffin files. This eliminates write amplification from random batch updates and GDPR compliance deletes, significantly reducing the overhead of maintaining fresh data. Organizations processing high-frequency updates will see immediate improvements in write performance and reduced storage costs from fewer small files.

Row-lineage

Enables precise change tracking at the row level. Your downstream systems can process changes incrementally, speeding up data pipelines and reducing compute costs for change data capture (CDC) workflows. This built-in capability eliminates the need for custom change tracking implementations.

Variant data type

With the variant data type, you can write semi-structured data like JSON directly in Iceberg tables without defining a fixed schema in advance. V3 compatible engines shred your semi-structured data into hidden columns as you write it, generating Parquet column statistics that query engines use for optimizations like file pruning. This reduces the data your analytical queries scan. S3 Tables provides ongoing table maintenance for variant columns, including compaction, so you can consolidate data from semi-structured sources into larger files that Iceberg engines can read efficiently.

Geospatial data types

The geometry and geography types store points, lines, and polygons as native columns instead of encoded strings or paired latitude and longitude values. geometry operates on a Cartesian plane. geography treats coordinates as spherical coordinates on a spheroid. Both types recognize spatial reference system identifier (SRID) 0 and SRID 4326. Workloads such as fleet tracking and asset mapping can filter on location at query time. For examples, see Using the geospatial data types.

Nanosecond-precision timestamps

timestamp(9) records event times to nanosecond precision without a time zone, and timestamptz(9) records them with a time zone. Telemetry, sensor fusion, and financial workloads can store timestamps at source precision instead of encoding them as integers and converting them on read. For examples, see Using nanosecond-precision timestamps.

Column default values

A field can carry an initial-default value, used to populate the field for records written before the field was added to the schema, and a write-default value, used when a writer omits the column. Column default values populate a newly added column for existing rows, with no backfill. For examples, see Using column default values.

Important

The V3 specification requires that all columns of unknown, variant, geometry, and geography type default to null. Non-null initial-default or write-default values are invalid for those four types.

Unknown type

The unknown type represents a column whose type has not yet been resolved. It must be optional with null defaults and is not stored in data files, which makes it useful for placeholder columns during schema evolution and when migrating from other table formats. For examples, see Using the unknown type.

Version Compatibility

V3 maintains backward compatibility with V2 tables. AWS services support both V2 and V3 tables simultaneously, allowing you to:

  • Run queries across both V2 and V3 tables

  • Upgrade existing V2 tables to V3 without data rewrites

  • Execute time travel queries that span V2 and V3 snapshots

  • Use schema evolution and hidden partitioning across table versions

Important

V3 is a one-way upgrade. Once a table is upgraded from V2 to V3, it cannot be downgraded back to V2 through standard operations.

Getting Started with V3

Prerequisites

Before working with V3 tables, ensure you have:

  • An AWS account with appropriate IAM permissions

  • Access to one or more AWS analytics services (EMR, Glue, Amazon SageMaker Unified Studio Notebooks, or S3 Tables)

  • To use the geometry and geography types in Apache Spark, set spark.sql.geospatial.enabled=true in your Spark configuration. On AWS Glue, set it through the job's --conf argument. The variant, unknown, nanosecond timestamp, and column default value features need no additional Spark configuration.

  • An S3 bucket for storing table data and metadata

  • A table bucket to get started with S3 Tables or a general purpose S3 bucket if you are building your own Iceberg infrastructure

  • AWS Glue catalog configured

Creating V3 Tables

Creating New V3 Tables

To create a new Iceberg V3 table, set the format-version table property to 3.

Using Spark SQL:

CREATE TABLE IF NOT EXISTS myns.orders_v3 ( order_id bigint, customer_id string, order_date date, total_amount decimal(10,2), status string, created_at timestamp ) USING iceberg TBLPROPERTIES ( 'format-version' = '3' )

Upgrading V2 Tables to V3

You can upgrade existing V2 tables to V3 atomically without rewriting data.

Using Spark SQL:

ALTER TABLE myns.existing_table SET TBLPROPERTIES ('format-version' = '3')
Important

V3 is a one-way upgrade. Once a table is upgraded from V2 to V3, it cannot be downgraded back to V2 through standard operations. Before you upgrade, verify that every engine that reads or writes the table supports V3. Amazon Athena can't read V3 tables. For engine support, see Troubleshooting.

What happens during upgrade:

  • A new metadata snapshot is created atomically

  • Existing Parquet data files are reused

  • Row-lineage fields are added to the table metadata

  • The next compaction will remove old V2 delete files

  • New modifications will use V3's Deletion Vector files

  • The upgrade does not perform a historical backfill of row-lineage change tracking records

Enabling Deletion Vectors

To take advantage of Deletion Vectors for updates, deletes, and merges, configure your write mode.

Using Spark SQL:

ALTER TABLE myns.orders_v3 SET TBLPROPERTIES ('format-version' = '3', 'write.delete.mode' = 'merge-on-read', 'write.update.mode' = 'merge-on-read', 'write.merge.mode' = 'merge-on-read' )

These settings ensure that update, delete, and merge operations create Deletion Vector files instead of rewriting entire data files.

Leveraging Row-lineage for Change Tracking

V3 automatically adds row-lineage metadata fields to track changes.

Using Spark SQL:

# Query with parameter value provided last_processed_sequence = 47 SELECT id, data, _row_id, _last_updated_sequence_number FROM myns.orders_v3 WHERE _last_updated_sequence_number > :last_processed_sequence

The _row_id field uniquely identifies each row, while _last_updated_sequence_number tracks when the row was last modified. Use these fields to:

  • Identify changed rows for incremental processing

  • Track data lineage for compliance

  • Optimize CDC pipelines

  • Reduce compute costs by processing only changes

Using the variant data type

Important

The variant data type is available only in specific AWS Regions. For the full list of supported Regions, see Availability.

With the variant data type, you can write semi-structured data like JSON directly in your Iceberg tables without defining a fixed schema in advance. You can write data faster while still getting efficient analytical query performance. Iceberg V3 compatible engines shred your semi-structured data into hidden columns as you write it. These hidden columns generate Parquet column statistics that query engines use for optimizations like file pruning.

Creating a table with a variant column using Spark SQL:

CREATE TABLE IF NOT EXISTS myns.events ( event_id bigint, event_timestamp timestamp, source string, event_data VARIANT ) USING iceberg TBLPROPERTIES ( 'format-version' = '3' )

Inserting semi-structured data into the variant column:

INSERT INTO myns.events VALUES ( 1, current_timestamp(), 'web-app', PARSE_JSON('{"user_id": "u-1234", "action": "page_view", "page": "/products", "duration_ms": 350}') ); INSERT INTO myns.events VALUES ( 2, current_timestamp(), 'mobile-app', PARSE_JSON('{"user_id": "u-5678", "action": "purchase", "items": [{"sku": "A100", "qty": 2}], "total": 49.99}') );

With S3 Tables, table maintenance for variant columns, including compaction, runs automatically. Compaction consolidates data from semi-structured sources from small files into larger files that Iceberg engines can read more efficiently, improving query performance over time.

Using the geospatial data types

Creating a table with geospatial columns using Spark SQL:

CREATE TABLE IF NOT EXISTS myns.vehicle_telemetry ( vehicle_id string, event_time timestamp, position geometry(4326), service_area geography(4326), firmware string ) USING iceberg TBLPROPERTIES ('format-version' = '3')

Inserting geospatial values:

INSERT INTO myns.vehicle_telemetry VALUES ( 'v-1024', TIMESTAMP '2026-09-18 14:22:31.123456', ST_SetSrid(ST_GeomFromWKT('POINT (1 2)'), 4326), ST_SetSrid(ST_GeogFromWKT('POLYGON ((0 0, 10 0, 10 10, 0 10, 0 0))'), 4326), 'fw-3.2.1' )

Filtering on location:

SELECT vehicle_id, event_time FROM myns.vehicle_telemetry WHERE ST_Intersects( position, ST_SetSrid(ST_GeomFromWKT('POLYGON ((0 0, 10 0, 10 10, 0 10, 0 0))'), 4326) )
Note

Existing columns that hold encoded coordinates, such as paired double values or WKT strings, are not converted when you upgrade a table to V3. To adopt the geospatial types for existing data, add a column of the new type and populate it.

Note

Geodesic spatial predicates over geography are engine-dependent. Spatial predicates in Apache Spark on Amazon EMR and AWS Glue operate on geometry.

Using nanosecond-precision timestamps

Creating a table with a nanosecond timestamp column using Spark SQL:

CREATE TABLE IF NOT EXISTS myns.trades ( trade_id bigint, executed_at timestamptz(9), symbol string, price decimal(18,8) ) USING iceberg TBLPROPERTIES ('format-version' = '3')

timestamp(9) stores a timestamp without a time zone. timestamptz(9) stores one with a time zone.

Note

Reading or writing a nanosecond column with an engine that supports only microsecond precision truncates the value. Confirm that your engine supports nanosecond precision before you rely on it.

Using column default values

Adding a column with a default value using Spark SQL:

ALTER TABLE myns.orders ADD COLUMN currency string DEFAULT 'USD'

Rows written before the currency column was added return USD rather than null, and no data files are rewritten. The value is stored once in the table's schema metadata as the field's initial-default and substituted at read time.

Note

Columns of unknown, variant, geometry, and geography type must default to null, so a non-null default value is invalid for them. For examples of those types, see Using the unknown type, Using the variant data type, and Using the geospatial data types.

Using the unknown type

Creating a table with an unknown column using Spark SQL:

CREATE TABLE IF NOT EXISTS myns.staging_events ( event_id bigint, payload string, reserved unknown ) USING iceberg TBLPROPERTIES ('format-version' = '3')

An unknown column must be optional, always reads as null, and is not written to data files. You can later evolve it to a concrete type.

Best Practices for V3

When to Use V3

Consider upgrading to or starting with V3 when:

  • You perform frequent batch updates or deletes

  • You need to meet GDPR or compliance delete requirements

  • Your workloads involve high-frequency upserts

  • You require efficient CDC workflows

  • You want to reduce storage costs from small files

  • You need better change tracking capabilities

Optimizing Write Performance

  • Enable Deletion Vectors for update-heavy workloads:

    SET TBLPROPERTIES ( 'write.delete.mode' = 'merge-on-read', 'write.update.mode' = 'merge-on-read', 'write.merge.mode' = 'merge-on-read' )
  • Configure appropriate file sizes:

    SET TBLPROPERTIES ( 'write.target-file-size-bytes' = '536870912' — 512 MB )

Optimizing Read Performance

  • Leverage row-lineage for incremental processing

  • Use time travel to access historical data without copying

  • Enable statistics collection for better query planning

Migration Strategy

When migrating from V2 to V3:

  • Test in non-production first - Validate upgrade process and performance

  • Upgrade during low-activity periods - Minimize impact on concurrent operations

  • Monitor initial performance - Track metrics after upgrade

  • Run compaction - Consolidate delete files after upgrade

  • Update documentation - Reflect V3 features in team documentation

Compatibility Considerations

  • Engine versions - Ensure all engines accessing the table support V3

  • Third-party tools - Verify V3 compatibility before upgrading

  • Backup strategy - Test snapshot-based recovery procedures

  • Monitoring - Update monitoring dashboards for V3-specific metrics

The V3 data types are supported only for tables that use the Parquet data file format. They are not supported for tables that use the ORC or Avro data file formats.

For V3 tables, the sort and Z-order compaction strategies don't support the variant, geometry, geography, and nanosecond-precision timestamp data types.

Considerations for compaction

Compaction writes shredded variant Parquet files by default. Older readers that do not support shredding might fail to read compacted files. You can disable shredding by setting the table property write.variant.shredding.enabled=false.

Troubleshooting

Common Issues

Error: "format-version 3 is not supported"
  • Check your query engine catalog for compatibility with Iceberg V3.

  • Ensure that you are using the latest AWS service versions.

  • Verify your engine version supports V3

    V3 support for Amazon AWS services is as follows:

    Service V3 Support V3 variant support
    EMR Spark Release 7.12+ Release 8.0+
    AWS Glue ETL Version 5.1+ Version 6.0+
    Amazon SageMaker Unified Studio Notebooks Yes No
    AWS Glue: Iceberg REST API, Table Maintenance Yes No
    Amazon S3 Tables: Iceberg REST API, Table Maintenance Yes Yes*
    Amazon Athena (Trino) No No
    Amazon Redshift Patch 204+ No

    *Partial Region availability

    S3 Tables support all V3 data types.

    Column default values are supported in S3 Tables. They are also supported in AWS Glue version 5.1 and later, and in Amazon Redshift.

Performance degradation after upgrade
  • Verify there are no compaction failures. See Logging and monitoring for S3 Tables for more details.

  • Check if Deletion Vectors are enabled. Ensure the following properties are set:

    SET TBLPROPERTIES ( 'write.delete.mode' = 'merge-on-read', 'write.update.mode' = 'merge-on-read', 'write.merge.mode' = 'merge-on-read' )
  • You can verify table properties with the following code:

    DESCRIBE FORMATTED myns.orders_v3
  • Review partition strategy. Over partitioning can lead to small files. Run the below query to get the average file size for your table:

    SELECT avg(file_size_in_bytes) as avg_file_size_bytes FROM myns.orders_v3.files
Incompatibility with third-party tools
  • Verify tool supports V3 specification

  • Consider maintaining V2 tables for unsupported tools

  • Contact tool vendor for V3 support timeline

Getting Help

  • AWS Support: Contact AWS Support for service-specific issues

  • Apache Iceberg Community: Iceberg Slack

  • AWS Documentation: AWS Analytics Documentation

Pricing

Availability

Apache Iceberg V3 support for deletion vectors and row lineage is available across all AWS Regions where Amazon EMR, AWS Glue Data Catalog, AWS Glue ETL, and S3 Tables operate.

Column default values and the geometry, geography, unknown, and nanosecond-precision timestamp data types are available in all AWS Regions where S3 Tables are available.

The variant data type in S3 Tables is available in the following AWS Regions: US East (N. Virginia), US East (Ohio), US West (Oregon), Asia Pacific (Mumbai), Asia Pacific (Seoul), Asia Pacific (Singapore), Asia Pacific (Sydney), Asia Pacific (Tokyo), Canada (Central), Europe (Frankfurt), Europe (Ireland), Europe (London), Europe (Paris), Europe (Stockholm), and South America (São Paulo).

Additional Resources