Amazon S3 Tables now support all Apache Iceberg V3 data types
Amazon S3 Tables now support all Apache Iceberg V3 data types
Overview
Amazon S3 Tables now offer complete support for the Apache Iceberg V3 specification. Users can create V3 tables or upgrade existing V2 tables to take advantage of V3 features like deletion vectors, row lineage, and new data types such as variant, nanosecond timestamps, unknown, geometry, and geography. Apache Iceberg has become the open standard for managing large analytics datasets, supporting petabyte-scale tables with schema evolution, hidden partitioning, and time travel queries while keeping data in open Parquet files in data lakes on object storage like Amazon S3.
Key V3 Capabilities
Starting today, Amazon S3 Tables support all V3 data types, including variant, nanosecond timestamps, geometry, geography, and unknown, along with deletion vectors and row lineage. These improvements solve long-standing limitations of V2 tables:
- Deletion vectors replace V2's positional delete files with a compact binary format - a 50,000-row compliance delete now writes a single deletion vector file instead of thousands of small deletes, significantly reducing compaction time and delete file overhead.
- Row lineage adds
_row_idand_last_updated_sequence_numberto each record automatically, allowing downstream pipelines to query changed rows without scanning the full table.
New Data Types
The following new data types allow native storage of semi-structured, geospatial, and nanosecond-precision data:
- Nanosecond timestamp(tz) - for nanosecond-precision timestamps
- Geometry - for geospatial data
- Geography - for geospatial data
- Unknown - for columns with no known type
- Variant - stores semi-structured data in columnar format, shredding into hidden columns and collecting statistics at write time
Example Usage
A retail analytics team can leverage V3's variant type to store diverse event shapes in a single table without predefined schemas:
CREATE TABLE my_catalog.namespace.clickstream (
event_id bigint,
event_time timestamp,
user_id string,
payload variant
) USING iceberg TBLPROPERTIES ('format-version' = '3')
Insert events with varying structures:
INSERT INTO my_catalog.namespace.clickstream VALUES
(1, current_timestamp(), 'user-42', PARSE_JSON('{"action": "purchase", "amount": 99.99, "items": ["laptop_stand"]'})),
(2, current_timestamp(), 'user-17', PARSE_JSON('{"action": "page_view", "url": "/products/webcam", "duration_ms": 4200}'));
Query the variant column directly without PARSE_JSON at read time:
SELECT event_id, user_id, variant_get(payload, '$.action', 'string') AS action,
variant_get(payload, '$.amount', 'double') AS amount
FROM my_catalog.namespace.clickstream
WHERE variant_get(payload, '$.action', 'string') = 'purchase'
AND variant_get(payload, '$.amount', 'double') > 50.00;
Enable deletion vectors for write operations by configuring merge-on-read mode:
ALTER TABLE my_catalog.namespace.clickstream
SET TBLPROPERTIES (
'write.delete.mode' = 'merge-on-read',
'write.update.mode' = 'merge-on-read',
'write.merge.mode' = 'merge-on-read'
);
When a compliance delete is executed, V3 writes a small deletion vector instead of rewriting data files, and S3 Tables compaction handles these deletion vector files automatically on the next maintenance cycle.
Upgrading from V2
Upgrading from V2 is straightforward and maintains backward compatibility. Existing V2 readers continue to work on upgraded tables until full adoption of V3 features is complete. To upgrade an existing table atomically without rewriting data:
ALTER TABLE my_catalog.namespace.existing_table
SET TBLPROPERTIES ('format-version' = '3');
On the next compaction cycle, S3 Tables will remove old V2 delete files and new modifications will use deletion vectors automatically. Note that this is a one-way operation - the Apache Iceberg specification does not support downgrading from V3 to V2.
Row Lineage for Incremental Pipelines
After a table has V3 data, row lineage enables efficient incremental processing:
SELECT *,
_row_id,
_last_updated_sequence_number
FROM my_catalog.namespace.clickstream
WHERE _last_updated_sequence_number > 42;
This returns only rows modified after sequence number 42, allowing downstream jobs to checkpoint the value and process only new changes on each run rather than scanning the full table.
AWS Ecosystem Integration
AWS offers the broadest native Apache Iceberg support among major cloud providers. V3 tables can be stored and optimized in Amazon S3 Tables, written with Amazon EMR Spark, integrated via AWS Glue, and analyzed with Amazon Redshift. Both S3 Tables and AWS Glue Data Catalog support the Iceberg REST Catalog (IRC) API, ensuring interoperability across engines regardless of the catalog endpoint.
Important Limitations
Several constraints apply to the new V3 data types:
- They require an engine built on Apache Spark 4.0 or later (e.g., AWS Glue 6.0 or later, or Amazon EMR release 8.1 or later).
- Only tables using the Parquet file format (not ORC or Avro) support the new V3 data types.
- Columns of type
variant,geometry,geography, ornanosecond timestampcannot be included in a table's sort order for compaction. However, tables containing these columns still compact under sort and Z-order strategies when the sort order uses columns of other types.
Availability
Amazon S3 Tables support for all Apache Iceberg V3 data types is now available in all AWS regions where S3 Tables is supported. The feature is available at no additional charge, with standard S3 Tables pricing applying. To get started, visit the Amazon S3 Tables documentation or create a table bucket from the Amazon S3 console. For API calls, search documentation, find regional availability, and check troubleshooting through the AWS MCP Server and relevant plugins.
Comments
No comments yet. Start the discussion.