Apache Iceberg 1.12.0 arrived on September 30, 2026. It is a broad release, with more than 800 merged changes listed in the project changelog. The practical story is easier to follow than that number suggests. Iceberg 1.12 makes several version 3 features usable in more engines, tightens the REST catalog contract, improves streaming reliability, and lands a substantial amount of format version 4 groundwork.
It also removes old integration surfaces. Spark 3.4 support is gone. Flink 2.0 support is gone. The Java implementation no longer writes position delete files that embed deleted row values. A group of APIs scheduled for removal in 1.12 has been deleted. This is an upgrade to test, not a dependency version to change casually.
This guide explains what changed, which changes matter in production, and how to validate an upgrade across Spark, Flink, Kafka Connect, REST catalogs, and shared Iceberg tables. It follows our earlier Iceberg 1.12 preview, but uses the final release tag and merged source as the authority.
In one minute: Iceberg 1.12.0 expands real engine support for Variant, geospatial types, row lineage, deletion vectors, encrypted metadata, and REST catalog operations. It adds Hilbert-curve clustering and a long list of correctness fixes. It also removes Spark 3.4, Flink 2.0, deprecated APIs, and the ability to write position deletes with row data. Format v4 work is present, but v4 is not a production format you should enable.
Apache Iceberg 1.12.0 at a Glance
Area
What changed
What operators should do
Compatibility
Spark 3.4 and Flink 2.0 support removed. Flink 2.2 and 2.3 added. Spark 4.2 included in source releases.
Inventory runtime artifacts and rebuild every integration in a test environment.
Delete files
Readers keep compatibility with old position deletes that carry row data, but writers and two maintenance actions reject them.
Find affected v2 tables before the upgrade and choose compaction or a v3 migration path.
Iceberg v3
Stronger support for Variant, geometry, geography, row lineage, deletion vectors, and encryption.
Test feature by feature. Opening a v3 table is not proof that an engine supports every v3 feature.
REST catalog
New unregister-table and function endpoints, catalog-side purge delegation, labels, Variant types, expression alignment, and path encoding fixes.
Test client and server capability negotiation, authorization, names with special characters, and purge ownership.
Maintenance
Hilbert clustering, better partition-statistics scans, safer expiration and orphan cleanup, and nonblocking REST metrics reporting.
Benchmark maintenance plans on representative tables and monitor rewritten bytes, files, and commits.
Format v4
Relative paths, Parquet or Avro manifests, content statistics, tracked-file adapters, and new readers are taking shape.
Treat this as implementation groundwork. Do not plan a production v4 migration yet.
The Upgrade Boundary: Support Removals and API Breaks
The fastest way to understand Iceberg 1.12 is to start with what no longer works. These changes were announced through deprecations, but a deprecation becomes an outage when an application still depends on it.
Spark 3.4 support is removed
The project removed its Spark 3.4 module. Iceberg 1.12 updates the supported Spark lines to 3.5.9, 4.0.4, and 4.1.3, and it includes Spark 4.2 work in the source release. If a job still loads an iceberg-spark-runtime-3.4 artifact, there is no 1.12 equivalent to substitute.
Do not solve this by mixing an older Iceberg runtime jar into a newer shared catalog without testing. Engines that share a table can run different library versions, but they must agree on every feature those tables contain. A Spark 3.4 writer on Iceberg 1.11 may keep working against a table used by a 1.12 reader, but it will not understand every behavior or feature the newer stack may introduce.
Flink support moves forward
Iceberg 1.12 adds Flink 2.2 and 2.3 while retaining Flink 1.20 and 2.1. Flink 2.0 is removed. That leaves four supported lines: 1.20, 2.1, 2.2, and 2.3. The important detail is not just compilation. Stateful streaming jobs need savepoint, checkpoint, serializer, and commit behavior tested on the exact Flink and Iceberg pair you will deploy.
Position deletes with row data become read-only legacy
Iceberg v2 allowed a position delete record to include the deleted row in addition to the file path and row position. The feature is usually called position deletes with row data, or PDWR. It was deprecated in Iceberg 1.11 and removed from Java writers in 1.12.
Existing files remain readable. The position_deletes metadata table retains its row column, so old records can still be inspected. New writers will not create those files. Two maintenance operations fail loudly when they encounter PDWR: rewrite_position_delete_files and rewrite_table_path. That choice prevents a maintenance job from silently changing semantics.
Before upgrading, scan the position_deletes metadata table for non-null row values. If they exist, either compact those deletes into data files while the table stays on v2, or validate a v3 path that replaces them with deletion vectors. The right choice depends on every engine that reads the table. Iceberg format upgrades are not a place to assume that “v3 supported” means feature parity. Our Iceberg v3 engine support matrix separates those capabilities.
Deprecated APIs are actually gone
The release removes deprecated methods and fields across Core, Data, Spark, Flink, Kafka Connect, BigQuery, and AWS integrations. Notable removals include SparkFilters, old SparkTableUtil methods, GenericAppenderFactory, BaseFileWriterFactory, DataReader, deprecated namespace encoding helpers, an old HadoopFileIO constructor, and older S3 signer classes and properties.
If you own custom catalog, file I/O, writer, or Spark extension code, compile it against 1.12 before scheduling the deployment. Runtime smoke tests alone will miss source incompatibilities in code paths that are loaded only by maintenance or failure handling.
Iceberg v3 Features Move from Specification to Engine Plumbing
Iceberg 1.12 does not flip a switch that makes all v3 features universal. What it does is close important gaps in the Java implementation and engine integrations. Four areas stand out.
Variant becomes practical in more pipelines
Variant is Iceberg's type for semi-structured values. Instead of forcing changing JSON objects into a fixed set of columns, Variant stores a typed binary value and can “shred” selected fields into Parquet columns for efficient filtering.
Iceberg 1.12 adds vectorized Parquet reads for Variant in Spark, Variant reads and writes through Avro in Flink, and Parquet shredding for Flink, Kafka Connect, and generic record writers. The release also fixes high-precision decimal shredding, UTF-8 ordering for string bounds, metrics behavior when a value column has no statistics, malformed binary parsing, default-writer conversion, and ORC filter pushdown on tables that contain Variant columns.
Those details matter because semi-structured data is rarely uniform. A happy-path test with small strings and integers does not exercise large decimals, null-heavy objects, encrypted columns, long strings, or schema changes. Add those cases to interoperability tests before a Variant writer joins a shared table.
Geometry and geography get a complete storage path
The v3 specification introduced geometry and geography. Iceberg 1.12 maps both to Parquet logical types, reads and writes Well-Known Binary values in Parquet, supports them in Avro, adds binary serialization in the API, calculates geometry bounding-box metrics, and preserves coordinate reference system information in writer metrics. Spark 4.1 gains Parquet read, write, and DML coverage with a fallback when vectorized reads are unavailable.
This is the first Java release where native geospatial columns are worth a serious end-to-end test. It is still not permission to upgrade a shared table without checking every consumer. Read the deep guide to Iceberg geometry and geography for CRS rules, serialization details, and adoption gates.
Row lineage and deletion-vector correctness improve
Row lineage assigns stable row identifiers and records the sequence at which a row was last updated. The release fixes last-updated-sequence inheritance, carries row-lineage and encryption key identifiers through snapshot value methods, fixes first-row-ID carryover during Spark manifest rewrites, and closes a direct-memory leak in Arrow vectorized readers.
Deletion-vector handling also becomes safer. Iceberg now rejects scans when a position delete or deletion vector claims a partition that does not match its data file. It exposes colocated deletion vectors through DataFile, preserves deletion-vector encryption metadata in merges, and strengthens v4 deletion-vector structures. These are correctness changes, not only optimizations.
Encryption reaches another operational milestone
Iceberg 1.12 commits manifest-list encryption keys with the snapshot that uses them. Spark manifest rewrites also encrypt newly written manifests, and merge paths preserve deletion-vector encryption metadata. A new delegate-style EncryptingFileIO makes encrypted I/O easier to compose with an existing implementation.
Encryption tests should include more than successful reads. Rotate keys, deny an old key, retry failed commits, rewrite manifests, expire snapshots, and prove that readers report a useful error when key access is missing. Our Iceberg table encryption guide covers the envelope and KMS boundaries in detail.
REST Catalog Changes Put More Policy on the Server
The Iceberg REST protocol keeps growing from a table metadata API into a broader catalog control plane. That direction is visible throughout 1.12.
Unregister table: The OpenAPI specification adds an endpoint that removes a table registration without deleting the table's files.
Functions: List and load endpoints bring SQL function metadata into the REST contract.
Catalog-side purge: Spark adds rest.catalog-purge, allowing DROP TABLE PURGE to delegate deletion to the REST service.
Catalog object identifiers: A shared identifier schema supports objects beyond tables and views.
Labels: Catalog metadata can carry labels for enrichment and classification.
Variant and expressions: The REST model adds Variant and aligns stored expressions with the newer expression specification.
Correct path encoding: REST path segments now use RFC 3986 percent-encoding, which matters for namespaces and object names that contain reserved characters.
Remote signing: The OpenAPI definition formalizes configuration for remote request signing.
Catalog-side purge is the most consequential operational change in that list. File deletion requires storage credentials, table context, and a clear audit boundary. Delegating purge lets the catalog enforce that policy rather than handing every Spark client the ability to delete table files. It also changes the failure domain. Operators need to know whether a failed request left the table registered, unregistered, partially deleted, or queued for asynchronous cleanup.
The unregister endpoint deserves the same care. Unregister is not drop. It removes catalog metadata while intentionally leaving files behind. A recovery runbook should record the last metadata location before unregistering anything, then verify that the table can be registered again.
If your catalog serves several engines, capability negotiation is essential. A client should not infer support from the server brand or version string. It should discover the endpoint or capability, send a valid request, and handle the server's documented error model. For protocol fundamentals, see what the Iceberg REST catalog is and how clients use it.
Spark: Better Clustering, Migration, and Maintenance
Spark receives features that will show up in daily table operations, not just compatibility plumbing.
Hilbert-curve clustering joins bin packing and Z-order
The rewrite_data_files action adds a Hilbert-curve clustering strategy. Like Z-order, a Hilbert curve maps several dimensions into a one-dimensional ordering so nearby values tend to land together. Hilbert curves often preserve locality well, but that does not make them a universal replacement for sorting or Z-order.
Benchmark with your actual filter combinations. Measure bytes scanned and files opened, but also measure rewrite cost, file-size distribution, and how quickly new writes erode the clustering. A better first query is not a win if a table must be rewritten too frequently.
Migration procedures can tolerate missing files
The Spark migrate and snapshot procedures add ignore_missing_files. This is useful when a legacy table's metastore references files that no longer exist. It is also dangerous if treated as a default. Ignoring a missing file converts an obvious failure into an intentional data omission. Capture every skipped path, reconcile it against the source, and require a data owner to accept the result.
Other Spark fixes include correct time-travel filtering after column renames, first-row-ID handling during manifest rewrites, encrypted manifest rewrites, a fix for orphan-file removal at root table locations, and a fix for Z-order on nullable booleans. These are good candidates for regression tests if your tables use the affected paths.
Flink: Streaming Correctness Gets the Attention It Needs
Streaming bugs are expensive because the failure may surface long after the write. Iceberg 1.12 addresses several of the failure modes that matter most in long-running Flink jobs.
Duplicate-commit prevention: The sink fixes a case where a restarted job with a changed job ID could commit the same work again.
Equality-delete conversion: A new pipeline can plan, resolve, write, and commit conversions from equality deletes to deletion vectors.
Identifier-aware routing: Dynamic sinks honor schema identifier fields when routing records.
Resource control: Dynamic sinks can set a slot-sharing group for finer resource management.
Variant: Flink can write shredded Variant and use Variant through Avro readers and writers.
Schema behavior: Flink SQL handles table comments and supports column positioning during alterations.
Reader stability: The release fixes wakeup-related thread and memory leaks and file-offset mismatches when files are skipped.
The equality-delete conversion work is particularly useful for CDC pipelines. Equality deletes are convenient to produce because the writer needs a key, not a known file and row position. Readers pay the cost of matching those keys. Converting them to deletion vectors can move work out of repeated reads and into controlled maintenance. Validate watermark behavior, checkpoint recovery, branch selection, and the cleanup of converted delete files before enabling it on a high-volume job.
Kafka Connect Stops Hiding Important Failures
Kafka Connect's most important change is simple: commit failures are surfaced instead of silently swallowed. The connector adds a metric for partial commit failures and bounded retries for transient commit exceptions. It also fixes rebalance scenarios that could commit files from a prior attempt and tracks control-topic offsets as a high-water mark.
Schema and conversion fixes cover null-valued records after a schema change, UUID conversion from Avro, decimal inference, MongoDB timestamp and date arrays, and Variant shredding. Removed deprecated members in TableReference and IcebergWriterResult may affect custom transforms or extensions at compile time.
After upgrading, force three failure drills: make a catalog commit fail, trigger a task rebalance during a pending commit, and restart from a known control-topic offset. Confirm that the connector reports the failure, avoids duplicate table data, and resumes from the expected point.
Scan Planning, Metrics, and Maintenance Improvements
The less visible changes in 1.12 may save more time than the headline features. Partition-statistics scans can now project columns and apply filters, reducing the amount of metadata an operator or optimizer must read. Snapshot expiration correctly reads delete manifests in cases that previously went wrong. A scan-based action removes dangling delete files. Adaptive split-sizing options are documented, making it easier to balance planning overhead against parallelism.
RESTMetricsReporter.report() no longer blocks the calling thread. Metrics are valuable only if emitting them cannot stall the work they describe. Still monitor queue size, delivery failures, and dropped reports. Moving work off the caller does not make a slow metrics endpoint harmless.
Reader and file-format fixes cover Arrow integer-to-long promotion, dictionary-encoded INT96 timestamps with offsets, large decimals, decimal default values, end-of-file handling in object-store streams, Parquet row-group size tracking, and faster decimal handling that avoids unnecessary BigInteger conversions. These changes deserve targeted tests when the affected types appear in critical tables.
For a durable operating model, combine these improvements with table-level signals. Our guide to monitoring Iceberg table health shows how metrics reports and metadata tables work together.
What the Format v4 Work Does and Does Not Mean
Iceberg 1.12 contains a lot of code labeled v4. It adds relative paths to the evolving specification, utilities for turning absolute locations into relative ones, Parquet or Avro manifests, optional table metadata locations, tracked-file builders and adapters, a v4 manifest reader, content-statistics structures and filters, and first-row-ID assignment. The project also added a draft Mumbling Bitmap specification and a read-only implementation.
This is implementation groundwork. It does not mean format version 4 is ready for production tables, and it does not mean that setting format-version=4 is a supported migration. The current work is valuable because it shows where the format is heading.
Relative paths can make table metadata less tied to one bucket or storage prefix.
Parquet manifests can bring richer types and reader tooling to manifest data.
Content statistics can carry structured index and average-size information alongside tracked files.
Modified entries and tracked files can express changes that do not fit the older added, existing, and deleted entry model cleanly.
Bitmap structures may provide compact ways to represent sets, but the Mumbling work is explicitly draft material.
The useful action today is to keep table movement and metadata assumptions out of application code. Do not parse v4 structures yourself, and do not build a roadmap around draft field names. Follow the specification and implementation through later releases.
Security and Supply-Chain Work
The project published an Iceberg security model that describes trust boundaries across catalogs, clients, storage, credentials, and metadata. CI now makes vulnerability scanning block relevant pull requests, and the release aligns and updates Jackson versions to address reported vulnerabilities. OAuth token refresh keeps optional parameters during non-exchange refreshes, and AWS REST signing can use assumed-role credentials.
A library upgrade does not secure an Iceberg deployment by itself. Use the security model to review who can alter catalog metadata, who can request temporary storage credentials, who can delete files, where tokens are cached, and how audit events connect a catalog action to an object-store action.
A Production Upgrade Plan for Iceberg 1.12.0
Upgrade by table behavior, not by dependency file. The following sequence catches the failures most likely to escape a basic smoke test.
1. Build an engine and feature inventory
List every Spark, Flink, Kafka Connect, Python, Rust, query-engine, and maintenance client that touches each catalog.
Record its runtime version, Iceberg library version, catalog client, authentication method, and file I/O implementation.
Record table format versions and the actual features present: equality deletes, position deletes, deletion vectors, Variant, geo types, encryption, branches, and views.
Identify custom code that implements Iceberg APIs or loads removed classes.
2. Find explicit blockers
Any Spark 3.4 job that expects a 1.12 runtime artifact.
Any Flink 2.0 job that expects continued support.
Any PDWR file that will reach a rewrite maintenance action.
Any custom code that imports removed APIs.
Any REST proxy or catalog that decodes paths differently from RFC 3986.
Any purge workflow that assumes clients, rather than the catalog, own file deletion.
3. Create a representative compatibility table
Use production-like object storage and a nonproduction catalog. Populate a table with nulls, renamed and reordered columns, several partition specs, a branch, equality and position deletes, decimal extremes, timestamps, and enough files to exercise planning. If you plan to adopt v3 features, create a separate v3 table for Variant, geo, row-lineage, and deletion-vector tests. Do not mix a format upgrade into the library upgrade unless the test is explicitly about both.
4. Test failure paths
Successful reads and writes are the smallest part of an upgrade test. Interrupt a commit after files are written, deny a storage credential, expire an OAuth token, restart a Flink job with a changed job ID, rebalance Kafka Connect, make a metrics endpoint slow, and retry an ambiguous catalog response. Then inspect snapshots, manifests, orphan files, control-topic offsets, and audit logs.
5. Benchmark maintenance separately
Run manifest rewrites, data-file rewrites, snapshot expiration, orphan cleanup, and dangling-delete removal against copies of representative tables. Compare file counts, bytes rewritten, wall time, planning time, commits, and object-store requests. Test Hilbert clustering only where multidimensional filters justify its ongoing rewrite cost.
6. Roll out by writer risk
Upgrade stateless readers first, then low-volume writers, then streaming writers, and finally maintenance jobs that can touch many files. Keep one known-good reader available during the rollout. Stop if snapshot counts, commit failures, orphan growth, delete-file ratios, or scan-planning latency move outside the established baseline.
Should You Upgrade Now?
Upgrade soon if you need Flink 2.2 or 2.3, Spark 4.2 source support, geospatial columns in Spark 4.1, broader Variant paths, Hilbert clustering, stronger encryption handling, REST catalog improvements, or one of the correctness fixes in your active workload.
Wait if a critical job remains on Spark 3.4 or Flink 2.0, if custom integrations have not compiled against the removed APIs, or if shared tables contain features that another engine cannot read. Waiting should be an explicit compatibility decision with a target date, not an indefinite freeze.
Most teams should separate three decisions: upgrading the Java library, upgrading an engine runtime, and upgrading a table's format version. Doing one at a time keeps the rollback surface understandable.
No. Iceberg 1.12 improves v3 implementation support, but a library upgrade does not automatically convert existing tables to format version 3. Treat a format upgrade as a separate, one-way compatibility decision.
Does Iceberg 1.12 support format v4?
The release contains significant v4 specification and implementation groundwork. That is not the same as a production-ready format v4 release. Do not create or migrate production tables to v4 based only on the presence of these classes and readers.
Can Iceberg 1.12 read old position deletes with row data?
Yes. Existing files remain readable. Iceberg 1.12 stops writing them, and the rewrite_position_delete_files and rewrite_table_path maintenance paths reject them.
Which Spark versions should be used with Iceberg 1.12?
The release updates supported Spark lines to 3.5.9, 4.0.4, and 4.1.3, and includes Spark 4.2 in source releases. Spark 3.4 support is removed. Verify the exact runtime artifact available for your Spark and Scala combination before deploying.
Which Flink versions are supported?
The supported Flink lines are 1.20, 2.1, 2.2, and 2.3. Flink 2.0 support is removed.
The Bottom Line
Apache Iceberg 1.12.0 is a maturity release with a wide blast radius. V3 types and metadata features work through more of the Java stack. REST catalogs gain stronger control-plane contracts. Streaming integrations close failure modes that matter in long-running jobs. Maintenance gets new tools and important correctness fixes.
The cost of that progress is a real compatibility boundary. Inventory Spark, Flink, delete files, and custom APIs before upgrading. Test shared tables by feature, then roll out from low-risk readers to high-impact writers and maintenance jobs.
For a deeper foundation, download Apache Iceberg: The Definitive Guide. It covers the table metadata model, catalogs, writes, reads, maintenance, and the design choices that make an upgrade like 1.12 easier to reason about.
Technical review date: September 30, 2026. Primary sources: the Apache Iceberg 1.12.0 release, the full 1.11.0 to 1.12.0 changelog, and the linked Apache Iceberg pull requests. Engine behavior can change in patch releases, so verify current documentation before enabling a feature in production.
Try Dremio Cloud free for 30 days
Deploy agentic analytics directly on Apache Iceberg data with no pipelines and no added overhead.
Intro to Dremio, Nessie, and Apache Iceberg on Your Laptop
Editor’s note, September 2026. This post was published in September 2023 and some product details have changed since. Nessie is still available and self-deployable under the Apache-2.0 licence, and it remains the clearest implementation of catalog-level branching. It is not an Apache Software Foundation project, and its development has slowed considerably. For how the current […]
Aug 16, 2023·Dremio Blog: News Highlights
5 Use Cases for the Dremio Lakehouse
With its capabilities in on-prem to cloud migration, data warehouse offload, data virtualization, upgrading data lakes and lakehouses, and building customer-facing analytics applications, Dremio provides the tools and functionalities to streamline operations and unlock the full potential of data assets.
Aug 24, 2026·Dremio Blog: Open Data Insights
Migrating to Apache Iceberg: Strategies for Every Source System
This is Part 15, the final article of a 15-part Apache Iceberg Masterclass. Part 14 covered hands-on Dremio Cloud. This article covers the three migration strategies and how to execute a zero-downtime migration using the view swap pattern. Most organizations do not start with Iceberg. They have years of data in Hive tables, data warehouses, CSV files, databases, […]