Dremio is now part of SAP
Dremio Blog

32 minute read · September 10, 2026

The Apache Polaris REST API: How Engines Talk to the Catalog

Alex Merced Alex Merced Head of DevRel, Dremio
The Apache Polaris REST API: How Engines Talk to the Catalog
Copied to clipboard

Apache Polaris implements the Iceberg REST Catalog specification, a set of HTTP endpoints that any Iceberg client can speak. The specification covers namespace and table management, an atomic commit protocol built on assertions rather than locks, server-side scan planning, and three different ways for a client to obtain storage access. Because the contract is the specification rather than a vendor SDK, an engine written against one implementation works against all of them.

“It speaks the Iceberg REST API” is the sentence every catalog vendor uses now. It has stopped meaning anything, because nobody says which parts.

The specification is 6,400 lines of OpenAPI. It covers far more than table lookup, and the parts that matter most are the ones people skip: how a commit avoids clobbering a concurrent writer, what the path prefix is doing in every URL, and the fact that credentials can arrive three different ways with a defined precedence between them.

This walks through what the API actually contains and how Polaris implements it.

What This Covers

The endpoint surface grouped by what it does, the commit protocol and its assertion model, what the prefix segment routes, how vended credentials and remote signing differ, server-side scan planning, and which parts of the spec are being deprecated.

The Endpoint Surface

THE ICEBERG REST CATALOG SURFACEDiscoveryGET /v1/configNamespacesGET|POST /v1/{prefix}/namespacesGET|HEAD|DELETE .../{namespace}POST .../{namespace}/propertiesTablesGET|POST .../namespaces/{ns}/tablesGET|POST|DELETE|HEAD .../tables/{table}POST .../register POST .../unregisterPOST /v1/{prefix}/tables/renameViewsGET|POST .../namespaces/{ns}/viewsGET|POST|DELETE .../views/{view}POST .../register-viewData accessGET .../tables/{table}/credentialsPOST .../tables/{table}/signScan planningPOST .../tables/{table}/planGET|DELETE .../plan/{plan-id}POST .../tables/{table}/tasksTransactionsPOST /v1/{prefix}/transactions/commit
One specification, and every engine that implements it can read and write the same tables.

It falls into seven groups, and reading them in order tells you the shape of the thing.

Discovery is one endpoint. A client calls GET /v1/config first, before anything else, and gets back catalog configuration plus any overrides it should apply. This is where the client learns things like which storage endpoint to use.

Namespaces handle the hierarchy. You can list, create, load metadata, check existence with a HEAD request, drop, and update properties. Dropping requires the namespace to be empty, which the spec states plainly.

Tables are the core. Create, load, commit updates, drop, check existence, rename. Two operations here get overlooked and both are useful: register adopts an existing metadata file into the catalog, and unregister removes a table from the catalog without deleting its data or metadata files.

Views mirror tables with their own create, load, drop, rename and register endpoints. Views are first-class in the spec, not an afterthought bolted onto tables.

register and unregister are the migration tools

If you have Iceberg tables written by something else and you want a catalog to know about them, register is how. You hand it a metadata file location and the catalog adopts the table without rewriting anything.

unregister is the reverse and it is the one people wish they had known about. It detaches a table from the catalog and leaves every file in place. That is what you want when moving a table between catalogs, and it is very much not what drop does.

A Table Load, Request by Request

Here is the actual sequence a client performs on first connection. Following it once makes the rest of the API obvious.

GET  /v1/config?warehouse=finance
     -> defaults and overrides the client should apply

GET  /v1/finance/namespaces
     -> [["analytics"], ["raw"], ["analytics","reporting"]]

GET  /v1/finance/namespaces/analytics%1Freporting/tables
     -> table identifiers in that namespace

GET  /v1/finance/namespaces/analytics%1Freporting/tables/revenue
     -> metadata-location, table metadata, storage-credentials

Two details in there catch people out. Multi-level namespaces are encoded with the unit separator character, not with dots, when they appear in a path segment. And the warehouse parameter on the config call is how a client tells the server which catalog it wants before it knows the prefix to use.

Everything after that first exchange is variations on the same shape: a path that identifies an object, a verb that says what to do with it, and a JSON body carrying the detail.

The Commit Protocol Is the Interesting Part

HOW A COMMIT STAYS ATOMICWriter Areads snapshot 7Writer Breads snapshot 7POST .../tables/{table}requirements + updates200 OKsnapshot 7 becomes 8409 Conflictassertion failedassert-ref-snapshot-id assert-table-uuid assert-current-schema-id assert-default-spec-id assert-createThe writer states what it believes to be true. The catalog checks, and the second writer is told to retry rather than silently winning.
Commits carry requirements, not just changes. That is what makes concurrent writes safe.

Committing a change to an Iceberg table through the REST API is not a write. It is a conditional write, and the condition travels with the request.

An update request carries two things: a list of updates describing what should change, and a list of requirements describing what the writer believes to be true right now. The catalog checks every requirement before applying any update. If one fails, the whole commit fails.

The requirement types are named for what they assert:

assert-create                     the table does not already exist
assert-table-uuid                 the table is the one I think it is
assert-ref-snapshot-id            a named ref still points at this snapshot
assert-current-schema-id          the schema has not changed underneath me
assert-default-spec-id            the partition spec has not changed
assert-default-sort-order-id      the sort order has not changed
assert-last-assigned-field-id     no fields were added since I read
assert-last-assigned-partition-id no partitions were added since I read

Two writers reading snapshot 7 and both trying to commit will both send assert-ref-snapshot-id naming snapshot 7. The first commit succeeds and the ref moves to snapshot 8. The second arrives, its assertion no longer holds, and it gets rejected. The client then refreshes and retries against snapshot 8.

This is optimistic concurrency, and it is why Iceberg does not need a lock manager for normal writes. No writer holds anything. They race, and the loser is told to try again with current information.

The practical consequence is that write contention shows up as retries in your job logs rather than as blocked writers or corrupted tables. If you see a lot of them, the answer is usually fewer concurrent writers to the same table rather than a catalog tuning knob.

Multi-table transactions exist

There is a POST /v1/{prefix}/transactions/commit endpoint that commits changes across several tables. Not every implementation supports it, and not every engine emits it, but it is in the specification. If you need two tables to move together, this is the mechanism rather than application-level coordination.

The Second API Nobody Mentions

The Iceberg REST Catalog specification covers tables, views and namespaces. It says nothing about who is allowed to touch them, because that is out of scope for a table format project.

So Polaris ships a second API alongside it. The management API handles principals, principal roles, catalog roles, grants and catalog registration. Two API surfaces, two audiences: the catalog API is what your query engines call thousands of times an hour, and the management API is what your platform team and your provisioning scripts call.

This split is worth understanding when you plan access. Nothing an engine does through the Iceberg API can change a permission, because permissions do not live in that API. A compromised Spark job cannot grant itself more access, it can only exercise what it already has.

It also means the management API is the one to lock down hardest. It is lower traffic and higher consequence.

What the Prefix Segment Actually Does

WHAT {prefix} IS FOROne clientone base URLCatalog service/v1/finance/...finance catalog/v1/marketing/...marketing catalog/v1/sandbox/...sandbox catalogThe prefix segment routes a request to one catalog inside the service. It is why a single endpoint can serve many catalogs without the client knowing.
Every path carries a prefix. It is the routing key, and most tutorials never explain it.

Nearly every path in the spec contains {prefix}. Most tutorials paste a URL with something in that slot and never explain it.

The prefix routes a request to a particular catalog inside the service. One deployment can host many catalogs, and the prefix is how a client says which one it means. That is why you can point Spark at a single base URL and have it work against several catalogs by changing configuration rather than infrastructure.

In Polaris this maps onto the catalog entity. Combined with realms, which partition at a level above catalogs, you get two independent axes of separation: realms for tenants, prefixes for catalogs within a tenant.

Three Ways to Get Storage Access

THREE WAYS A CLIENT GETS STORAGE ACCESSstorage-credentialsfield on LoadTableResultcheck this firstconfigfallback mapolder servers onlyremote signingPOST .../signcatalog signs each requestThe spec is explicit: clients must check storage-credentials before falling back to config. Remote signing suits environments where a credential must never leave the catalog.
Vended credentials are the common path. Remote signing exists for the stricter cases.

This is the part where implementations diverge most, and the spec is unusually direct about precedence.

The modern path is the storage-credentials field on the load result. It carries credentials for S3, ADLS, GCS and others in a structured array. The spec instructs clients to check this field first, before looking anywhere else.

The older path is the config map. Some servers still return credentials there, and clients fall back to it when storage-credentials is absent. If you are debugging a client that cannot authenticate, knowing which of the two your server populates resolves it quickly.

There is also a dedicated GET .../tables/{table}/credentials endpoint for loading vended credentials separately, which is how a long-running client refreshes before expiry rather than reloading the whole table.

Remote signing is the strict option

The third path is different in kind. With POST .../tables/{table}/sign, the client never receives a credential at all. It asks the catalog to sign each storage request, and the catalog returns a signed URL or headers.

That is slower and chattier, and it suits environments where a storage credential must never exist outside the catalog process. If your security review has ever objected to handing temporary keys to a compute cluster, this is the answer to that objection.

Error Responses Are Part of the Contract

The specification defines error shapes, and treating them as noise is how you end up with clients that retry the wrong things.

A 409 Conflict on a commit means an assertion failed. Someone else committed first. This is retryable, and the correct retry refreshes table state before trying again rather than resending the same requirements.

A 404 on a table load means the table is not in the catalog. That is different from the metadata file being missing in storage, which surfaces later and looks like an I/O failure. Distinguishing them saves real debugging time.

A 403 means authorization failed at the catalog. In Polaris that maps to a missing privilege, and the privilege it maps to is usually more specific than people expect. A principal can load a table's metadata and still be refused a credential for its data, because those are separate grants.

A 400 on create or rename is often name validation. Polaris enforces entity name rules at the REST layer and rejects names containing control characters or any of /:*?"<>|#+`, along with names that are empty or that start or end with whitespace.

Server-Side Scan Planning

A newer part of the spec moves scan planning to the catalog. Instead of the client reading manifests to work out which files it needs, it submits a scan and the catalog does the work.

The flow uses four endpoints. POST .../tables/{table}/plan submits a scan and returns a plan id. GET .../plan/{plan-id} fetches the result, since planning can be asynchronous. POST .../tables/{table}/tasks fetches the concrete scan tasks. And DELETE .../plan/{plan-id} cancels planning that is no longer needed.

Why bother? Because manifest reading is I/O against object storage, and a client planning a scan over a large table issues a lot of small reads. A catalog that already holds or caches that metadata can answer faster, and a thin client stops needing the full Iceberg library to plan a query.

Support varies. Check whether your engine emits these calls before designing around them.

Versioning and What That Means for You

The specification evolves. Server-side scan planning arrived after the initial release, functions endpoints are newer still, and the OAuth token endpoint is on its way out.

The config endpoint is how clients and servers negotiate around this. A server tells the client what it supports and what defaults to apply, and a well-behaved client adjusts. That is why the first call in every session is GET /v1/config rather than a version header.

Practically, mixed versions are normal. You will run an engine built against one revision of the spec against a catalog implementing another. Things mostly work, because the endpoints that matter have been stable, and the newer additions are additive.

Where it bites is optional features. If your engine does not emit scan planning calls, server-side planning does nothing for you no matter what the catalog supports. Check both ends before building a plan around a capability.

What Is Being Removed

/v1/oauth/tokens is marked in the spec as deprecated for removal. It exists so a client can obtain a token through an OAuth2 flow against the catalog itself.

The direction is clear: authentication belongs to an identity provider, not to the catalog. If your client configuration currently obtains tokens from the catalog endpoint, that is worth migrating before it stops working. Polaris supports external identity providers, with Keycloak as the documented path, and that is where this is heading.

Worth checking your own configs for this one. It tends to be buried in a Spark properties file written two years ago and never revisited.

Metrics Reporting

There is a POST .../tables/{table}/metrics endpoint that lets clients report scan metrics back to the catalog. It is optional and easy to ignore, and it is also the only standardised way a catalog learns how tables are actually being queried.

That data is what makes automated maintenance possible. A catalog that knows which partitions get scanned and which files get skipped can make better decisions about compaction and clustering than one working from table size alone.

What Implements This Today

The list matters because it is the actual measure of whether a specification is real.

On the engine side, Apache Spark, Apache Flink, Trino, StarRocks, Apache Doris and Dremio all speak the Iceberg REST Catalog API. PyIceberg brings it to Python without a JVM, which has done more for adoption than any of the JVM engines.

On the server side, Apache Polaris is one implementation. There are others, including managed services and other open source projects. That is the point. A specification with one implementation is a product with documentation.

Writing your own client is realistic

Because the contract is HTTP and JSON, a minimal read-only client is a few hundred lines. Call config, list namespaces, load a table, parse the metadata location, read the Iceberg metadata from storage. People do this to build catalog browsers, lineage tools and cost reporters without pulling in the full Iceberg library.

Common Questions

Is the Iceberg REST Catalog API the same as Polaris?

No. The API is a specification maintained by the Apache Iceberg project. Polaris is one implementation of it, plus management endpoints of its own for principals, roles and grants that the Iceberg spec does not cover.

What is the difference between drop and unregister?

Drop removes the table from the catalog and can remove its files. Unregister removes the catalog's reference and leaves every file in place. Use unregister when moving a table between catalogs.

Why did my commit return 409?

Another writer committed to the same table between your read and your write, so one of your requirements no longer held. Refresh and retry. Frequent 409s usually mean too many concurrent writers on one table rather than a misconfiguration.

Do I need server-side scan planning?

Not today. It is in the specification and support is uneven across engines. It matters most for very large tables and for thin clients that would rather not read manifests themselves.

Should I still use /v1/oauth/tokens?

No. It is marked deprecated for removal in the specification. Move authentication to an external identity provider before it disappears.

Can two different vendors' engines write to the same catalog?

Yes, and that is the whole point of the specification. The commit protocol is designed for concurrent writers that have never heard of each other, which is why it uses assertions rather than locks.

Why the Specification Matters More Than the Implementation

Here is the argument this whole post is building toward.

An open table format gives you portable files. It does not give you portable access. If the only thing that can find your Iceberg tables is one vendor's catalog, your data is open in the same way a locked filing cabinet is transparent.

The REST specification is what closes that gap. Because the contract is HTTP and JSON rather than a Java SDK, any language can implement a client, and any catalog can implement a server. Apache Spark, Apache Flink, Trino, StarRocks, Apache Doris and Dremio all speak it today.

That is the test worth applying to any catalog you are evaluating. Not whether it claims REST support, but whether a client from a different vendor can create a table in it, and whether your engine can read a table someone else created.

Where Dremio Fits

Dremio co-created Apache Polaris and donated it to the Apache Software Foundation. Dremio's Open Catalog is built on Polaris, so it presents the same Iceberg REST endpoints described here.

The practical effect is that a table created through Dremio is immediately readable by Spark or Flink without an export, and a table created by an external pipeline shows up in Dremio without a sync job. Both are speaking the same specification to the same catalog.

Dremio's federated query engine then sits above that, so a query can join an Iceberg table to a table in PostgreSQL or Snowflake without moving either, and Reflections can accelerate the result by up to 100x. The catalog underneath stays a standard implementation that anything else can talk to.

Read the Spec Once

The Iceberg REST Catalog specification is public, versioned, and more readable than its line count suggests. An hour with it will tell you more about what your catalog can and cannot do than any vendor comparison.

Start with the commit endpoint and the requirement types. Once you understand that a commit is an assertion plus a change, the rest of the design follows, including why concurrent writers are safe and why your retry logic looks the way it does.

One last thing worth saying about specifications generally. The reason this one works is that it was written by a project that had no engine to sell. Apache Iceberg defines the format and the catalog contract, and then gets out of the way of whoever implements them.

A catalog API designed by a vendor would have been shaped by that vendor's engine. This one was not, which is why six engines from different companies can all write to the same table and none of them has an advantage in doing so.

Related technical guides: what Apache Polaris is and how it governs Iceberg tables; using Apache Polaris with Trino.

Keep learning

For a full treatment of the catalog layer, download Apache Polaris: The Definitive Guide by Alex Merced, Andrew Madson, and Tomer Shiran, free from Dremio.

To see these ideas applied in practice, explore Dremio’s Open Catalog, built on Apache Polaris.

Try Dremio Cloud free for 30 days

Deploy agentic analytics directly on Apache Iceberg data with no pipelines and no added overhead.