Dremio is now part of SAP
Dremio Blog

37 minute read · September 10, 2026

Choosing an Apache Iceberg Catalog: Polaris, Gravitino, Unity, Lakekeeper and More

Alex Merced Alex Merced Head of DevRel, Dremio
Choosing an Apache Iceberg Catalog: Polaris, Gravitino, Unity, Lakekeeper and More
Copied to clipboard

Choosing an Apache Iceberg catalog comes down to three questions, and the feature matrices most vendors publish do not help with any of them. Do you need more than one engine to write to the same tables? Do you need governance you do not control to be predictable? And do you need to keep the catalogs you already run while you migrate? Answer those and the field narrows from seven options to about two.

Every Iceberg catalog comparison you will read is a table with checkmarks. REST API: yes. Multi-engine: yes. Credential vending: yes. After twenty rows everything looks equivalent, which is exactly what the vendor publishing the table wanted.

The differences that matter are not features. They are who decides the roadmap, whether anyone is still writing code, and what happens to the catalog you are already running.

This compares Apache Polaris, Apache Gravitino, Unity Catalog, Lakekeeper, Nessie, AWS Glue Data Catalog and Hive Metastore on those terms, with the evidence for each claim.

First, a Distinction Almost Everyone Blurs

LICENCE IS NOT GOVERNANCECATALOGLICENCEWHO SETS THE ROADMAPApache PolarisApache-2.0ASF Top-Level ProjectApache GravitinoApache-2.0ASF Top-Level ProjectHive MetastoreApache-2.0ASF Top-Level Project, part of Apache HiveUnity CatalogApache-2.0LF AI & Data, sandbox stageLakekeeperApache-2.0Vakamo, a single companyNessieApache-2.0Small group, Dremio in originAWS Glue Data CatalogProprietaryAmazon Web ServicesSix of the seven ship under the same licence. The licence says what you may do with the code. It says nothing about who decides what goes into it.
Apache-2.0 is a licence. The Apache Software Foundation is a governance body. They are not the same claim.

Six of the seven catalogs here ship under the Apache-2.0 licence. Only AWS Glue does not. That shared licence is used in marketing to imply something it does not mean.

Apache-2.0 is a licence. It governs what you may do with the code. Anyone can apply it to anything, and plenty of single-vendor projects do.

The Apache Software Foundation is a governance body. It decides how decisions get made: public mailing lists, recorded votes, a project management committee, and a requirement that no single company controls the project. Getting there means going through incubation and graduating.

Three of the six open source catalogs here are ASF projects: Polaris, Gravitino, and Hive Metastore, which ships as part of Apache Hive. Unity Catalog sits under the LF AI & Data Foundation at sandbox stage, the earliest tier that foundation has. Lakekeeper is built and maintained by Vakamo, and Nessie is maintained by a small group with roots at Dremio.

None of that is a criticism. Lakekeeper in particular is a well-built piece of software, and a single-vendor project can be excellent. But if your reason for choosing an open catalog is that you do not want one company deciding your data platform's future, the licence does not give you that. The governance does.

Why this matters more for a catalog than for most software

You can swap a query engine in a quarter. Swapping a catalog means re-registering every table and rebuilding every grant, and it means doing it while jobs are running.

The catalog is the component with the longest practical lifespan and the highest switching cost. That is exactly the component where you want the roadmap decided somewhere you can watch, by people you can email.

What the foundation actually enforces

It is tempting to treat foundation membership as a badge. It is closer to a set of ongoing requirements, and the requirements are the reason it predicts project health.

An ASF project has to make decisions on public lists where any user can read the thread and reply to it. Releases pass by recorded vote. The project management committee is required to recruit new committers on merit, which is the mechanism that keeps a single employer from quietly owning the commit bit. Projects that stop attracting contributors get moved to the Attic, and that outcome is visible rather than silent.

None of that guarantees good software. What it does is make community diversity a condition of continued existence rather than a nice-to-have, and diversity of contributors is what keeps a project alive after the company that started it changes strategy. Look at what happened to Nessie for the version of that story where the company moved on.

The catalog that wins will write the next standard

Here is the part that makes governance more than a philosophical preference.

Apache Iceberg specifies the REST catalog interface for Iceberg tables. That part is settled, it is maintained by a foundation project with no engine to sell, and every catalog in this article implements it. Interoperability for Iceberg tables is not up for grabs.

Everything else in the lakehouse is up for grabs. Models, vector indexes, unstructured assets, functions, ML features, semantic definitions, lineage. There is no ratified interface for any of them, and every catalog is extending in that direction right now. Whichever catalog reaches critical mass will have its extensions treated as the de facto standard, the way a dominant implementation always does.

So the question is not only which catalog you run. It is who gets to define the API that the rest of the ecosystem builds against for the next decade, and whether that definition happens on a public mailing list or in a product planning meeting you will never see.

Second, Check Whether Anyone Is Still Building It

Star counts are a popularity metric from years ago. The question worth asking is how many different people are writing code right now, and it takes about a minute to check.

HOW MANY PEOPLE WROTE THE LAST 100 COMMITSApache Gravitino23 peopleUnity Catalog21 peopleApache Polaris17 peopleLakekeeper10 peopleNessie3 peopleDistinct human authors in the last 100 commits, measured from public commit history in September 2026. Four of these projects have a real bench. One has three people.
Star counts measure attention. The size of the committer bench measures whether anyone is still building.

Here is what the public commit history showed in September 2026, counting the distinct human authors behind the last 100 commits on each project:

Apache Gravitino   3,212 stars   23 people
Unity Catalog      3,517 stars   21 people
Apache Polaris     2,053 stars   17 people
Lakekeeper         1,448 stars   10 people
Nessie             1,509 stars    3 people

Four of these have a real bench. One does not.

Seventeen people have committed to Polaris in its last hundred commits, drawn from more than one employer, which is the shape you want under a component you cannot easily replace. A project with a bench that size survives any single contributor changing jobs. Gravitino and Unity Catalog are in similar territory, and Lakekeeper at ten is respectable for a project that size.

Nessie is the outlier. Its repository is not archived and it received commits the day I checked, so a quick glance suggests health. Look at who wrote them and it is three people.

Run this check yourself

Do not take my numbers. They will be stale by the time you read this, and the method matters more than the result. This prints every author in the last hundred commits with a count, so you can see the bot entries and ignore them:

curl -s 'https://api.github.com/repos/OWNER/REPO/commits?per_page=100' \
  | jq -r '.[].commit.author.name' | sort | uniq -c | sort -rn

Count the human names that remain. If there are only two or three, you are looking at a project in maintenance. That is not automatically disqualifying, but you should know it before you build on it.

The Seven Options, Honestly

Apache Polaris

An ASF Top-Level Project, co-created by Dremio and donated to the foundation. Implements the Iceberg REST Catalog specification, with RBAC over catalogs, namespaces, tables, views and policies, credential vending on AWS, Azure and GCP, realms for multi-tenancy, and federation to Hive Metastore, BigQuery Metastore and other REST catalogs.

Engines that speak to it include Apache Spark, Apache Flink, Trino, StarRocks, Apache Doris and Dremio. Available as software you run and as a managed service inside commercial platforms.

Choose it when you want foundation governance, multi-engine writes, and a migration path that does not require moving everything at once.

Apache Gravitino

Also an ASF Top-Level Project, and the most active of the group by human commit count. Its emphasis is broader than Iceberg: it aims to be a metadata catalog across multiple data sources and formats, with Iceberg REST support as one part of that.

Choose it when your problem is genuinely heterogeneous metadata across many systems rather than Iceberg specifically. If you only have Iceberg tables, its extra scope is surface area you will not use.

Unity Catalog

Open sourced by Databricks, Apache-2.0, and the highest star count in the group. Real development activity with twenty-one people committing in the recent history. It sits under the LF AI & Data Foundation at sandbox stage, that foundation's earliest tier.

Choose it when you are substantially invested in Databricks and want consistency between the managed product and your open deployments. Be clear-eyed that roadmap direction follows one company's commercial interest, which is fine as long as that interest stays aligned with yours.

Lakekeeper

An independent Iceberg REST Catalog implementation written in Rust. Apache-2.0, actively developed, ten people committing recently, and built and maintained by Vakamo.

Choose it when you want a lean, fast implementation and you are comfortable depending on a smaller project. The Rust implementation means a small memory footprint and no JVM, which matters if you are running catalogs per environment.

Nessie

Nessie brought git-style branching and merging to the catalog layer, and it remains the clearest implementation of that idea. You can still self-deploy it, it is Apache-2.0, and the repository is not archived.

Two things to be straight about. Nessie is not an Apache Software Foundation project, despite the licence and despite how often it gets written as though it were. And development has slowed considerably: three people wrote the human-authored commits in its last hundred.

The modern Dremio platform is built on Apache Polaris, not Nessie. If you are running Nessie today it keeps working, and if you depend specifically on catalog-level branching it still does that better than the alternatives. For a new deployment, the activity numbers should give you pause.

AWS Glue Data Catalog

A managed service, not open source, and only on AWS. It works, plenty of production lakehouses run on it, and it integrates with the rest of AWS without effort.

Choose it when you are entirely on AWS and value not operating a service over portability. Avoid it when you have engines outside AWS or expect to.

Hive Metastore

The one most teams already have. It ships as part of Apache Hive, which means it is Apache-2.0 and governed as an ASF Top-Level Project, the same standing as Polaris and Gravitino. Anyone who tells you Hive Metastore is closed or proprietary is wrong.

Its problem is not openness, it is design. It predates Iceberg by roughly a decade. It speaks Thrift rather than REST, it has no credential vending, its permission model predates everything the current catalogs do, and it works with Iceberg only through a catalog implementation layered on top. It also tends to arrive with a relational database you now have to keep alive.

Few teams choose Hive Metastore for a new Iceberg deployment in 2026. The realistic question is not whether to keep it, it is how to stop depending on it without a big bang migration. Federation is the answer, which is the next section.

Third, What Happens to What You Already Run

WHICH CATALOGS FRONT OTHER CATALOGSYour enginesone REST endpointPolaris or Gravitinofederating catalogHive Metastorelegacy, still in useAWS Gluemanaged, AWS onlyAnother REST catalogany implementationFederation is what makes adoption incremental. You do not migrate a hundred jobs on a weekend, you put a catalog in front of the old one and move namespaces when convenient.
Federation turns a migration into a sequence of small moves.

This is the question that decides most real evaluations, and it rarely appears in feature matrices.

You have a Hive Metastore with several hundred tables and a hundred jobs pointing at it. No comparison chart addresses what to do about that.

Polaris and Gravitino can both federate. You put the new catalog in front of the old one, clients connect to a single endpoint, and the catalog resolves each namespace to whichever backing store owns it. New tables get created natively. Old ones move when someone has time.

Without federation, adoption is a migration project with a cutover date, and cutover dates on catalogs are how you end up with a frozen quarter.

What Switching Actually Costs

Every comparison assumes you are choosing from nothing. Most readers are not. Here is what moving between catalogs really involves, because it should weight your decision heavily.

The tables themselves do not move

This is the part people get wrong in both directions. Iceberg data and metadata files sit in object storage. A catalog holds a pointer to the current metadata file for each table. Switching catalogs means re-pointing, not rewriting.

The Iceberg REST specification has a register operation that adopts an existing metadata file into a catalog. There is also an Iceberg Catalog Migrator tool for doing this in bulk. So the data movement cost is genuinely zero.

The grants do move, and painfully

Permissions are the real cost. Every catalog models them differently, and there is no standard for exporting a permission graph and importing it somewhere else. Polaris uses principals, principal roles, catalog roles and privileges. Glue uses IAM and Lake Formation. Hive Metastore predates most of this.

Budget for rebuilding your permission model by hand, and for a period where two systems both think they are authoritative. That window is where incidents happen.

The client configuration moves too

Every Spark job, every Trino catalog definition, every notebook with a hardcoded catalog URI. Federation reduces this to changing one endpoint rather than a hundred, which is most of why federation matters.

Can You Run More Than One?

Yes, and plenty of organisations do. It is worth saying because the question is usually framed as a single irreversible choice.

Different teams can run different catalogs against the same object storage, as long as they do not both claim the same tables. The failure mode to avoid is two catalogs holding pointers to one table, because Iceberg's commit safety depends on a single authority for that pointer. Two catalogs means two authorities and no coordination between them.

So the rule is straightforward: one catalog owns a table. Multiple catalogs can coexist across different tables, and federation is how you present them as one thing to clients.

A Decision Path

A DECISION PATH THAT ACTUALLY NARROWS THINGSDo you need more than one engine to write?No: your engine's built-in catalog is fineDo you need foundation governance?No: Lakekeeper or Unity Catalog are credibleDo you need to front catalogs you already run?No: Polaris covers the common caseDo you need git-style branching over the catalog?No: stop hereMost teams stop at question one or two. The catalogs differ far less than the marketing suggests once you have answered them.
Four questions. Most evaluations should end after the second one.

Work through it in order and most of the field falls away quickly.

If only one engine writes, your engine's built-in catalog is genuinely fine. The whole argument for a standalone catalog is multi-engine consistency. Do not take on an extra service to solve a problem you do not have.

If foundation governance does not matter to you, Lakekeeper and Unity Catalog are both credible and actively built. Pick on operational fit.

If you need to front existing catalogs, you are choosing between Polaris and Gravitino, and the tiebreaker is scope. Iceberg-focused points at Polaris, broad metadata across many systems points at Gravitino.

If you need catalog-level branching, Nessie is still the answer, with the activity caveat above understood and accepted.

The Criteria Worth Scoring

If you do need a matrix, score these rather than feature checkmarks. Each one has a real consequence attached.

Who can change the roadmap

A foundation project changes through public proposals and recorded votes. A single-vendor project changes when that vendor's product strategy changes. Both are legitimate. Only one lets you see it coming.

How many people are writing code

Three human authors is a bus-factor problem. Twenty is a project. This is checkable in a minute and it is the single most predictive number available to you.

Whether it federates

If it cannot sit in front of what you already run, adoption becomes a migration with a cutover date. Polaris and Gravitino federate. Most of the others do not.

What it does about storage credentials

Ask whether the catalog vends short-lived scoped credentials or expects your engines to hold standing keys. This is the difference between the catalog enforcing access and merely describing it.

How permissions are modelled

Specifically, whether metadata access and data access are separable. Being able to let a discovery tool enumerate every table without giving it a key to any of them turns out to matter a lot once you have more than a handful of integrations.

What it costs to leave

Register and unregister operations mean tables move cheaply. Permission graphs do not. Ask how you would export the grants before you import the first table.

What Not to Decide On

Star counts. Unity Catalog has more stars than Polaris and Nessie has more than Lakekeeper. Neither fact predicts anything about whether the project will be maintained in three years.

Feature checkmarks. Every catalog in this list implements the Iceberg REST specification. The specification is what makes them interchangeable, so a table showing that they all support it is a table showing they all did the required work.

Benchmark numbers. The catalog is not in the data path. It answers metadata questions during query planning. Unless your catalog is pathologically slow, it is not your bottleneck, and vendors publishing catalog benchmarks are competing on a dimension that rarely limits anyone.

The word Apache in a name or a licence file. Governance matters a great deal, as the whole first section argues. The word does not tell you whether you have it. Nessie is Apache-2.0 and not an ASF project. Hive Metastore carries no Apache in the name most people use for it and is an ASF project. Gravitino graduated from incubation and is one too. Check the project page for a PMC and a public list, not the string in the repository name.

Common Questions

Is Nessie an Apache project?

No. Nessie is Apache-2.0 licensed, which is a licence, not a governance model. It is not an Apache Software Foundation project. Apache Polaris and Apache Gravitino are.

Is Nessie dead?

No. The repository is active and not archived, and you can still self-deploy it. Development has slowed a lot though: three people wrote the human-authored commits in its last hundred, against seventeen on Polaris and twenty-three on Gravitino. It still does catalog-level branching better than anything else. Weigh that against the maintenance profile.

Does Dremio still use Nessie?

The modern Dremio platform is built on Apache Polaris. Dremio's Open Catalog is a Polaris implementation.

Can I switch catalogs later?

Yes, and it is cheaper than people expect on the data side and more expensive on the permissions side. Tables register into a new catalog without moving files. Grants have to be rebuilt by hand because no standard exists for moving them.

Do I need a catalog at all?

If more than one engine writes to your tables, yes. The catalog is what makes the metadata pointer swap atomic, and without it concurrent writers can silently overwrite each other. If exactly one engine ever writes, its built-in catalog is enough.

What about the Hadoop catalog?

It stores the current metadata pointer in a file in object storage. It works for a single writer and it is not recommended for production with multiple writers, because object stores do not all give the atomic file update guarantees it depends on.

The Specification Is Doing the Heavy Lifting

Worth naming what makes this comparison possible at all. The Iceberg REST Catalog specification is maintained by the Apache Iceberg project, which has no engine to sell.

Because the contract is HTTP and JSON rather than a vendor SDK, any of these catalogs can serve any compliant client. That is why switching is a re-registration rather than a rewrite, and why an engine built against one implementation works against the others.

It also means the catalogs compete on operations, governance and durability rather than on lock-in. That is a healthier market than the one Iceberg replaced, and it is worth remembering when a comparison table tries to make the specification itself sound like a differentiator.

Note the boundary though. The specification covers Iceberg tables. It says nothing about models, unstructured assets, functions or semantic definitions, and those are exactly where the catalogs are extending. On Iceberg tables you are protected by a ratified contract. On everything else you are betting on whoever writes the interface first, which is why the governance question in the first section is not academic.

Where Dremio Fits

Dremio co-created Apache Polaris and donated it to the Apache Software Foundation. Dremio's Open Catalog is built on Polaris, so tables created through Dremio are readable by Spark, Flink, Trino, StarRocks and Apache Doris without an export step.

Worth stating the obvious bias here: Dremio has a horse in this race and this article appears on Dremio's blog. The numbers above are checkable, and the method for checking them is in the article for exactly that reason.

The stronger argument is not that you should use Polaris because Dremio built it. It is that Dremio gave it away, to a foundation, where a project management committee that includes people from other companies decides what happens next. That is a costlier signal than a licence file.

The Test That Settles It

Take the catalog you are considering. Create a table with one engine. Read it with a different engine from a different vendor. Then revoke a grant and confirm the read stops working.

That exercise takes an afternoon and tells you more than every comparison table ever published, including this one. If any step fails, the catalog is not doing the job you are hiring it for, whatever the feature list says.

Keep learning

For a full treatment of the catalog layer, download Apache Polaris: The Definitive Guide by Alex Merced, Andrew Madson, and Tomer Shiran, free from Dremio.

To see these ideas applied in practice, explore Dremio’s Open Catalog, built on Apache Polaris.

Try Dremio Cloud free for 30 days

Deploy agentic analytics directly on Apache Iceberg data with no pipelines and no added overhead.