Featured Articles
Popular Articles
-
Dremio Blog: Open Data Insights
Apache Iceberg 1.12.0: What’s New, Breaking Changes, and Upgrade Guide
-
Dremio Blog: Open Data InsightsApache Polaris 1.8.0: What’s New, Breaking Changes, and Upgrade Guide
-
Dremio Blog: Open Data InsightsApache Iceberg Table Encryption: KMS, Envelopes, and Parquet Encryption
-
Dremio Blog: Open Data InsightsApache Iceberg Views: Portable View Metadata Across SQL Engines
Browse All Blog Articles
-
Dremio Blog: Open Data Insights
Choosing an Apache Iceberg Catalog: Polaris, Gravitino, Unity, Lakekeeper and More
Choosing an Apache Iceberg catalog comes down to three questions, and the feature matrices most vendors publish do not help with any of them. Do you need more than one engine to write to the same tables? Do you need governance you do not control to be predictable? And do you need to keep the catalogs you already run while you migrate? Answer those and the field narrows from seven options to about two. -
Dremio Blog: Open Data Insights
Access Control in Apache Polaris: RBAC and Credential Vending
Apache Polaris controls access through role-based access control with four levels: a privilege is granted to a catalog role, the catalog role is granted to a principal role, and the principal role is assigned to a principal. Privileges never attach to a person directly. The two privileges that matter most, TABLE_READ_DATA and TABLE_WRITE_DATA, are the only ones that cause Polaris to issue short-lived storage credentials, and they sit outside the metadata grant that covers everything else. -
Dremio Blog: Open Data Insights
Apache Polaris Architecture Explained
Apache Polaris is a catalog service for Apache Iceberg. It answers the question “which files make up this table right now, and are you allowed to touch them?” Its architecture has four moving parts: an entity model that nests catalogs, namespaces and tables inside a realm, a persistence layer that stores those entities and their grants, a REST service that speaks the Iceberg REST Catalog API, and a credential broker that hands engines short-lived, path-scoped access to object storage. -
Dremio Blog: Open Data Insights
State of the Open Lakehouse, September 2026
The state of the open lakehouse in September 2026 is a stack that stopped arguing about whether it won and started dealing with the consequences of winning. -
Dremio Blog: Open Data Insights
Migrating to Apache Iceberg: Strategies for Every Source System
This is Part 15, the final article of a 15-part Apache Iceberg Masterclass. Part 14 covered hands-on Dremio Cloud. This article covers the three migration strategies and how to execute a zero-downtime migration using the view swap pattern. Most organizations do not start with Iceberg. They have years of data in Hive tables, data warehouses, CSV files, databases, […] -
Dremio Blog: Open Data Insights
Hands-On with Apache Iceberg Using Dremio Cloud
This article is a practical walkthrough of working with Iceberg on Dremio Cloud, covering table creation, data ingestion, optimization, semantic layer construction, and AI-powered analytics. -
Dremio Blog: Open Data Insights
Approaches to Streaming Data into Apache Iceberg Tables
Streaming ingestion bridges this gap by committing data to Iceberg tables at regular intervals. The challenge is that frequent commits create the small file problem, and managing that trade-off between data freshness and table health is the central concern of streaming to Iceberg. -
Dremio Blog: Open Data Insights
Using Apache Iceberg with Python and MPP Query Engines
This article covers the two main ways to access Iceberg data: directly from Python libraries and through MPP (massively parallel processing) query engines. -
Dremio Blog: Various Insights
Apache Iceberg v4: An Efficiency Rewrite of the Table Format
Iceberg 1.11.0, released in May 2026, contains the full implementation of the v3 spec(ification), including features like deletion vectors, the VARIANT type, and row lineage. But what comes next? V4, obviously, as four comes after 3. However, it's still early days on the v4 spec, so we are a ways off from the next version […] -
Dremio Blog: Open Data Insights
Apache Iceberg Metadata Tables: Querying the Internals
This article covers the metadata tables that let you inspect Iceberg table internals using standard SQL. Iceberg exposes its internal metadata as queryable virtual tables. -
Dremio Blog: Open Data Insights
Maintaining Apache Iceberg Tables: Compaction, Expiry, and Cleanup
This article covers the four maintenance operations that keep Iceberg tables healthy and the three approaches to running them. -
Dremio Blog: News Highlights
Apache Ossie (Incubating): The New Name for Open Semantic Interchange
Apache Ossie is currently undergoing incubation at The Apache Software Foundation (ASF). If you've been following the Open Semantic Interchange project — the open specification for semantic layer and ontology — there's an important update. The project has been accepted into the Apache Incubator under a new name: Apache Ossie (Incubating). The spec, the community, and […] -
Dremio Blog: Open Data Insights
How Data Lake Table Storage Degrades Over Time
An Iceberg table that works well on day one will not work well on day 365 without maintenance. Every append, update, and delete operation adds files and metadata. -
Dremio Blog: Various Insights
What’s The Deal With Apache Parquet?
Apache Parquet is the recommended file format used in every modern data platform, and for good reason. But what are those reasons? And would it really matter if you stuck with CSV? The short answer is "YES". The slightly longer answer is "Yes, because columns". The full answer is below, so keep on reading to […] -
Dremio Blog: Open Data Insights
When Catalogs Are Embedded in Storage
This article examines a newer approach: embedding the catalog directly inside the storage layer. Traditional Iceberg architectures have three components: the query engine, a standalone catalog, and object storage.
- « Previous Page
- 1
- 2
- 3
- 4
- 5
- …
- 48
- Next Page »
