The Dremio Blog
Dremio Blog: Open Data Insights
-
Dremio Blog: Open Data InsightsMigrating to Apache Iceberg: Strategies for Every Source System
This is Part 15, the final article of a 15-part Apache Iceberg Masterclass. Part 14 covered hands-on Dremio Cloud. This article covers the three migration strategies and how to execute a zero-downtime migration using the view swap pattern. Most organizations do not start with Iceberg. They have years of data in Hive tables, data warehouses, CSV files, databases, […] -
Dremio Blog: Open Data InsightsHands-On with Apache Iceberg Using Dremio Cloud
This article is a practical walkthrough of working with Iceberg on Dremio Cloud, covering table creation, data ingestion, optimization, semantic layer construction, and AI-powered analytics. -
Dremio Blog: Open Data InsightsApproaches to Streaming Data into Apache Iceberg Tables
Streaming ingestion bridges this gap by committing data to Iceberg tables at regular intervals. The challenge is that frequent commits create the small file problem, and managing that trade-off between data freshness and table health is the central concern of streaming to Iceberg. -
Dremio Blog: Open Data Insights
Using Apache Iceberg with Python and MPP Query Engines
This article covers the two main ways to access Iceberg data: directly from Python libraries and through MPP (massively parallel processing) query engines. -
Dremio Blog: Open Data Insights
Apache Iceberg Metadata Tables: Querying the Internals
This article covers the metadata tables that let you inspect Iceberg table internals using standard SQL. Iceberg exposes its internal metadata as queryable virtual tables. -
Dremio Blog: Open Data Insights
Maintaining Apache Iceberg Tables: Compaction, Expiry, and Cleanup
This article covers the four maintenance operations that keep Iceberg tables healthy and the three approaches to running them. -
Dremio Blog: Open Data Insights
How Data Lake Table Storage Degrades Over Time
An Iceberg table that works well on day one will not work well on day 365 without maintenance. Every append, update, and delete operation adds files and metadata. -
Dremio Blog: Open Data Insights
When Catalogs Are Embedded in Storage
This article examines a newer approach: embedding the catalog directly inside the storage layer. Traditional Iceberg architectures have three components: the query engine, a standalone catalog, and object storage. -
Dremio Blog: Open Data Insights
What Are Lakehouse Catalogs? The Role of Catalogs in Apache Iceberg
A lakehouse catalog is the component that answers one question: "Where is the current metadata for this table?" Without a catalog, every engine would need to independently locate and track metadata files. With a catalog, there is a single source of truth that coordinates reads, writes, and access control across all engines. -
Dremio Blog: Open Data Insights
Enterprise Agentic Analytics Explained
Learn how agentic workflows for enterprise analytics connect AI agents, governed data and multi-step analysis to improve complex business decisions. -
Dremio Blog: Open Data Insights
Writing to an Apache Iceberg Table: How Commits and ACID Actually Work
Understanding the write process is critical because it explains why Iceberg can provide ACID guarantees on top of object storage, something that seems impossible when you consider that S3, ADLS, and GCS have no built-in transaction support. -
Dremio Blog: Open Data Insights
Agentic Lakehouse: The Architecture Built for AI-Native Analytics
The Agentic Lakehouse is not a new name for the same architecture. It represents a genuine shift in what a data platform is responsible for. A traditional lakehouse is a managed repository. An Agentic Lakehouse is an active participant in AI workflows: it provides context, enforces governance, and optimizes itself autonomously. -
Dremio Blog: Open Data Insights
Text-to-SQL vs Agentic Analytics: What the Upgrade Requires
Text-to-SQL on a governed semantic layer is significantly more reliable than text-to-SQL on a raw production schema. The semantic layer constrains what the model can access, provides business-friendly terminology, and enforces metric definitions. The accuracy improvement is material. -
Dremio Blog: Open Data Insights
Semantic Layer vs Data Catalog: What’s the Difference?
The convergence of AI agents, open table formats, and semantic tooling is making this architecture decision more consequential than it was a few years ago. AI agents that query through ungoverned raw tables or that cannot discover what data exists are not reliable. -
Dremio Blog: Open Data Insights
Hidden Partitioning: How Iceberg Eliminates Accidental Full Table Scans
The most expensive mistake in data lake querying is the accidental full table scan: a query that reads every file because the user did not correctly reference the partition columns. In Hive, this happens constantly. In Iceberg, it is structurally impossible because users never reference partition columns at all.
- 1
- 2
- 3
- …
- 15
- Next Page »