Browse all Subsurface content →
Menu
Data Lake Storage
Amazon S3
Azure Data Lake Storage
Google Cloud Storage
File Formats
Apache Parquet
Table Formats
Apache Iceberg
Delta Lake
Metastores
AWS Glue
Hive Metastore
Nessie
Data Lake Engines
Dremio
Apache Spark
Interfaces
Apache Arrow Flight
In-Memory Formats
Apache Arrow
Data Catalog
Amundsen
Business Intelligence
Power BI
Tableau
Apache Superset
Events
Contact Us
Glossary
Subsurface LIVE Winter 2022 sessions are now online!
Nearly 60 live keynotes and breakout sessions.
Watch On Demand Now
Data Lake Storage
Amazon S3
Azure Data Lake Storage
Google Cloud Storage
File Formats
Apache Parquet
Table Formats
Apache Iceberg
Delta Lake
Metastores
AWS Glue
Hive Metastore
Nessie
Data Lake Engines
Dremio
Apache Spark
Interfaces
Apache Arrow Flight
In-Memory Formats
Apache Arrow
Data Catalog
Amundsen
Business Intelligence
Power BI
Tableau
Apache Superset
Search for:
Data Lake Storage
Amazon S3
Azure Data Lake Storage
Google Cloud Storage
File Formats
Apache Parquet
Table Formats
Apache Iceberg
Delta Lake
Metastores
AWS Glue
Hive Metastore
Nessie
Data Lake Engines
Dremio
Apache Spark
Interfaces
Apache Arrow Flight
In-Memory Formats
Apache Arrow
Data Catalog
Amundsen
Business Intelligence
Power BI
Tableau
Apache Superset
Data Lake Storage
File Formats
Table Formats
Metastores
Data Lake Engines
Interfaces
File Formats
A file format is a standard way that information is encoded for storage in a computer file. It specifies how bits are used to encode information in a digital storage medium.
Optimized Row Columnar (ORC)
AVRO
Apache Parquet
Apache Iceberg
October 27, 2022
Puffins and Icebergs: Additional Stats for Apache Iceberg Tables
A short introduction to the new file format called Puffin in Apache Iceberg that helps with additional table statistics
Read more
File Formats
Apache Parquet
March 2, 2022
1 Stone, 3 Birds: Finer – Grained Encryption @ Apache Parquet
Read more
Table Formats
Nessie
Metastores
File Formats
Data Lake Engines
CSV
Apache Spark
Apache Iceberg
September 27, 2021
Project Nessie: Transactional Catalog for Data Lakes with Git-like semantics
Read more
Table Formats
Nessie
Metastores
In-Memory Formats
Apache Parquet
Apache Iceberg
Apache Arrow Flight
Apache Arrow
July 22, 2021
Panel: Open Data Architecture
Read more
Load More