# Alex Merced - 2025 Guide to Architecting an Iceberg Lakehouse (Highlights)

## Metadata
**Review**:: [readwise.io](https://readwise.io/bookreview/61213934)
**Source**:: #from/readwise #from/reader
**Zettel**:: #zettel/fleeting
**Status**:: #x
**Authors**:: [[Alex Merced]]
**Full Title**:: 2025 Guide to Architecting an Iceberg Lakehouse
**Category**:: #articles #readwise/articles
**Category Icon**:: 📰
**Document Tags**:: #favorite #sok
**URL**:: [iceberglakehouse.com](https://iceberglakehouse.com/posts/2024-12-2025-guide-architecting-an-iceberg-lakehouse/)
**Host**:: [[iceberglakehouse.com]]
**Highlighted**:: [[2026-06-10]]
**Created**:: [[2026-06-13]]
## Note
- Catalog: Nessie, AWS Glue, Hive, Polaris, Lakekeeper, Gravitino, Dremio Catalog, Snowflake Open Catalog
- Intestion: Apache Spark, Flink, Kafka Connect
- Integration: Dremio, Trino, Presto
- Consumption: Tableau, Power BI, dbt
- Storage: AWS S3, Azure Data Lake Storage, Google Cloud Storage
## Highlights
- Before we begin architecting your Apache Iceberg Lakehouse, it’s essential to perform a self-audit to clearly define your requirements. ([View Highlight](https://read.readwise.io/read/01ktssrv3cte3fcjjwa94wtg77)) ^1024217562
↩︎
- Where is my data currently?
- Which of my data is the most accessed by different teams?
- Which of my data is the highest cost generator?
- Which data platforms will I still need if I standardize on Iceberg?
- What are the SLAs I need to meet?
- What tools are accessing my data, and which of those are non-negotiables?
- What are my regulatory barriers?
- your data will be stored as **Parquet files** with **Iceberg metadata** ([View Highlight](https://read.readwise.io/read/01ktstete1wp026wvkhsrgn8jd)) ^1024219014
- **VAST Data**: Designed for high-performance workloads, leveraging technologies like NVMe over Fabrics. ([View Highlight](https://read.readwise.io/read/01ktstwjg5vkqhrb0pcznk78p9)) ^1024220326
- **MinIO**: An open-source, high-performance object storage system compatible with S3 APIs, ideal for hybrid environments. ([View Highlight](https://read.readwise.io/read/01ktstwgf02mbqjdpca5qypg5e)) ^1024220324
- **Apache Spark**: Ideal for large-scale batch processing and ETL workflows. ([View Highlight](https://read.readwise.io/read/01ktsv3a1bbf8h75sa7vg5d83f)) ^1024220552
- **Apache Kafka** or **Apache Flink**: Excellent choices for real-time streaming data ingestion. ([View Highlight](https://read.readwise.io/read/01ktsv37yp7e6525z8b1wx2ngb)) ^1024220550
- **Batch Ingestion Tools**:
Examples include **Fivetran**, **Airbyte**, **AWS Glue**, and **ETleap**. ([View Highlight](https://read.readwise.io/read/01ktsv50w4be22rf5jyfk1h0ea)) ^1024220601
- **Streaming Ingestion Tools**:
Examples include **Upsolver**, **Delta Stream**, **Estuary**, **Confluent**, and **Decodable** ([View Highlight](https://read.readwise.io/read/01ktsv55k5yts08zq2bjf2hg35)) ^1024220607
- Dremio allows you to connect and query all your data sources in one place. Even if your datasets haven’t yet migrated to Iceberg, you can combine them with Iceberg tables seamlessly. ([View Highlight](https://read.readwise.io/read/01ktsv92wx3xpxf6vjp2rt9jdf)) ^1024220752