David Morgan-Gumm speaks with Alex Merced, Head of DevRel at Dremio, tech educator, and author, about the open source foundations behind modern data engineering. They break down Apache governance, Apache Arrow, Iceberg, Polaris, and how these pieces fit together in lakehouse architectures. Alex also shares his path into technical writing, community building, and why AI is changing how creators and data teams work.
Key topics
Timestamps
00:00 - Introduction to SQL Squared and guest welcome
03:46 - What Apache means in open source governance
06:06 - How Apache projects become community driven
07:59 - Community over code and why that matters
09:52 - Dremio’s roots in Arrow, Drill, and Calcite
10:21 - Hadoop, MapReduce, and the need for better query engines
11:46 - Why Apache Arrow standardizes in-memory data
13:55 - Where Arrow shows up in modern tools
15:22 - Why Iceberg emerged from Hadoop and Parquet limitations
17:10 - From folders of files to tables with metadata
18:10 - How Hudi, Iceberg, and Delta Lake solve versioning differently
19:07 - Iceberg metadata trees and manifest lists
20:04 - File skipping and why it speeds up analytics
23:26 - Iceberg as the table, Polaris as the catalogue
24:24 - Polaris as vendor-neutral phone book and governance layer
25:47 - Querying the same Iceberg tables from different engines
26:47 - Why lakehouse architecture is future-proof
28:12 - AI makes data silos harder to live with
30:00 - How Dremio helps migrate without a hard cutover
31:28 - Table-by-table migration with views and abstraction
33:47 - Why the role of DevRel matters in data infrastructure
34:32 - Alex’s early tech roots, from GeoCities to RPG Maker
36:27 - Years in finance, training, and building a personal brand
38:20 - Boot camp teaching, video content, and the move into DevRel
41:41 - Why Alex started writing an Iceberg book
43:10 - The timing of the Iceberg book and rising industry interest
45:37 - Tabular, Polaris, and the catalogue gap
48:03 - Snowflake, Databricks, and the open catalogue race
50:15 - Alex’s YouTube channels and weekly newsletters
51:10 - AI, creativity, and the tension for working artists
54:39 - Using AI as a tool, not a shortcut
55:53 - What makes strong technical communities
58:36 - Different community types and how programming should match the audience
62:05 - Closing thoughts on community and open ecosystems
Support the show!
Enjoying The sql_squared Podcast? The best way to support us is by subscribing to our YouTube channel!
The sql_squared podcast is your guide to navigating the ever-evolving world of data. We go beyond the code to explore the tools, techniques, and trends that shape the data landscape, from SQL Server and cloud platforms to AI and developer productivity. Join us as we chat with experts from the community to help you learn, grow, and make the right decisions on your data journey.
sql_squared: Open Data Engineering w/ Alex Merced
1:01:00
Evolution of Database Infrastructure & Chaos Engineering w/ Andrew Pruski - sql_squared
1:08:30
The Future of Data & Building Inclusive Tech Communities w/ James Reeves
56:16
Evolving Roles: A Journey to Data Expert w/Simon Frazer
45:10
Agile and Innovative Data Use in Recruitment w/ Mike Hodson
1:06:59
Mastering Azure Migrations & DevOps w/ Elliot Leighton-Woodruff
1:03:14
sql_squared: Ep. 1 - Alex Dean & Introduction
54:15
sql_squared: Ep. 4 - The Data Community & Database Testing w/ Lee Brownhill
57:13
sql_squared: Ep. 3 - Database DevOps, Blogs & Speaking w/ Tonie Huizer
48:09
sql_squared: Ep. 2 - The usage of data in Web Applications /w Matthew Williams
45:05