Architecting a Modern Data Lake with Dipti Borkar from Ahana
Listen now
Description
In this episode of Building The Backend we hear from Dipti Borkar cofounder @ Ahana  a managed service for Presto on AWS, where we talk all about the data lake, how it should be structured and where the industry is going.   Below are top 3 value bombs:  Presto is an open source distributed SQL query engine originally created by Facebook, mainly used to run SQL queries on data lakes but can be connected to relational data stores as well. Ahana is a managed Presto service on AWS with 3x price/performance. When optimizing your data lake, it’s normally best to store the data in Parquet or ORC format vs JSON or CSV as they are columnar formats that can have indexes built in. Data Lake Houses are continuing to gain popularity by bringing the benefits of your data lake and data warehouse together with the help of tools like  Databricks DeltaLake and Apache HUDI.
More Episodes
In this episode we speak with Justin Borgman, Chairman & CEO at Starburst, which is based on open source Trino (formerly PrestoSQL) and was recently valued at $3.35 billion after securing their series D funding.  In this episode we discuss convergence of DW’s / DL's, why data lakes fail and...
Published 03/15/22
In this episode we speak with Paul Singman Developer Advocate at Treeverse / LakeFS. LakeFS is an open source project  that allows you to transform your object storage into a Git-like repository.  Top 3 takeaways LakeFS enables use cases like debugging to quickly view historical versions of your...
Published 03/01/22