Azure Databricks Data Engineering: Build a Lakehouse
Master PySpark, Delta Lake and Native Dashboards by building a Real Estate Market Tracker from scratch in Azure Databricks
★★★★★ 4.7
(4 reviews)
2 hours of content
What you'll learn
- Azure Infrastructure. Deploy Data Lakes (Gen2), Key Vaults, and Databricks Workspaces using the Azure Portal.
- The Medallion Architecture. Architect a professional Bronze (Raw), Silver (Clean), and Gold (Aggregated) data flow.
- PySpark & SQL Mastery. Write robust transformations to handle messy JSON data, enforce schemas, and deduplicate records.
- Delta Lake Internals. Master "Time Travel," ACID transactions, and Schema Enforcement to treat files like database tables.
- Orchestration. Replace manual runs with automated Databricks Workflows (Jobs) that run on a Cron schedule.
- Native BI. Build stunning, auto-updating dashboards directly inside Databricks using SQL visualizations.
Covers building a lakehouse on Azure Databricks with a medallion architecture, using PySpark and SQL to clean and aggregate data. Learners configure Data Lake Gen2, Key Vault, and Databricks workspaces, work with Delta Lake time travel and ACID transactions, and automate workflows with Databricks Jobs. They also create native SQL dashboards for business intelligence.
Who this course is for
Ideal for data engineers and analysts with basic Python and SQL knowledge who want to build production-style Azure Databricks pipelines.
#azure-databricks
#pyspark
#delta-lake
#data-engineering
#sql
#medallion-architecture