databricks blk
410  Reviews star_rate star_rate star_rate star_rate star_outline

Data Engineering with Databricks

Data professionals from all walks of life will benefit from this comprehensive introduction to the components of the Databricks Lakehouse Platform that directly support putting ETL pipelines into...

Read More
$1,500 USD GSA  $1,360.20
Course Code DEWD
Duration 2 days
Available Formats Classroom, Virtual

Data professionals from all walks of life will benefit from this comprehensive introduction to the components of the Databricks Lakehouse Platform that directly support putting ETL pipelines into production. You will leverage SQL and Python to define and schedule pipelines that incrementally process new data from a variety of data sources to power analytic applications and dashboards in the Lakehouse. This course offers hands-on instruction in Databricks Data Science & Engineering Workspace, Databricks SQL, Delta Live Tables, Databricks Repos, Databricks Task Orchestration, and the Unity Catalog.

  • This course will prepare you to take the Databricks Certified Data Engineer Associate exam.

Skills Gained

  • Leverage the Databricks Lakehouse Platform to perform core responsibilities for data pipeline development
  • Use SQL and Python to write production data pipelines to extract, transform, and load data into tables and views in the Lakehouse
  • Simplify data ingestion and incremental change propagation using Databricks-native features and syntax, including Delta Live Tables
  • Orchestrate production pipelines to deliver fresh results for ad-hoc analytics and dashboarding

Prerequisites

  • Basic knowledge of SQL query syntax, including writing queries using SELECT, WHERE, GROUP BY, ORDER BY, LIMIT, and JOIN
  • Basic knowledge of SQL DDL statements to create, alter, and drop databases and tables
  • Basic knowledge of SQL DML statements, including DELETE, INSERT, UPDATE, and MERGE
  • Experience with or knowledge of data engineering practices on cloud platforms, including cloud features such as virtual machines, object storage, identity management, and metastores
  • Basic familiarity with Python variables, functions, and control flow (preferred)

Course Details

Course Outline

Day 1

  • Delta Lake
  • Relational entities on Databricks
  • ETL with Spark SQL
  • Incremental data processing with Structured Streaming and Auto Loader

Day 2

  • Medallion architecture in the data lakehouse
  • Delta Live Tables
  • Task orchestration with Databricks Jobs
  • Databricks SQL
  • Managing Permissions in the lakehouse
  • Productionizing dashboards and queries on Databricks SQL
|
View Full Schedule