Course Outline
Core Databricks Platform and Lakehouse Concepts
- Overview of Databricks Lakehouse architecture and key components
- Managing workspaces and data catalogs
Navigating the Databricks Workspace and Notebooks
- Workspace navigation and development via notebooks
- Organizing code into reusable notebook structures
Apache Spark Architecture and Execution Models
- Spark runtime architecture and execution mechanisms
- Lazy evaluation and Directed Acyclic Graphs (DAG) for jobs
PySpark DataFrames and API Usage
- Understanding DataFrame abstractions and schemas
- Essential DataFrame operations and column expressions
Converting SQL to PySpark DataFrames
- Mapping core SQL clauses to DataFrame methods
- Utilizing window functions and aggregations in PySpark
Data Ingestion and Output in Databricks
- Accessing data from standard file and database sources
- Writing and partitioning data within the Lakehouse
Delta Lake and Table Administration
- Delta tables and ACID transaction support
- Time travel features and schema evolution
Data Cleaning and Transformation Techniques
- Data cleansing and data type conversions
- Creating reusable transformation logic
Custom Functions and Code Modularity
- Implementing Python UDFs and pandas UDFs
- Encapsulating procedural logic into functions
Performance Optimization and Tuning
- Strategies for partitioning and caching
- Identifying bottlenecks using the Spark UI
Basics of Structured Streaming
- Differences between batch and streaming processing
- Streaming DataFrames and fundamental aggregations
Job Scheduling and Workflow Management in Databricks
- Scheduling notebooks as automated jobs and tasks
- Designing multi-stage workflows with dependencies
Unity Catalog and Governance Frameworks
- Unity Catalog structure and namespace management
- Access controls and data lineage tracking
Testing, Debugging, and Production Best Practices
- Unit testing for PySpark logic
- Debugging techniques and code quality standards
Complete Financial Services Implementations
- Developing end-to-end banking ETL pipelines
- Translating legacy SQL processes into PySpark
Transitioning SQL Workloads to PySpark
- Migration strategies and planning frameworks
- Step-by-step conversion of SQL workflows to PySpark
Requirements
- Proficiency in Python programming, covering functions and data types.
- Solid knowledge of SQL, including joins, aggregations, and subqueries.
- No previous experience with Databricks or PySpark is necessary.
Target Audience
- Data engineers, data analysts, and general data professionals.
- Teams transitioning existing SQL-based workflows to Databricks and PySpark.
Testimonials (1)
I liked that it was practical. Loved to apply the theoretical knowledge with practical examples.