Inside the books
Two volumes. 37 chapters. ~1600 pages of practitioner depth.
Every chapter is opinionated, hands-on, and written for engineers who already know the basics. No "what is Spark." Yes "here is the exact pattern that works at scale, here is the failure mode, here is what it costs."
Volume 3 · Ch 22–37 · ~800 pages
The Production Lakehouse Playbook
- Part I (Ch 22–24)
Platform Foundations
From OSS Spark to the Databricks Data Intelligence Platform. Workspaces, notebooks, Git folders, the Assistant. Compute: classic, serverless, SQL warehouses, AI Runtime, GPUs.
- Part II (Ch 25–30)
Unity Catalog & Governance
UC architecture, the three-level namespace, volumes. Access control with ABAC and governed tags. Identity, service principals, OAuth, secret scopes. Managed tables with Delta + Iceberg/UniForm. Liquid Clustering + Predictive Optimization. System tables and HMS → UC migration.
- Part III (Ch 31–36)
Data Engineering at Scale
Ingestion with Auto Loader and Lakeflow Connect. Lakeflow SDP (Spark Declarative Pipelines). Lakeflow Jobs orchestration. Declarative Automation Bundles, CLI, SDK, Terraform. CI/CD with GitHub Actions + OIDC. Performance tuning with Photon.
- Part IV (Ch 37)
Bridge into Volume 4
UC Metrics and the data-product mindset that makes the AI layer in Volume 4 trustworthy.
Volume 4 · Ch 38–58 · ~800 pages
The AI Lakehouse Playbook
- Part V (Ch 38–42)
Analytics & Natural-Language BI
Databricks SQL in production. External BI (Tableau, Power BI, Looker, dbt). AI/BI Dashboards. AI/BI Genie. The AI SQL function family, ai_query, ai_parse_document, ai_classify, ai_extract, ai_gen, ai_forecast, vector_search.
- Part VI (Ch 43–46)
Mosaic AI: Models & Retrieval
Model Serving and the AI Gateway. Foundation Model APIs and external models. Vector Search and RAG patterns. MLflow 3, experiments, UC Model Registry, traces, evaluation.
- Part VII (Ch 47–49)
The ML Lifecycle
Feature Store on Unity Catalog. MLOps, promotion as alias movement, champion/challenger, traffic splitting. Lakehouse Monitoring for data and model drift.
- Part VIII (Ch 50)
Distributed Training
TorchDistributor, DeepSpeed, Ray on Databricks, and serverless GPU for fine-tuning.
- Part IX (Ch 51–52)
Agentic AI
Agent Bricks, classification and information extraction. The Multi-Agent Supervisor with MCP, agent evaluation, and deployment.
- Part X (Ch 53–54)
Operational Data + Capstone
Lakebase, managed Postgres in the lakehouse. End-to-end retail-intelligence capstone that pulls every V3 and V4 piece together.
- Part XI (Ch 55)
Close
Certification paths (Data Engineer Associate/Professional, ML Associate/Professional) and what comes next on the platform.
Free sample topics
Long-form practitioner previews of the book's content are published as free topic guides: Unity Catalog, Lakeflow SDP, Asset Bundles, and CI/CD on Databricks from Volume 3. RAG on Databricks, Agent Bricks, Multi-Agent Systems, and MLflow 3 from Volume 4.
Two volumes
Get the series.
Volume 3 is the platform foundation. Volume 4 is the AI build-out. Read either standalone, or both for the complete picture.
Volume 3
The Production Lakehouse Playbook
Unity Catalog, Lakeflow, and the Databricks Data Intelligence Platform, the production playbook for engineers who already know Spark.
Buy on Amazon$29.00 Kindle · $39.99 Paperback · 16 chapters
Volume 4
The AI Lakehouse Playbook
Mosaic AI, Agent Bricks, Lakebase, and the production Lakehouse, the 2026 field manual for shipping AI systems on Databricks.
Buy on Amazon$32.00 Kindle · $39.99 Paperback · 21 chapters