Microsoft Fabric
Fabric Data Platform Template
A reusable end-to-end Microsoft Fabric reference implementation: ingestion, Bronze, Silver and Gold layers, serving, Git, CI/CD and observability, built with practical engineering patterns.
- Status
- Planned
- Level
- Intermediate
Planned project: this page describes the intended design. No code, repository or demo has been published yet.
Project overview
Who this is for
- Students who want to build a complete Fabric platform step by step, not just isolated demos
- Data engineers designing or reviewing a Fabric platform for their organisation
- Architects who need a reference for environments, deployment and operational boundaries
What you will learn
- How workspaces, Lakehouses, Warehouses, notebooks and pipelines fit together in one platform
- How to implement Bronze, Silver and Gold layers with clear responsibilities and quality gates
- How to load data incrementally and keep pipelines idempotent
- How to choose partitioning and file-size settings for Delta tables
- How to log runs and debug failures across layers
- How to work with Git integration and promote changes from dev to test to production
Tech stack
- Microsoft Fabric
- Fabric Lakehouse
- Fabric Warehouse
- Fabric Notebooks
- Fabric Data Pipelines
- SQL analytics endpoint
- Delta Lake
- Parquet
- PySpark
- Spark SQL
- Power BI
- Git
- GitHub
- Bicep
On this page
The problem
Most Microsoft Fabric examples show one item at a time: a notebook that reads a CSV, a pipeline that copies a table, a report on top of a Lakehouse. Real platforms are harder because the pieces have to work together. Layers need owners and contracts. Loads must be incremental and safe to rerun. Changes must move between environments without manual edits. And someone has to be able to tell, at 7 a.m., which step failed and why.
This project is a single, coherent reference implementation of those concerns. Students can follow it from an empty workspace to a working platform, and professionals can reuse it as a set of patterns for their own environments.
Architecture
Sources
- Sample operational database
- CSV and Parquet files
- REST API
Ingestion
- Data pipelinesCopy and orchestration
- NotebooksAPI and file ingestion
Bronze
- Bronze LakehouseRaw Delta tables + landed files, with batch metadata
Silver
- Silver LakehouseValidated, deduplicated, conformed entities
Gold
- Gold WarehouseDimensional model and serving tables
Serving
- SQL analytics endpoint
- Semantic model
- Reports
- API
Cross-cutting concerns
- Git integration
- CI/CD
- Run logging
- Monitoring
- Configuration
- Workspace security
The layers follow the pattern described in Medallion Architecture Explained. Each layer is a separate item with its own permissions, so the boundary between raw and curated data is enforced by the platform, not only by naming. Gold is a Warehouse in the template because its consumers expect T-SQL and multi-table transactions. A Gold Lakehouse is documented as an alternative.
How it works
- TriggerSchedule or manual run
- Ingest batchNew data since the watermark
- Bronze appendBatch id, load time, source
- Silver MERGEBy business key, with quarantine
- Gold refreshOnly affected facts and aggregates
- Quality gateRow counts, keys, reconciliation
A run is driven by one orchestration pipeline with a run identifier. Ingestion lands new source data in Bronze with the run and batch identifiers attached. Silver notebooks read only the new Bronze batches, validate them, quarantine failing records, and merge the rest by business key. Gold refreshes only the facts and aggregates affected by the changed keys. A final quality gate compares row counts and key totals between layers before the run is marked successful.
Configuration such as connection names, workspace ids and source lists lives in parameters and a configuration table, not in notebook code. This is what lets the same code run in dev, test and production.
Implementation walkthrough
The template will be built and published in milestones. Each one ends with something that runs:
- Workspaces and Git: dev, test and production workspaces, naming conventions, Git integration on dev.
- Bronze ingestion: a pipeline and a notebook landing three sample sources with ingestion metadata.
- Silver transformations: validation rules, deduplication, quarantine table and MERGE by business key.
- Gold model: a small dimensional model in the Warehouse and a semantic model on top.
- Incremental runs: watermarks, idempotent reruns and late-arriving data.
- Observability: a run log table, failure alerts and a simple operational report.
- Deployment: promotion from dev to test to production, with environment-specific configuration.
- Performance: partitioning and file-size choices, compaction and table maintenance at larger volumes.
- Feature branchChange in a dev workspace
- Pull requestReview and checks
- DevGit-connected workspace
- TestDeployed, with test data
- ProductionDeployed, approved release
Student path
For students
Prerequisites
- Access to a Microsoft Fabric capacity or trial, and permission to create workspaces
- Basic SQL; some Python helps but is not required at the start
- A GitHub account for the Git milestones
Concepts to understand first
Read Medallion Architecture Explained before milestone 2. The Fabric learning path introduces workspaces, Lakehouses, notebooks and Delta tables in the order this project uses them.
Guided steps
Follow the milestones in order. Each milestone will have its own instructions, a checklist of what should exist when you finish, and a short “how do I know it worked” section.
Exercises
- Add a fourth source to Bronze without changing the Silver code.
- Write a validation rule that quarantines orders with a missing customer, and check the quarantine table.
- Rerun the same batch twice and prove that Silver and Gold do not change.
Try this next
Replace the Gold Warehouse with a Gold Lakehouse and compare the developer experience, then read the planned Lakehouse vs Warehouse material.
Expected outcomes
You will be able to explain what each Fabric item in the platform does, build and rerun a layered pipeline, and use Git to move a change between environments.
Professional considerations
For professionals
Architecture decisions
One Lakehouse per layer keeps permissions and ownership explicit; a single Lakehouse with schemas is simpler for small teams. The template documents both and uses the first. Gold in a Warehouse favours T-SQL consumers; Gold in a Lakehouse favours Spark-heavy teams.
Environments and deployment
Dev, test and production are separate workspaces, ideally on separate capacities for production isolation. Only dev is edited directly. Environment differences live in configuration, so deployment never requires changing code.
Scale
Bronze tables are partitioned by load date only when volumes justify it; Silver and Gold start unpartitioned. Incremental processing, compaction and table maintenance are part of the design rather than added after the first slow run.
Security
Workspace roles separate engineers from consumers. Consumers read Gold through the SQL analytics endpoint or the semantic model and never get access to Bronze. Secrets are not stored in notebooks.
Observability and failure modes
Every run writes a row per step to a run log: run id, batch id, rows read and written, duration and outcome. This answers which batch failed, in which layer, and what changed. Failure modes covered include source unavailability, schema drift, duplicate deliveries, partial runs and reruns.
Cost and maintainability
Incremental processing keeps compute proportional to change rather than to data size. Shared notebook utilities for logging, configuration and MERGE keep layer code short and consistent.
Alternatives considered
- A single all-in-one notebook per source: simple at first, but hard to test and impossible to rerun per layer.
- Full reloads everywhere: correct and easy, but cost and duration grow with data size.
- Fewer layers: valid for simple workloads; the template uses three because it is a reference for multi-source platforms.
Common mistakes
- Building all three layers before the first end-to-end run works with one source.
- Hard-coding workspace ids or connection names, which breaks every deployment.
- Merging into Silver without deduplicating the incoming batch first.
- Letting reports read Silver “temporarily”, which makes it a contract nobody planned.
- Logging only failures. Successful runs need row counts too, or there is nothing to compare against.
Next improvements
- Publish the repository and milestone 1.
- Add a sample dataset large enough to show partitioning and compaction effects.
- Add a workshop version of the student path with timings and checkpoints.
Repository, demo and downloads
Nothing has been published yet. The repository, demo and downloads will be linked here when they exist.
Have feedback or want to collaborate? Contact me
