Checklist Microsoft Fabric
Fabric Notebook Performance Checklist
A practical pre-run, read, transform, write and orchestration checklist for faster and more reliable Microsoft Fabric notebooks.
- Type
- Checklist
- Level
- Intermediate
- Updated
- Format
- Printable
In short
Identify the true slow stage, minimize unnecessary data and shuffles, choose session reuse deliberately, and validate Delta output before calling a notebook optimized.
Who it is for
- Data engineers operating Fabric notebooks and pipelines
- Teams reviewing a slow or unreliable PySpark workload
What it helps you do
- Separate session startup from Spark execution
- Reduce avoidable I/O, shuffle and repeated work
- Validate orchestration and Delta writes with evidence
Use this checklist with a representative run and comparable input volume. Record session acquisition, major Spark stages, and output writing separately. A shorter total duration is useful only when row counts, business grain, and rerun behavior remain correct.
Before running
- Confirm the notebook has one clear transformation purpose and owner.
- Record the input boundary, expected output grain, and rerun behavior.
- Remove unused imports and avoid unnecessary inline package installation.
- Confirm whether starter, reused high-concurrency, or isolated compute fits the workload.
- Parameterize environment-specific Lakehouses, dates, paths, and correlation IDs.
- Capture a baseline using representative volume before changing configuration.
While reading
- Select only columns required by the transformation and its quality checks.
- Apply selective filters close to the source and confirm them in the executed plan.
- Avoid reading complete history for an incremental run.
- Check partition pruning, Delta data skipping, input file count, and small-file patterns.
- Measure rows or bytes entering the first wide transformation.
While transforming
- Avoid repeated actions such as multiple counts or displays on the same long lineage.
- Do not use
collect()ortoPandas()for results that may exceed driver memory. - Identify shuffle-heavy joins, groups, sorts, distinct operations, and repartitioning.
- Check task duration and partition sizes for skew rather than tuning one global setting blindly.
- Broadcast only measured small reference data.
- Cache only expensive DataFrames reused by multiple actions; materialize and unpersist them deliberately.
While writing
- Choose append, overwrite, or merge from the data contract and rerun requirements.
- Avoid unnecessary repartitioning, especially forcing all output through one partition.
- Monitor file count and write-heavy stages; compact according to downstream read patterns.
- Evaluate Optimize Write, V-Order, Z-Order, and table maintenance by workload and consumer.
- Validate output row counts, keys, partitions, and important measures.
- Confirm a retry cannot duplicate or partially replace the target unexpectedly.
Pipeline and orchestration
- Enable and use session tags for compatible pipeline notebooks when reuse is appropriate.
- Keep stable session tags separate from per-run correlation identifiers.
- Separate incompatible Lakehouses, environments, compute configurations, identities, or resource-heavy workloads.
- Log pipeline activity, notebook, acquisition, first action, stage, and write durations.
- Use Fabric monitoring and session details to identify the true slow layer.
- Change one understood variable, rerun comparable data, and retain correctness controls.
Go deeper: related articles
Microsoft Fabric
How to Make Microsoft Fabric Notebooks Faster
A practical workflow for reducing Fabric notebook startup, read, shuffle, transformation and Delta write time without guessing.
Microsoft Fabric
Using Session Tags in Fabric Pipelines
Use Microsoft Fabric pipeline session tags to group compatible notebook activities for Spark session reuse while preserving workload boundaries.
Microsoft Fabric
How to Reduce Fabric Notebook Startup Time
Diagnose Microsoft Fabric Spark session acquisition and reduce avoidable notebook startup overhead through compatible pools, environments and reuse.
DevOps & Observability
Monitoring Microsoft Fabric Pipelines and Notebooks
Separate Fabric orchestration, session startup, Spark compute, data read, shuffle and write time with connected telemetry.
Planned articles on these topics
Build it: related projects
Related talks
From Notebook to Production: Git and CI/CD in Microsoft Fabric
How a Fabric notebook or pipeline gets from a developer's workspace to production safely: Git integration, dev/test/prod workspaces, deployments, environment configuration and release practices.
Formats: Talk · Workshop · Webinar · Internal session
Medallion Architecture Without Over-Engineering
Bronze, Silver and Gold as responsibilities rather than mandatory boxes: when all three layers help, when fewer are better, and how Delta, incremental processing, file design and data quality fit in.
Formats: Talk · Webinar · Workshop · Internal session
What the community is discussing
- AI-assisted development in FabricMicrosoft Fabric Community
- Delta Lake interoperability, schema and write behaviorReddit
- Moving notebooks between Dev, Test and ProductionReddit