Fabric Start Here2 prévus

  • PrévuMicrosoft FabricDébutant

    Create Your First Microsoft Fabric Workspace

    Create a development workspace with sensible capacity, access and ownership decisions from the start.

    Sujets prévus

    Create a development workspace with sensible capacity, access and ownership decisions from the start.

    • creating a workspace
    • capacity basics
    • permissions and roles
    • what belongs in a development workspace
    • Microsoft Fabric
    • Fabric Workspace
  • PrévuDevOps & ObservabilityDébutant

    Connect Microsoft Fabric to Git

    Connect a Fabric workspace to a repository and understand branches, synchronization, versioned artifacts and current limitations.

    Sujets prévus

    Connect a Fabric workspace to a repository and understand branches, synchronization, versioned artifacts and current limitations.

    • repository structure and branches
    • connecting and synchronizing a workspace
    • what Fabric versions in Git
    • common integration limitations
    • Microsoft Fabric
    • Git
    • DevOps

Fabric Foundations1 prévus

  • PrévuMicrosoft FabricDébutant

    What Is Microsoft Fabric? A Practical Overview

    Understand how OneLake, workspaces, Lakehouses, Warehouses, notebooks, pipelines, SQL endpoints and semantic models fit together.

    Sujets prévus

    Understand how OneLake, workspaces, Lakehouses, Warehouses, notebooks, pipelines, SQL endpoints and semantic models fit together.

    • OneLake and the Fabric workspace model
    • Lakehouse, Warehouse, notebooks and pipelines
    • SQL analytics endpoints and semantic models
    • how the pieces form an end-to-end platform
    • Microsoft Fabric
    • Data Architecture

Fabric Notebooks3 prévus

  • PrévuMicrosoft FabricDébutant

    Create Your First Fabric Notebook

    Use PySpark to load, inspect and summarize a dataset in your first Fabric notebook.

    Sujets prévus

    Use PySpark to load, inspect and summarize a dataset in your first Fabric notebook.

    • creating and attaching a notebook
    • reading CSV data with PySpark
    • displaying rows and schema
    • counting and grouping records
    • Microsoft Fabric
    • Fabric Notebook
    • PySpark
    • Data Ingestion
  • PrévuMicrosoft FabricDébutant

    Explore Data with PySpark in Microsoft Fabric

    Practice the DataFrame operations used to explore and shape data in Fabric notebooks.

    Sujets prévus

    Practice the DataFrame operations used to explore and shape data in Fabric notebooks.

    • select and filter
    • groupBy and agg
    • orderBy and join
    • using display for rapid exploration
    • Microsoft Fabric
    • Fabric Notebook
    • PySpark
  • PrévuMicrosoft FabricDébutant

    Use SQL Inside a Fabric Notebook

    Query notebook data with Spark SQL and decide when SQL or PySpark is the clearer tool.

    Sujets prévus

    Query notebook data with Spark SQL and decide when SQL or PySpark is the clearer tool.

    • creating temporary views
    • running Spark SQL queries
    • switching between SQL and DataFrames
    • choosing PySpark versus Spark SQL
    • Microsoft Fabric
    • Fabric Notebook
    • Spark SQL
    • PySpark

Fabric Lakehouse2 prévus

  • PrévuMicrosoft FabricDébutant

    Create Your First Fabric Lakehouse

    Build a Lakehouse and understand its Files, Tables and SQL analytics endpoint surfaces.

    Sujets prévus

    Build a Lakehouse and understand its Files, Tables and SQL analytics endpoint surfaces.

    • creating the Lakehouse
    • Files versus Tables
    • Delta table basics
    • querying through the SQL analytics endpoint
    • Microsoft Fabric
    • Fabric Lakehouse
    • SQL Endpoint
    • Delta Lake
  • PrévuDelta LakeDébutant

    Write Your First Delta Table in Fabric

    Persist a DataFrame as a managed Delta table and query it from Spark and SQL.

    Sujets prévus

    Persist a DataFrame as a managed Delta table and query it from Spark and SQL.

    • writing a managed Delta table
    • overwrite and append modes
    • querying the saved table
    • where the table appears in the Lakehouse
    • Microsoft Fabric
    • Fabric Lakehouse
    • Delta Lake

Fabric Medallion Architecture9 prévus

  • PrévuMicrosoft FabricIntermédiaire

    Designing a Gold Layer for Reporting and APIs

    How to design the Gold layer that reports, semantic models and APIs consume: modelling choices, stable contracts, serving engines and how to change it without breaking consumers.

    Sujets prévus

    The curated layer is the contract with consumers. Its design determines how easy it is to change the platform underneath.

    • what belongs in the curated layer
    • dimensional vs wide tables
    • naming and data types as a contract
    • versioning changes
    • serving through SQL endpoint or Warehouse
    • data quality checks
    • Gold
    • Microsoft Fabric
    • Data Architecture
    • Medallion Architecture
  • PrévuMicrosoft FabricIntermédiaire

    When You Should NOT Use All Three Medallion Layers

    A critical look at the Medallion pattern: when two layers or a single curated model are enough, and the cost that each unnecessary layer adds.

    Sujets prévus

    The Medallion pattern is often applied by default, producing extra copies of data without a clear purpose. This article looks at concrete cases where fewer layers are the better design, and at the signals that justify adding one.

    • what each layer must provide to justify itself
    • Bronze + Gold designs
    • Silver + Gold designs
    • a single curated Lakehouse or Warehouse
    • storage, compute and operational cost per layer
    • migrating from three layers to fewer, and back
    • Medallion Architecture
    • Data Architecture
    • Microsoft Fabric
  • PrévuMicrosoft FabricIntermédiaire

    Designing a Bronze Layer in Microsoft Fabric

    How to design a Bronze layer that preserves source fidelity: landing formats, ingestion metadata, schema drift, retention and replay.

    Sujets prévus

    Bronze is only useful if it can answer what a source actually sent and let a batch be replayed. That depends on decisions made at ingestion time that are hard to add later.

    • files vs Delta tables vs both
    • ingestion metadata columns
    • tolerating schema drift
    • append-only design and batch identifiers
    • retention and storage cost
    • replaying a batch into Silver
    • Medallion Architecture
    • Bronze
    • Microsoft Fabric
    • Data Ingestion
  • PrévuMicrosoft FabricIntermédiaire

    Designing a Silver Layer for Clean and Conformed Data

    How to turn source-shaped Bronze data into clean, validated, conformed Silver entities that several consumers can share.

    Sujets prévus

    Silver is where most data quality work happens and where business keys are established. A weak Silver layer pushes cleanup into every Gold model.

    • cleaning and standardization rules
    • choosing and enforcing business keys
    • deduplication strategies
    • conformed entities across sources
    • quarantine tables for failed records
    • testing Silver transformations
    • Medallion Architecture
    • Silver
    • Microsoft Fabric
    • Delta Lake
  • PrévuMicrosoft FabricAvancé

    Incremental Processing Across Bronze, Silver and Gold

    How changes propagate from Bronze to Silver to Gold without full reloads: change detection between layers, affected keys and recomputing only impacted aggregates.

    Sujets prévus

    Loading new data from a source is only the first step. Each later layer also needs to process only what changed, or the platform still pays for a full rebuild downstream.

    • layer-to-layer change detection
    • watermarks vs Delta change data feed
    • identifying affected business keys
    • recomputing affected Gold aggregates
    • late and corrected data across layers
    • reconciliation runs
    • Medallion Architecture
    • Incremental Loads
    • Delta Lake
    • Microsoft Fabric
  • PrévuMicrosoft FabricAvancé

    Delta MERGE Patterns in a Medallion Architecture

    Which write pattern belongs in which layer: append in Bronze, MERGE for Silver entities and history, and targeted overwrite or MERGE for Gold.

    Sujets prévus

    MERGE is powerful but expensive, and it is often used in layers where an append or a partition overwrite would be simpler and cheaper.

    • write patterns by layer
    • Silver upserts by business key
    • slowly changing dimensions in Silver
    • Gold refresh: MERGE vs replaceWhere
    • keeping MERGE idempotent
    • when MERGE becomes the bottleneck
    • Medallion Architecture
    • MERGE
    • Delta Lake
    • Upsert
  • PrévuMicrosoft FabricAvancé

    Partitioning Strategies Across Medallion Layers

    Why the right partitioning often differs by layer, from load-date partitions in Bronze to unpartitioned Gold tables, and how to decide.

    Sujets prévus

    Copying one partitioning scheme into every layer is a common source of small files and missed pruning. Each layer is written and read differently.

    • load date vs business date
    • partitioning Bronze for ingestion and replay
    • large Silver facts
    • why most Gold tables need no partitions
    • clustering as an alternative
    • checking that pruning actually happens
    • Medallion Architecture
    • Partitioning
    • Delta Lake
    • Performance
  • PrévuMicrosoft FabricIntermédiaire

    Small Files and Compaction in Medallion Architectures

    Where small files come from in each layer and how to plan OPTIMIZE, V-Order and VACUUM per layer instead of applying one schedule everywhere.

    Sujets prévus

    Frequent Bronze appends, Silver MERGE operations and Gold rebuilds create small files for different reasons, so the maintenance plan should differ too.

    • sources of small files by layer
    • measuring file counts and sizes
    • OPTIMIZE and V-Order in Fabric
    • VACUUM and retention by layer
    • scheduling maintenance around loads
    • avoiding conflicts with writers
    • Medallion Architecture
    • File Sizing
    • Table Maintenance
    • Delta Lake
  • PrévuMicrosoft FabricAvancé

    Building an End-to-End Fabric Medallion Pipeline

    A complete, production-minded Medallion pipeline in Fabric: incremental ingestion, quality gates, MERGE, maintenance, logging and deployment. It builds on the first end-to-end pipeline tutorial.

    Sujets prévus

    The first end-to-end pipeline tutorial shows the flow; this capstone adds what a production pipeline needs, bringing the rest of the series together in one working design.

    • architecture and workspace layout
    • incremental Bronze ingestion
    • Silver validation and quarantine
    • Gold models and refresh
    • quality gates between layers
    • run logging and alerting
    • table maintenance
    • deploying with Git and pipelines
    • Medallion Architecture
    • Microsoft Fabric
    • Data Pipelines
    • Incremental Loads

Fabric Warehouse1 prévus

  • PrévuMicrosoft FabricIntermédiaire

    Lakehouse vs Warehouse in Microsoft Fabric: How I Decide

    The criteria I use to choose between a Fabric Lakehouse and a Fabric Warehouse: team skills, write patterns, T-SQL needs and governance.

    Sujets prévus

    Both store Delta tables in OneLake, which makes the choice confusing. The difference is mostly in how data is written and who works with it.

    • what each item is underneath
    • write paths: Spark vs T-SQL
    • transactions and DML support
    • security model
    • team skills
    • using both together
    • Microsoft Fabric
    • Fabric Lakehouse
    • Fabric Warehouse
    • Data Architecture

Fabric Pipelines2 prévus

  • PrévuMicrosoft FabricIntermédiaire

    Build Your First End-to-End Fabric Pipeline

    Move CSV or Parquet through a notebook-driven Bronze, Silver and Gold flow into a queryable serving layer.

    Sujets prévus

    Move CSV or Parquet through a notebook-driven Bronze, Silver and Gold flow into a queryable serving layer.

    • ingesting CSV and Parquet
    • orchestrating a notebook
    • Bronze, Silver and Gold responsibilities
    • serving through Warehouse or SQL endpoint
    • Microsoft Fabric
    • Data Pipelines
    • Medallion Architecture
    • Data Ingestion
  • PrévuMicrosoft FabricIntermédiaire

    Make Your Fabric Notebook Reusable with Parameters

    Pass dates and business identifiers from a pipeline into one reusable Fabric notebook.

    Sujets prévus

    Pass dates and business identifiers from a pipeline into one reusable Fabric notebook.

    • declaring notebook parameters
    • passing values from a pipeline
    • validating dates and identifiers
    • using parameters safely with Delta tables
    • Microsoft Fabric
    • Fabric Notebook
    • Data Pipelines

Fabric Git & CI/CD4 prévus

  • PrévuDevOps & ObservabilityIntermédiaire

    CI/CD for Microsoft Fabric: A Practical Deployment Strategy

    A practical deployment approach for Fabric items using Git integration, deployment pipelines and automation, including what still needs manual steps.

    Sujets prévus

    Fabric deployment options are changing quickly and each has gaps. The article describes a strategy that works today and where it needs workarounds.

    • Git integration
    • deployment pipelines
    • APIs and automation tools
    • parameterising environment differences
    • what cannot be deployed automatically yet
    • release process
    • CI/CD
    • Microsoft Fabric
    • DevOps
  • PrévuDevOps & ObservabilityIntermédiaire

    Designing Dev, Test and Production Deployments for Fabric

    How to promote Fabric items between environments: configuration per environment, data separation, approvals and rollback.

    Sujets prévus

    Promoting changes through environments is where most deployment problems appear: connection strings, IDs and data that differ per environment.

    • environment-specific configuration
    • connections and IDs
    • test data
    • approvals and gates
    • rollback options
    • keeping environments in sync
    • CI/CD
    • Microsoft Fabric
    • DevOps
  • PrévuDevOps & ObservabilityIntermédiaire

    Git Integration in Microsoft Fabric

    Use Fabric Git integration deliberately across branches, workspace artifacts and team workflows.

    Sujets prévus

    Use Fabric Git integration deliberately across branches, workspace artifacts and team workflows.

    • supported workspace artifacts
    • branching and collaboration flow
    • workspace synchronization
    • limitations and conflict handling
    • Microsoft Fabric
    • Git
    • DevOps
  • PrévuDevOps & ObservabilityIntermédiaire

    What Actually Gets Deployed in Fabric?

    Separate deployable workspace artifacts from data, connections and environment-specific configuration.

    Sujets prévus

    Separate deployable workspace artifacts from data, connections and environment-specific configuration.

    • notebooks and pipelines
    • Git-managed workspace content
    • data that is not deployed
    • environment differences and deployment limitations
    • Microsoft Fabric
    • CI/CD
    • DevOps

Fabric Production Engineering3 prévus

  • PrévuMicrosoft FabricIntermédiaire

    Incremental Loads in Microsoft Fabric Without Reprocessing Everything

    Patterns for loading only new and changed data into Fabric, with watermarks, change tracking and Delta MERGE, and how to handle late or corrected data.

    Sujets prévus

    Full reloads stop being viable as data grows. Incremental loads are more efficient but bring their own correctness problems.

    • watermark-based loads
    • change tracking and CDC sources
    • Delta MERGE for changes
    • late-arriving data
    • handling deletes
    • periodic reconciliation
    • Microsoft Fabric
    • Incremental Loads
    • Data Ingestion
    • Delta Lake
  • PrévuDevOps & ObservabilityIntermédiaire

    Logging Your Fabric Notebooks Properly

    Design an execution log that makes notebook and pipeline runs traceable in production.

    Sujets prévus

    Design an execution log that makes notebook and pipeline runs traceable in production.

    • execution and pipeline identifiers
    • start, finish and status fields
    • rows processed and error messages
    • writing reliable operational records
    • Microsoft Fabric
    • Fabric Notebook
    • Logging
    • Observability
  • PrévuDevOps & ObservabilityIntermédiaire

    Debugging Fabric Notebooks and Pipelines

    Use a repeatable workflow to diagnose Spark, schema, parameter, dependency and pipeline failures.

    Sujets prévus

    Use a repeatable workflow to diagnose Spark, schema, parameter, dependency and pipeline failures.

    • Spark exceptions and unresolved columns
    • schema mismatches and Delta concurrency
    • parameter and dependency failures
    • a practical troubleshooting sequence
    • Microsoft Fabric
    • Fabric Notebook
    • Debugging
    • Data Pipelines

Fabric Performance at Scale7 prévus

  • PrévuMicrosoft FabricAvancé

    Fabric Performance Troubleshooting: A Practical Workflow

    A step-by-step workflow for slow Fabric workloads: deciding whether the problem is capacity, Spark, SQL, storage layout or the query itself.

    Sujets prévus

    Fabric has several engines and a shared capacity, so 'it is slow' can have many causes. A workflow narrows it down before anything is changed.

    • defining the symptom precisely
    • capacity vs workload problems
    • Spark job diagnostics
    • SQL query diagnostics
    • storage layout and file health
    • documenting findings
    • Microsoft Fabric
    • Performance
    • Debugging
  • PrévuDelta LakeAvancé

    MERGE into Delta Tables: Performance Patterns That Matter

    How Delta MERGE works internally and the patterns that keep it fast: narrowing the target, partition and file pruning, and source deduplication.

    Sujets prévus

    MERGE is the main upsert tool for Delta tables and also one of the most common causes of slow jobs. Most of the cost is in finding and rewriting files.

    • how MERGE finds matching files
    • narrowing the target with predicates
    • partition and file pruning
    • deduplicating the source
    • file rewrite cost
    • low-shuffle and optimised merge features
    • Delta Lake
    • MERGE
    • Performance
    • Upsert
  • PrévuDelta LakeIntermédiaire

    Choosing a Good Partition Column for Delta Lake

    Criteria for choosing a partition column: cardinality, query filters, write patterns and data volume per partition.

    Sujets prévus

    A good partition column helps reads, writes and maintenance at the same time. The article turns that into concrete criteria.

    • data volume per partition
    • cardinality
    • matching common filters
    • write and MERGE patterns
    • date columns and granularity
    • clustering as an alternative
    • Delta Lake
    • Partitioning
    • Performance
  • PrévuDelta LakeIntermédiaire

    Small Files in Delta Lake: Why They Hurt Performance

    Where small files come from in Delta tables, how they slow down reads, writes and the transaction log, and how to measure the problem.

    Sujets prévus

    Small files accumulate quietly through frequent writes and over-partitioning. By the time queries are slow, there may be millions of them.

    • sources of small files
    • overhead per file
    • impact on reads and MERGE
    • transaction log growth
    • measuring file counts and sizes
    • prevention at write time
    • Delta Lake
    • File Sizing
    • Performance
  • PrévuParquetIntermédiaire

    Choosing the Right Parquet File Size

    How file size affects read parallelism, metadata overhead and write cost, and how to pick a target size for your engines.

    Sujets prévus

    Files that are too small add overhead and files that are too large limit parallelism. The right size depends on the engines reading them.

    • what file size affects
    • too small vs too large
    • engine-specific guidance
    • controlling output size when writing
    • file size vs row group size
    • measuring the effect
    • Parquet
    • File Sizing
    • Performance
  • PrévuMicrosoft FabricIntermédiaire

    Working with Millions of Rows in Microsoft Fabric

    Build a practical baseline for file layout, Spark execution and Delta writes at multi-million-row scale.

    Sujets prévus

    Build a practical baseline for file layout, Spark execution and Delta writes at multi-million-row scale.

    • measuring before tuning
    • file and partition layout
    • Spark transformations and shuffles
    • efficient Delta writes
    • Microsoft Fabric
    • Large Scale Data
    • Performance
    • Delta Lake
  • PrévuMicrosoft FabricAvancé

    Working with Billions of Rows in Microsoft Fabric

    Plan storage, partitioning, incremental processing and capacity for billion-row Fabric workloads.

    Sujets prévus

    Plan storage, partitioning, incremental processing and capacity for billion-row Fabric workloads.

    • workload and capacity planning
    • partition and file sizing
    • incremental processing
    • measuring production-scale behavior
    • Microsoft Fabric
    • Billions of Rows
    • Performance
    • Partitioning

Microsoft Fabric Notebook Performance5 prévus

  • PrévuMicrosoft FabricIntermédiaire

    When to Reuse a Notebook Session vs Start a New One

    Choose between shared and isolated Fabric notebook sessions based on compatibility, concurrency, state, reliability and cost.

    Sujets prévus

    Session reuse can reduce acquisition overhead and improve utilization, while isolation can protect incompatible or resource-heavy workloads.

    • Compatibility and single-user boundaries
    • Shared resources versus isolated REPL state
    • Sequential and concurrent orchestration patterns
    • Reliability, monitoring, cost and workload isolation tradeoffs
    • Microsoft Fabric
    • Fabric Notebook
    • Performance
    • PySpark
  • PrévuMicrosoft FabricIntermédiaire

    Common Performance Mistakes in PySpark Notebooks

    Diagnose repeated actions, driver collection, shuffle-heavy logic, skew, inappropriate caching and output anti-patterns in Fabric notebooks.

    Sujets prévus

    Convenient interactive patterns can trigger repeated jobs, driver pressure, shuffles, skew and small-file output when moved unchanged into production.

    • Repeated actions and long lineage
    • `collect`, `toPandas`, Python UDFs and driver limits
    • Join, aggregation, repartition and skew mistakes
    • Cache lifecycle and plan-driven debugging
    • Microsoft Fabric
    • Fabric Notebook
    • Performance
    • PySpark
    • Troubleshooting
  • PrévuMicrosoft FabricIntermédiaire

    How to Read Less Data in Fabric Notebooks

    Reduce Fabric notebook I/O through projection, early filtering, incremental boundaries, partition pruning and Delta data skipping.

    Sujets prévus

    The cheapest byte to process is one the notebook never reads, but source syntax must be confirmed against the executed plan and table design.

    • Column projection and early predicates
    • Incremental time and key boundaries
    • Partition pruning, data skipping and file layout
    • Plan and metric checks that prove less data was scanned
    • Microsoft Fabric
    • Fabric Notebook
    • Performance
    • PySpark
    • Delta Lake
  • PrévuMicrosoft FabricIntermédiaire

    How to Write Delta Tables Efficiently from a Fabric Notebook

    Design reliable Fabric notebook Delta writes around mode, partition counts, file size, compaction, V-Order and rerun behavior.

    Sujets prévus

    Write performance and future read performance are connected through distribution, file layout, table properties and idempotent pipeline design.

    • Append, overwrite and merge contracts
    • Repartition, coalesce and small-file tradeoffs
    • Optimize Write, compaction, V-Order and Z-Order
    • Output validation, concurrency and safe reruns
    • Microsoft Fabric
    • Fabric Notebook
    • Performance
    • PySpark
    • Delta Lake
    • Table Maintenance
  • PrévuMicrosoft FabricIntermédiaire

    How to Debug a Slow Fabric Notebook

    Separate session acquisition, Spark execution, data skew, shuffle, I/O and writes when a Fabric notebook runs slowly.

    Sujets prévus

    Notebook elapsed time spans multiple layers, so changing code or capacity without locating the slow layer often moves the symptom rather than fixing it.

    • Startup versus first action versus Spark stages
    • Plans, jobs, tasks, skew, spill and shuffle
    • Read and Delta write diagnostics
    • Comparable baselines, monitoring and regression evidence
    • Microsoft Fabric
    • Fabric Notebook
    • Performance
    • PySpark
    • Debugging
    • Troubleshooting

SQL Server Performance6 prévus

  • PrévuSQL ServerDébutant

    How SQL Server Executes a Query: From Parsing to Execution Plan

    A walk through what SQL Server does between receiving a query and returning rows: parsing, binding, optimisation, plan caching and execution.

    Sujets prévus

    Most performance problems are easier to reason about once you know which stage of query processing they come from. This article builds that mental model before the rest of the series goes deeper into plans, statistics and indexing.

    • parsing and binding
    • the query optimiser and cost-based decisions
    • trivial plans vs full optimisation
    • plan caching and reuse
    • execution: operators, iterators and memory grants
    • where each stage shows up when troubleshooting
    • SQL Server
    • Execution Plans
    • Performance
  • PrévuSQL ServerIntermédiaire

    Reading SQL Server Execution Plans Without Guessing

    A systematic way to read execution plans: estimated vs actual rows, the operators that matter, warnings, and how to find the real cost instead of the highest percentage.

    Sujets prévus

    Execution plans are often read top to bottom looking for the biggest percentage, which regularly points at the wrong operator. The article proposes a repeatable reading order based on row estimates and data flow.

    • estimated vs actual plans
    • reading data flow right to left
    • cardinality estimate mismatches
    • the operators that usually matter
    • plan warnings: spills, implicit conversions, missing indexes
    • why cost percentages mislead
    • a checklist for every plan
    • SQL Server
    • Execution Plans
    • Performance
  • PrévuSQL ServerIntermédiaire

    When a Missing Index Recommendation Is Actually a Bad Idea

    Why SQL Server's missing index suggestions are a starting point and not an instruction, and how to evaluate them against the whole workload.

    Sujets prévus

    Missing index hints are generated per query, ignore existing indexes and know nothing about write cost. Applying them blindly leads to overlapping indexes and slower writes.

    • how missing index suggestions are produced
    • known limitations of the DMVs
    • overlap with existing indexes
    • column order and included columns
    • write and storage cost
    • consolidating suggestions across a workload
    • SQL Server
    • Indexing
    • Execution Plans
  • PrévuSQL ServerIntermédiaire

    SQL Server Statistics: Why Good Indexes Can Still Produce Bad Plans

    How statistics drive cardinality estimates, why they go stale on large tables, and how sampling and update thresholds affect plan quality.

    Sujets prévus

    A correct index is not enough if the optimiser misjudges how many rows it will read. Statistics are usually where that misjudgement starts.

    • what statistics contain: histogram and density
    • auto-update thresholds on large tables
    • sampling rate and skewed data
    • ascending keys
    • cardinality estimator versions
    • maintenance strategies
    • SQL Server
    • Statistics
    • Execution Plans
    • Performance
  • PrévuSQL ServerAvancé

    Parameter Sniffing in SQL Server: Diagnose It Before You Fix It

    How to confirm that a regression really is parameter sensitivity before reaching for OPTION(RECOMPILE) or plan forcing, and how to choose between the available fixes.

    Sujets prévus

    Parameter sniffing gets blamed for many slow queries and is often fixed with a blanket recompile hint. The article focuses on proving it first and then picking the least invasive fix.

    • how plans are compiled for parameter values
    • symptoms that look like sniffing but are not
    • confirming with Query Store and plan comparison
    • skewed data distributions
    • fix options and their costs
    • Parameter Sensitive Plan optimisation
    • SQL Server
    • Execution Plans
    • Query Store
    • Performance
  • PrévuSQL ServerIntermédiaire

    SQL Server Query Store for Production Performance Investigations

    Using Query Store as the first stop when production performance changes: configuration, finding regressions, comparing plans and forcing plans safely.

    Sujets prévus

    When someone reports that the system got slower, the first question is what changed. Query Store keeps the history needed to answer it, provided it is configured sensibly.

    • configuration that works in production
    • finding regressed queries
    • comparing plans over time
    • wait statistics in Query Store
    • plan forcing and its risks
    • storage and cleanup
    • SQL Server
    • Query Store
    • Performance
    • Debugging

SQL Server Troubleshooting & Diagnostics15 prévus

  • PrévuSQL ServerIntermédiaire

    How to Find Long-Running Queries in SQL Server

    Find active requests above a configurable duration and interpret elapsed time alongside CPU, reads, writes, status and waits.

    Sujets prévus

    A long duration can describe legitimate ETL, reporting, index maintenance, a large update, or a request that is merely blocked. The useful diagnosis compares elapsed time with work and waits instead of declaring every old request unhealthy.

    • A safe threshold parameter in seconds or minutes
    • CPU, logical and physical reads, writes, status, wait, command and SQL text
    • Examples for ETL, reporting, maintenance, updates and blocked requests
    • How to distinguish work in progress from time spent waiting
    • SQL Server
    • Performance
    • Troubleshooting
    • DMV
  • PrévuSQL ServerIntermédiaire

    How to Find the Queries Using the Most CPU in SQL Server

    Rank current and cached SQL Server workload by total CPU and average CPU per execution, with the right cache-lifetime caveats.

    Sujets prévus

    Total CPU identifies cumulative consumers, while average CPU highlights individually expensive executions. Both are needed, and neither is durable history when its evidence comes from the plan cache.

    • Current workload and `sys.dm_exec_query_stats`
    • Total worker time, execution count, total CPU and average CPU
    • Separate total and per-execution rankings
    • Cache eviction, recompilation, restart, failover and cache-reset caveats
    • SQL Server
    • Performance
    • Troubleshooting
    • DMV
  • PrévuSQL ServerIntermédiaire

    How to Find Queries Doing the Most Reads in SQL Server

    Rank cached SQL Server queries by total and average logical reads and connect the result to plans, scans, estimates and indexing.

    Sujets prévus

    Logical reads measure pages processed and often expose expensive access patterns more reliably than duration, which changes with blocking and server load. High reads still need workload and plan context.

    • Total and average logical-read rankings from `sys.dm_exec_query_stats`
    • Physical reads, scans and large range operations
    • Connections to indexing, estimates and large-table access
    • Plan-cache lifetime and reset limitations
    • SQL Server
    • Performance
    • Troubleshooting
    • DMV
    • Indexing
  • PrévuSQL ServerIntermédiaire

    How to See SQL Server Index Sizes and Usage

    Inspect index definitions, size, row count, seeks, scans, lookups, updates and observation-window limitations.

    Sujets prévus

    Indexes trade read access paths for storage and write maintenance. Zero seeks in a DMV observation window does not prove that an index is safe to remove.

    • Keys, included columns, index type, size and row count
    • Seeks, scans, lookups and updates
    • Restart, failover, attach/detach and reset conditions
    • Observation windows, overlapping indexes and write cost
    • SQL Server
    • Performance
    • Troubleshooting
    • DMV
    • Indexing
  • PrévuSQL ServerIntermédiaire

    How to Find Missing Index Recommendations in SQL Server

    Read SQL Server missing-index DMV recommendations and evaluate benefit, overlap, width and write cost before changing production.

    Sujets prévus

    Missing-index DMVs produce optimization hints, not complete index designs. Recommendations can overlap, become extremely wide, disappear after restart, and omit the cost imposed on writes.

    • Equality, inequality and included columns
    • Estimated improvement and observation-window limitations
    • Duplicate and overlapping recommendations
    • A safe plan-first workflow for consolidation, testing and validation
    • SQL Server
    • Performance
    • Troubleshooting
    • DMV
    • Indexing
  • PrévuSQL ServerIntermédiaire

    How to Check SQL Server Statistics

    Inspect statistics metadata, update dates, samples and modification counters, then use histograms only when the evidence calls for them.

    Sujets prévus

    Statistics influence cardinality estimates and plan choices, but age alone does not prove they are harmful. The investigation should connect sampling and modifications to estimate errors in a specific workload.

    • `sys.stats` and `sys.dm_db_stats_properties`
    • Auto-created versus index statistics, sample sizes and modification counters
    • `DBCC SHOW_STATISTICS`, density and histogram interpretation
    • Large-table considerations without blindly updating every statistic
    • SQL Server
    • Performance
    • Troubleshooting
    • Statistics
  • PrévuSQL ServerIntermédiaire

    How to Check Whether a Stored Procedure Is Still Being Used

    Combine cached procedure execution statistics, Query Store and application evidence before deciding that a stored procedure is unused.

    Sujets prévus

    Zero executions in `sys.dm_exec_procedure_stats` only means the procedure is absent from the current cache evidence. It does not prove that a monthly job, rare recovery path, or external integration never calls it.

    • Cached execution count and last execution time
    • Restart, recompile, cache eviction and reset caveats
    • Query Store, application tracing and sufficient observation windows
    • A safe evidence checklist before deprecation or removal
    • SQL Server
    • Troubleshooting
    • DMV
    • Query Store
  • PrévuSQL ServerIntermédiaire

    How to Use Query Store to Find SQL Server Performance Regressions

    Compare Query Store time windows, resource consumers and plan changes to investigate SQL Server regressions safely.

    Sujets prévus

    Query Store connects present symptoms to historical runtime statistics and plans. Comparing equivalent time windows can reveal regression after a deployment, upgrade, or compatibility-level change.

    • Duration, CPU, reads, execution count and top resource consumers
    • Comparing incident and baseline windows
    • Plan changes and regressed-query investigation
    • Careful plan forcing, upgrades and compatibility-level changes
    • SQL Server
    • Performance
    • Troubleshooting
    • Query Store
  • PrévuSQL ServerIntermédiaire

    How to Identify Which Login, Host or Application Is Generating SQL Server Load

    Group active SQL Server work by login, host, application and database to explain API, report, batch and integration spikes.

    Sujets prévus

    An instance-wide spike becomes actionable when it can be attributed to a workload owner. Login, host and program values are useful signals, though client-supplied identity must be interpreted carefully.

    • Grouping active requests by login, host, program and database
    • Request count, CPU, reads, writes and elapsed time
    • API spikes, reporting systems, batch jobs and integrations
    • Anonymous examples using `ApiService`, `ReportingUser` and `SalesDb`
    • SQL Server
    • Performance
    • Troubleshooting
    • DMV
  • PrévuSQL ServerAvancé

    How to Investigate Very Large SQL Server Tables

    Investigate hundred-million and billion-row tables through workload, storage, distribution, maintenance and retention evidence.

    Sujets prévus

    Row count alone does not select a physical design. Very large tables demand a workload-led review of access, writes, retention, distribution and operational constraints.

    • Table and index size, row distribution and retention
    • Most-used queries, logical reads and write patterns
    • Statistics and index-maintenance constraints
    • Partition elimination and columnstore suitability without automatic prescriptions
    • SQL Server
    • Performance
    • Troubleshooting
    • Indexing
    • Statistics
    • Partitioning
    • Large Scale Data
    • Billions of Rows
  • PrévuSQL ServerAvancé

    How to Troubleshoot a Large UPDATE in SQL Server

    Diagnose a large SQL Server update through log use, locking, row count, predicates, plans, statistics, indexes and rollback risk.

    Sujets prévus

    A large update can be correct yet operationally disruptive through transaction-log growth, blocking, index maintenance and rollback exposure. Batching changes those tradeoffs but is not automatically safer or faster.

    • Row count, predicate selectivity, plan and statistics
    • Transaction log, locks, blocking and index write cost
    • One massive transaction versus controlled batches
    • Batch-ordering, correctness, recovery and rollback tradeoffs
    • SQL Server
    • Performance
    • Troubleshooting
    • Blocking
    • Indexing
    • Statistics
  • PrévuSQL ServerIntermédiaire

    How to Compare SQL Server Performance Before and After a Change

    Design an evidence-based SQL Server comparison for upgrades, compatibility levels, indexes, deployments and infrastructure changes.

    Sujets prévus

    Averages can improve while the slowest user requests regress. A useful comparison controls workload and time windows and reports distribution, work and frequency together.

    • Upgrade, compatibility level, index, deployment and infrastructure cases
    • Average, median and P95 duration
    • CPU, logical reads, writes and execution count
    • Query Store baselines, comparable windows and interpreting mixed outcomes
    • SQL Server
    • Performance
    • Troubleshooting
    • Query Store
  • PrévuSQL ServerIntermédiaire

    SQL Server Wait Types: What Is the Query Waiting For?

    Use SQL Server request and aggregate waits as evidence while avoiding simplistic wait-to-fix mappings.

    Sujets prévus

    A wait type identifies where execution currently cannot proceed; it does not independently explain why. The query, plan, resource, workload and time window supply the diagnosis.

    • Current request waits and aggregate observation windows
    • `LCK_M_*`, `PAGEIOLATCH_*`, `WRITELOG` and `RESOURCE_SEMAPHORE`
    • `CXPACKET`, `CXCONSUMER`, `ASYNC_NETWORK_IO` and `SOS_SCHEDULER_YIELD`
    • Evidence-driven follow-up rather than “wait X means fix Y”
    • SQL Server
    • Performance
    • Troubleshooting
    • DMV
  • PrévuSQL ServerIntermédiaire

    How to Inspect Open Transactions in SQL Server

    Find long-running and sleeping-session transactions, their owners, database and potential blocking and transaction-log impact.

    Sujets prévus

    An old transaction can retain locks and delay log truncation even when its session is sleeping. Joining transaction, session and request state makes that hidden ownership visible.

    • `sys.dm_tran_session_transactions` and `sys.dm_tran_active_transactions`
    • `sys.dm_tran_database_transactions` and database log impact
    • Transaction begin time, session owner, request and blocking context
    • Why sleeping connections with open transactions require special attention
    • SQL Server
    • Performance
    • Troubleshooting
    • DMV
    • Blocking
  • PrévuSQL ServerIntermédiaire

    SQL Server Troubleshooting Queries: Quick Reference

    A concise companion index for the SQL Server diagnostic queries in this series and the performance-tuning checklist resource.

    Sujets prévus

    During an incident, readers need a trustworthy route to the right diagnostic rather than one enormous script with no interpretation. This reference will index the focused articles and resource checklist without duplicating their reasoning.

    • Current requests, sessions, blocking and open transactions
    • CPU, reads, table size, index usage and statistics
    • Procedure usage and Query Store history
    • The SQL Server Performance Tuning Checklist resource and future companion resources
    • SQL Server
    • Performance
    • Troubleshooting
    • DMV
    • Blocking
    • Indexing
    • Statistics
    • Query Store

SQL Server at Scale9 prévus

  • PrévuSQL ServerAvancé

    SQL Server Indexing for Tables with Hundreds of Millions of Rows

    Index design when tables are too large to rebuild casually: clustered key choice, covering indexes, columnstore, maintenance windows and the write cost of every index.

    Sujets prévus

    On very large tables an index decision is expensive to reverse: building, rebuilding or dropping it takes hours and affects every write. The article covers how to choose indexes deliberately at that scale.

    • choosing the clustered key
    • nonclustered and covering indexes
    • rowstore vs columnstore
    • filtered indexes
    • write amplification and insert cost
    • online index operations and maintenance windows
    • measuring index usage before adding more
    • SQL Server
    • Indexing
    • Large Scale Data
    • Performance
  • PrévuSQL ServerIntermédiaire

    Batching Large UPDATE and DELETE Operations in SQL Server

    Patterns for changing or deleting millions of rows without long blocking, transaction log growth or lock escalation.

    Sujets prévus

    A single statement that touches millions of rows can fill the transaction log, escalate locks and block the application. Batching fixes that, but only if the batches are designed carefully.

    • why single large statements hurt
    • choosing a batch size
    • key-range batching vs TOP
    • transaction log and recovery model
    • lock escalation
    • progress tracking and restartability
    • SQL Server
    • Large Scale Data
    • Performance
    • Concurrency
  • PrévuSQL ServerAvancé

    Moving Hundreds of Millions of Rows Without Destroying SQL Server Performance

    Options for copying or migrating very large tables, from batched inserts to partition switching and bulk load, and how to keep the source system usable meanwhile.

    Sujets prévus

    Large data moves compete with the production workload for I/O, log and locks. The approach needs to be chosen based on downtime tolerance and how much the source can absorb.

    • choosing between batching, bulk load and switching
    • minimal logging requirements
    • throttling and scheduling
    • indexes during the load
    • validating row counts and checksums
    • restarting after failure
    • SQL Server
    • Large Scale Data
    • Data Ingestion
    • Performance
  • PrévuSQL ServerIntermédiaire

    SQL Server Partitioning: When It Helps and When It Does Not

    What table partitioning really gives you, mainly data management, and why it is not an automatic query performance feature.

    Sujets prévus

    Partitioning is often introduced to make queries faster and ends up making some of them slower. Its real strengths are loading, archiving and maintenance.

    • partition functions and schemes
    • partition elimination
    • aligned and non-aligned indexes
    • maintenance per partition
    • cases where partitioning hurts
    • alternatives to consider first
    • SQL Server
    • Partitioning
    • Large Scale Data
  • PrévuSQL ServerAvancé

    Designing Partition Keys for Very Large SQL Server Tables

    How to choose the partition column and boundaries for very large tables, and how that choice interacts with the clustered index and query patterns.

    Sujets prévus

    The partition key is hard to change later and affects loading, querying and maintenance. Choosing it means balancing all three.

    • candidate keys: dates, tenants, ingestion batches
    • boundary granularity
    • RANGE LEFT vs RANGE RIGHT
    • clustered index alignment
    • query patterns that miss elimination
    • managing boundaries over time
    • SQL Server
    • Partitioning
    • Large Scale Data
    • Indexing
  • PrévuSQL ServerAvancé

    Partition Switching in SQL Server for High-Volume Data Loads

    Using staging tables and partition switching to load or replace large slices of data as a near-instant metadata operation.

    Sujets prévus

    Partition switching turns a long, blocking load into a metadata operation, but it has strict requirements that are easy to miss.

    • requirements for switching
    • staging table design
    • constraints and indexes
    • switch-in and switch-out patterns
    • sliding window maintenance
    • locking during the switch
    • SQL Server
    • Partitioning
    • Data Ingestion
    • Large Scale Data
  • PrévuSQL ServerIntermédiaire

    SQL Server MERGE vs INSERT/UPDATE: What I Prefer in Production

    A comparison of MERGE with separate INSERT and UPDATE statements: readability, concurrency behaviour, known issues and the pattern I prefer in production.

    Sujets prévus

    MERGE looks like the natural upsert statement, but its concurrency behaviour and history of issues lead many teams to avoid it. The article explains the trade-offs and states a preference with reasons.

    • what MERGE does in one statement
    • concurrency and race conditions
    • known issues and caveats
    • separate UPDATE then INSERT
    • performance comparison method
    • my production default and why
    • SQL Server
    • MERGE
    • Upsert
  • PrévuSQL ServerIntermédiaire

    Building Reliable Upsert Patterns Without SQL Server MERGE

    Upsert patterns that behave correctly under concurrency, for single rows and for set-based batches, without relying on MERGE.

    Sujets prévus

    Upserts that work in testing can produce duplicates or primary key violations under concurrent load. The article shows patterns that hold up.

    • single-row vs set-based upserts
    • UPDATE then INSERT with the right locking
    • unique constraints as a safety net
    • handling deletes
    • idempotency
    • testing under concurrency
    • SQL Server
    • Upsert
    • Concurrency
  • PrévuSQL ServerIntermédiaire

    Optimizing INSERT Performance for Large SQL Server Workloads

    What limits insert throughput, from logging and indexes to page splits and client batching, and how to remove each bottleneck.

    Sujets prévus

    Slow inserts are rarely caused by the INSERT statement itself. The bottleneck is usually in logging, indexes, constraints or the client.

    • row-by-row vs batched inserts
    • bulk insert and minimal logging
    • indexes and constraints during load
    • page splits and key choice
    • transaction log throughput
    • client-side batching
    • SQL Server
    • Data Ingestion
    • Performance
    • Large Scale Data

Fabric in Production6 prévus

  • PrévuMicrosoft FabricAvancé

    Microsoft Fabric Architecture for High-Volume Data Platforms

    Architecture decisions for a Fabric platform that has to handle large volumes: storage layout, compute choices, workspace boundaries and the limits to plan for.

    Sujets prévus

    Fabric makes it easy to start but the early layout decisions matter once volumes grow. The article lays out the decisions worth making deliberately up front.

    • OneLake storage layout
    • Lakehouse, Warehouse and SQL endpoint roles
    • workspace boundaries
    • capacity planning
    • ingestion paths
    • known limits and how to plan around them
    • Microsoft Fabric
    • Data Architecture
    • Large Scale Data
  • PrévuMicrosoft FabricAvancé

    Designing a Multi-Tenant Data Platform in Microsoft Fabric

    Isolation options for serving many tenants from Fabric, from shared tables with row-level security to workspace per tenant, and their operational cost.

    Sujets prévus

    Tenant isolation affects security, cost, deployment and performance at the same time. The right model depends on how many tenants there are and how different they are.

    • isolation models compared
    • row-level security on shared tables
    • workspace or item per tenant
    • onboarding automation
    • noisy neighbours and capacity
    • deployment across tenants
    • Microsoft Fabric
    • Multi-Tenancy
    • Data Architecture
  • PrévuMicrosoft FabricIntermédiaire

    Microsoft Fabric Capacity: Understanding Where Your Compute Goes

    How Fabric capacity units are consumed by different workloads, how smoothing and throttling work, and how to attribute usage to items and jobs.

    Sujets prévus

    Capacity is shared by every workload in it, so one heavy job can affect everything else. Understanding how usage is measured is the basis for sizing and troubleshooting.

    • capacity units and SKUs
    • interactive vs background operations
    • smoothing and bursting
    • throttling stages
    • the Capacity Metrics app
    • attributing usage to items
    • Microsoft Fabric
    • Capacity
    • Performance
  • PrévuMicrosoft FabricIntermédiaire

    How I Structure Fabric Workspaces for Dev, Test and Production

    A workspace layout for separating environments in Fabric, with naming, permissions and how items reference each other across workspaces.

    Sujets prévus

    Environment separation in Fabric starts with workspace design. Getting it wrong makes deployments and permissions painful later.

    • workspace per environment and per domain
    • naming conventions
    • permissions and roles
    • cross-workspace references and shortcuts
    • capacity assignment per environment
    • what not to share between environments
    • Microsoft Fabric
    • DevOps
    • CI/CD
  • PrévuMicrosoft FabricIntermédiaire

    Fabric Lakehouse Design Patterns for Large Data Platforms

    Patterns for organising tables, files, schemas and shortcuts in a Fabric Lakehouse that has to scale to many sources and teams.

    Sujets prévus

    A Lakehouse with a few tables is easy. One with hundreds needs conventions for layout, ownership and naming.

    • Tables vs Files areas
    • schemas and naming
    • shortcuts and when to use them
    • one Lakehouse or several
    • table ownership
    • documentation and discoverability
    • Microsoft Fabric
    • Fabric Lakehouse
    • Data Architecture
    • Delta Lake
  • PrévuMicrosoft FabricIntermédiaire

    Microsoft Fabric Pipelines vs Notebooks: Choosing the Right Tool

    When to use Data Factory pipelines, when to use notebooks, and how to combine them so orchestration and transformation stay clearly separated.

    Sujets prévus

    Pipelines and notebooks overlap enough that teams end up mixing logic between them. Clear roles make platforms easier to debug and deploy.

    • orchestration vs transformation
    • copy activities vs Spark
    • parameters and dynamic content
    • error handling in each
    • cost and startup time
    • a combined pattern
    • Microsoft Fabric
    • Data Pipelines
    • Python

Production Data Engineering3 prévus

  • PrévuMicrosoft FabricIntermédiaire

    Designing Idempotent Fabric Data Pipelines

    How to design pipelines that can be rerun safely after a failure without duplicating or losing data.

    Sujets prévus

    Pipelines fail, and someone will rerun them. If a rerun creates duplicates or skips data, every failure becomes a data quality incident.

    • what idempotency means for data loads
    • deterministic batch identifiers
    • overwrite vs append vs merge
    • watermarks and their storage
    • partial failure handling
    • testing reruns
    • Microsoft Fabric
    • Data Pipelines
    • Upsert
  • PrévuMicrosoft FabricIntermédiaire

    Handling Schema Changes Safely in Microsoft Fabric

    How to absorb source schema changes in Fabric without breaking pipelines or consumers: detection, evolution options and contracts.

    Sujets prévus

    Source systems change columns without warning. A platform needs to decide which changes it absorbs automatically and which should stop the pipeline.

    • types of schema change
    • detecting changes early
    • Delta schema evolution options
    • protecting the curated layer
    • communicating changes to consumers
    • testing with changed schemas
    • Microsoft Fabric
    • Schema Evolution
    • Delta Lake
  • PrévuMicrosoft FabricAvancé

    How to Design Fabric Pipelines for Hundreds of Daily Data Files

    Ingestion design for many small files a day: discovery, batching, parallelism, tracking what was processed and avoiding small-file problems downstream.

    Sujets prévus

    Processing files one at a time does not scale, and processing them all at once is hard to recover. The design has to handle both throughput and traceability.

    • file discovery and landing zones
    • batching files per run
    • parallelism limits
    • tracking processed files
    • handling bad or duplicate files
    • compacting output
    • Microsoft Fabric
    • Data Pipelines
    • Data Ingestion
    • File Sizing

Delta Lake at Scale12 prévus

  • PrévuDelta LakeDébutant

    Delta Tables Explained for SQL Engineers

    Delta Lake concepts mapped to what SQL Server engineers already know: transactions, the log, updates, deletes and what has no equivalent.

    Sujets prévus

    Engineers coming from SQL Server bring useful intuition and some wrong assumptions. Mapping concepts side by side speeds up the transition.

    • files plus a transaction log
    • ACID guarantees and their scope
    • updates and deletes as file rewrites
    • no indexes in the SQL Server sense
    • time travel
    • what to unlearn
    • Delta Lake
    • Microsoft Fabric
    • SQL Server
  • PrévuDelta LakeIntermédiaire

    How Delta Lake Actually Stores Data

    A look inside a Delta table folder: Parquet data files, the _delta_log, checkpoints and how a read reconstructs the current version.

    Sujets prévus

    Most Delta performance issues make sense once you look at the files on disk. This article opens the folder and explains each part.

    • Parquet data files
    • the _delta_log JSON commits
    • checkpoints
    • add and remove actions
    • how readers find the current snapshot
    • deletion vectors
    • Delta Lake
    • Parquet
    • File Sizing
  • PrévuDelta LakeAvancé

    When Delta MERGE Becomes Slow and How to Diagnose It

    A diagnostic approach for slow MERGE jobs: reading the operation metrics, finding the expensive phase and matching it to a fix.

    Sujets prévus

    A MERGE that took minutes last month can take hours today without any code change. The table history and job metrics usually explain why.

    • reading DESCRIBE HISTORY metrics
    • files scanned vs files rewritten
    • skew in the join
    • growth of the target table
    • small files
    • fix options by cause
    • Delta Lake
    • MERGE
    • Debugging
    • Performance
  • PrévuDelta LakeIntermédiaire

    INSERT vs MERGE in Delta Tables

    When appends are enough and when MERGE is required, and how append-then-deduplicate patterns compare on cost and correctness.

    Sujets prévus

    MERGE is often used where a plain append would do, at a much higher cost. Choosing correctly depends on the data's change pattern.

    • append-only data
    • data that changes
    • append then deduplicate
    • replaceWhere and partition overwrite
    • cost comparison method
    • correctness risks of each
    • Delta Lake
    • MERGE
    • Data Ingestion
  • PrévuDelta LakeIntermédiaire

    Designing Delta Tables for Hundreds of Millions of Rows

    Design choices that matter once a Delta table holds hundreds of millions of rows: partitioning, clustering, file size and write patterns.

    Sujets prévus

    At this size the table is still manageable, but early design choices start to show. It is the right time to set conventions before the table grows further.

    • whether to partition yet
    • Z-ORDER and liquid clustering
    • target file size
    • write patterns
    • query patterns
    • maintenance schedule
    • Delta Lake
    • Large Scale Data
    • Partitioning
    • File Sizing
  • PrévuDelta LakeAvancé

    Designing Delta Tables for Billions of Rows

    Practical considerations for partitioning, file size, ingestion, MERGE operations and query performance when Delta tables reach very large scale.

    Sujets prévus

    At billions of rows, mistakes in table design turn into long-running jobs and expensive maintenance. The article covers the decisions that matter most at that scale.

    • partitioning strategy
    • file sizing
    • ingestion patterns
    • MERGE performance
    • concurrency
    • maintenance
    • query patterns
    • monitoring
    • Delta Lake
    • Billions of Rows
    • Large Scale Data
    • Partitioning
  • PrévuDelta LakeIntermédiaire

    Partitioning Delta Tables: The Mistakes That Create More Problems Than They Solve

    Common Delta partitioning mistakes, such as high-cardinality keys and tiny partitions, what they cause, and how to recover from them.

    Sujets prévus

    Over-partitioning is one of the most common causes of slow Delta tables. It is also expensive to undo once data is written.

    • partitioning by high-cardinality columns
    • partitions that are too small
    • partitioning on columns nobody filters by
    • symptoms in file counts and query plans
    • repartitioning an existing table
    • when not to partition
    • Delta Lake
    • Partitioning
    • File Sizing
  • PrévuDelta LakeIntermédiaire

    Delta Table Compaction and File Optimization Strategies

    Compaction options for Delta tables, including OPTIMIZE, optimised writes, auto compaction and V-Order, and how to schedule them.

    Sujets prévus

    Compaction fixes small files but costs compute and can conflict with writers. The strategy should match how the table is written and read.

    • OPTIMIZE and bin-packing
    • optimised writes
    • auto compaction
    • V-Order in Fabric
    • targeting partitions
    • scheduling and conflicts
    • Delta Lake
    • Table Maintenance
    • File Sizing
    • Performance
  • PrévuDelta LakeAvancé

    Delta Lake Concurrency: What Happens When Multiple Jobs Write at Once

    How Delta's optimistic concurrency control works, which operations conflict, and how to design jobs that write to the same table safely.

    Sujets prévus

    Two jobs writing to the same table can both succeed, or one can fail with a conflict. Knowing which is which avoids surprises in production.

    • optimistic concurrency control
    • conflict types
    • isolation levels
    • partition-level separation
    • retry behaviour
    • coordinating writers
    • Delta Lake
    • Concurrency
  • PrévuDelta LakeAvancé

    Handling Concurrent Append Conflicts in Delta Tables

    Why ConcurrentAppendException happens, how to read it, and the design and retry patterns that prevent it.

    Sujets prévus

    Concurrent append conflicts often appear only under production load. The fix is usually in how writers' predicates overlap.

    • what the exception means
    • overlapping read predicates
    • making predicates explicit
    • partitioning to separate writers
    • safe retry logic
    • serialising when necessary
    • Delta Lake
    • Concurrency
    • Debugging
  • PrévuDelta LakeIntermédiaire

    Delta Schema Evolution: Safe Patterns for Production Systems

    Using mergeSchema, column mapping and explicit migrations to evolve Delta tables without breaking writers or readers.

    Sujets prévus

    Automatic schema evolution is convenient and also risky: an unexpected column can spread silently. The article covers when to allow it and when to require a migration.

    • schema enforcement
    • mergeSchema and autoMerge
    • adding, renaming and dropping columns
    • column mapping
    • type changes
    • explicit migrations
    • Delta Lake
    • Schema Evolution
  • PrévuDelta LakeIntermédiaire

    Delta Table Maintenance: OPTIMIZE, VACUUM and Production Tradeoffs

    Planning routine maintenance for Delta tables: what OPTIMIZE and VACUUM do, retention and time travel trade-offs, and how often to run them.

    Sujets prévus

    Without maintenance, Delta tables slow down and storage grows. With careless maintenance, time travel and concurrent readers can break.

    • what OPTIMIZE changes
    • what VACUUM deletes
    • retention and time travel
    • risks to concurrent readers
    • scheduling per table
    • monitoring maintenance
    • Delta Lake
    • Table Maintenance
    • Performance

Understanding Parquet8 prévus

  • PrévuParquetDébutant

    Parquet Files Explained for Database Engineers

    The Parquet format explained in database terms: columnar layout, row groups, pages, encodings, compression and metadata.

    Sujets prévus

    Parquet is the storage format under Delta Lake and Fabric. Understanding it explains much of their performance behaviour.

    • row vs columnar storage
    • file structure: row groups, column chunks, pages
    • encodings
    • compression codecs
    • footer metadata and statistics
    • what Parquet does not do
    • Parquet
    • Performance
  • PrévuParquetDébutant

    Why Parquet Is Fast for Analytics

    The specific mechanisms that make Parquet fast for analytical queries: column pruning, predicate pushdown, encoding and compression.

    Sujets prévus

    Parquet is described as fast, but the speed comes from specific mechanisms, and each one depends on how the file was written.

    • column pruning
    • predicate pushdown with statistics
    • dictionary and run-length encoding
    • compression
    • vectorised reading
    • when Parquet is not fast
    • Parquet
    • Performance
  • PrévuParquetIntermédiaire

    The Small-File Problem in Data Lakes

    Why data lakes accumulate small files, what that costs in listing, metadata and query time, and the ingestion patterns that avoid it.

    Sujets prévus

    Small files are a general data lake problem, not only a Delta one. Most of them are created by ingestion patterns that can be changed.

    • how small files are created
    • cost in listing and metadata
    • cost in query planning and execution
    • batching at ingestion
    • compaction jobs
    • monitoring file counts
    • Parquet
    • File Sizing
    • Data Ingestion
  • PrévuParquetIntermédiaire

    Partitioned Parquet: How Much Partitioning Is Too Much?

    Hive-style partitioning of Parquet datasets: how directory partitions help pruning, and the point at which they create more files than benefit.

    Sujets prévus

    Directory partitioning is easy to add and easy to overdo. The cost shows up as many small files and slow listing.

    • Hive-style directory partitioning
    • partition pruning
    • data per partition
    • multi-level partitioning
    • signs of over-partitioning
    • alternatives
    • Parquet
    • Partitioning
    • File Sizing
  • PrévuParquetAvancé

    Parquet Partitioning Strategies for Very Large Datasets

    Partitioning and sorting strategies for very large Parquet datasets, combining directory partitions with in-file ordering for effective pruning.

    Sujets prévus

    At very large scale, directory partitioning alone is too coarse. Sorting data within files is what makes statistics-based pruning effective.

    • coarse directory partitions
    • sorting within files
    • row group statistics for pruning
    • bucketing
    • write cost of sorting
    • validating pruning
    • Parquet
    • Partitioning
    • Large Scale Data
  • PrévuParquetAvancé

    Parquet Row Groups and Why They Matter for Performance

    What row groups are, how their size and statistics affect pruning and memory use, and how writers decide row group boundaries.

    Sujets prévus

    Row groups are the unit of pruning and parallelism inside a Parquet file. Their size is often left at defaults that do not suit the workload.

    • row group structure
    • min/max statistics
    • row group size vs file size
    • memory use when reading and writing
    • writer settings
    • inspecting row groups
    • Parquet
    • Performance
  • PrévuParquetIntermédiaire

    Schema Evolution with Parquet Files

    What happens when Parquet files with different schemas are read together, how engines reconcile them, and the changes that are safe.

    Sujets prévus

    Parquet files carry their own schema, so a dataset can contain several versions. Reading them together works only within certain limits.

    • schema in each file footer
    • adding and removing columns
    • type changes
    • schema merging when reading
    • differences between engines
    • why table formats help
    • Parquet
    • Schema Evolution
  • PrévuParquetAvancé

    Reading Billions of Rows from Parquet Efficiently

    Techniques for reading very large Parquet datasets: column selection, predicate pushdown, partition pruning, batching and choosing an engine.

    Sujets prévus

    Reading billions of rows is mostly about not reading most of them. The article covers how to make sure the engine skips what it can.

    • selecting only needed columns
    • predicate pushdown
    • partition and row group pruning
    • streaming and batch reading
    • choosing an engine
    • verifying what was actually read
    • Parquet
    • Billions of Rows
    • Performance
    • Python

Fabric SQL Engineering5 prévus

  • PrévuFabric SQLDébutant

    Fabric SQL Endpoint Explained for SQL Server Engineers

    What the Lakehouse SQL analytics endpoint is, how it differs from a SQL Server database, and which familiar features behave differently.

    Sujets prévus

    The SQL endpoint looks like a SQL Server database but works very differently underneath. Knowing the differences avoids wrong expectations.

    • what the endpoint is
    • read-only over Delta tables
    • metadata sync
    • supported and unsupported T-SQL
    • security
    • where it fits in an architecture
    • SQL Endpoint
    • Microsoft Fabric
    • SQL Server
  • PrévuFabric SQLIntermédiaire

    Lakehouse SQL Endpoint vs Fabric Warehouse

    A side-by-side comparison of the Lakehouse SQL endpoint and the Fabric Warehouse for querying and serving data.

    Sujets prévus

    Both expose T-SQL over Delta tables, but they differ in writes, transactions and table management. The choice affects how the platform is built.

    • write capabilities
    • transactions
    • table management
    • performance characteristics
    • security
    • when to use each
    • SQL Endpoint
    • Fabric Warehouse
    • Fabric Lakehouse
  • PrévuFabric SQLIntermédiaire

    Why a Fabric SQL Endpoint Can Be Slower Than You Expect

    The usual reasons SQL endpoint queries are slower than expected: table layout, file health, statistics, data types and query patterns.

    Sujets prévus

    Teams moving from SQL Server often expect similar performance from the same query. The reasons for differences are usually in the underlying Delta tables.

    • cold vs warm queries
    • file count and V-Order
    • statistics
    • string and data type choices
    • query patterns that do not translate
    • capacity effects
    • SQL Endpoint
    • Performance
    • Delta Lake
  • PrévuFabric SQLAvancé

    Troubleshooting Fabric SQL Endpoint Performance

    A workflow for diagnosing slow SQL endpoint queries using query insights, plans and table health checks.

    Sujets prévus

    Once a slow query is reported, there needs to be a repeatable way to find out whether the problem is the query, the table or the capacity.

    • reproducing the problem
    • query insights views
    • reading the plan
    • checking table and file health
    • capacity contention
    • fixes by cause
    • SQL Endpoint
    • Performance
    • Debugging
  • PrévuFabric SQLAvancé

    Designing Reporting APIs on Top of Fabric SQL Endpoints

    Design considerations for serving data through an API backed by a Fabric SQL endpoint: authentication, query shape, caching, latency and limits.

    Sujets prévus

    Serving an API from an analytical endpoint is possible but has different latency and concurrency characteristics from an OLTP database.

    • authentication and identities
    • connection management
    • query shape and pagination
    • caching
    • latency and concurrency expectations
    • when to use a different store
    • SQL Endpoint
    • API
    • Data Architecture
    • Python

Observability & Debugging7 prévus

  • PrévuSQL ServerIntermédiaire

    Finding the Queries That Are Really Hurting Your SQL Server

    How to identify the queries with the biggest real impact, by total resource use and frequency, instead of chasing the single slowest query.

    Sujets prévus

    The slowest query is rarely the most expensive one overall. A fast query that runs a million times a day can matter far more.

    • average vs total cost
    • CPU, reads, duration and waits
    • DMVs vs Query Store
    • grouping by query hash
    • separating the workload by application
    • building a short, repeatable report
    • SQL Server
    • Performance
    • Query Store
    • Observability
  • PrévuSQL ServerIntermédiaire

    How to Troubleshoot Blocking and Long-Running Sessions in SQL Server

    A practical workflow for finding the head blocker, understanding what it is waiting on, and deciding whether to wait, kill or fix.

    Sujets prévus

    Blocking incidents are stressful because the pressure is to kill something quickly. A clear workflow makes it possible to act fast without losing the evidence needed to prevent the next one.

    • finding the head blocker
    • sessions, requests and open transactions
    • lock types and lock escalation
    • isolation levels and row versioning
    • capturing evidence before killing a session
    • preventing recurrence
    • SQL Server
    • Concurrency
    • Debugging
  • PrévuSQL ServerAvancé

    SQL Server Deadlocks: A Practical Investigation Workflow

    How to capture and read deadlock graphs, identify the access pattern that causes the cycle, and choose a fix that removes it.

    Sujets prévus

    Deadlocks are usually handled with retries, which hides the cause. Reading the deadlock graph properly usually points to a specific access order or missing index.

    • capturing deadlocks with Extended Events
    • reading a deadlock graph
    • common deadlock patterns
    • access order and index design
    • retry logic and where it belongs
    • verifying the fix
    • SQL Server
    • Concurrency
    • Debugging
  • PrévuMicrosoft FabricAvancé

    Troubleshooting Fabric Capacity Spikes in Production

    A workflow for finding which item and operation caused a capacity spike, and for preventing the next one.

    Sujets prévus

    When capacity is throttled, every user notices. Finding the cause quickly depends on knowing where to look and what the metrics mean.

    • recognising a spike vs sustained load
    • drilling down in Capacity Metrics
    • common causes
    • scheduling and concurrency of jobs
    • short-term mitigation
    • long-term prevention
    • Microsoft Fabric
    • Capacity
    • Debugging
    • Observability
  • PrévuMicrosoft FabricIntermédiaire

    What I Monitor in a Production Microsoft Fabric Platform

    The signals worth tracking on a production Fabric platform: pipeline outcomes, data freshness, capacity, table health and cost.

    Sujets prévus

    Without deliberate monitoring, problems are found by report users. The article lists what to watch and how to alert on it without noise.

    • pipeline success and duration
    • data freshness
    • capacity usage
    • Delta table health
    • cost trends
    • alerting without noise
    • Microsoft Fabric
    • Observability
    • Logging
  • PrévuDevOps & ObservabilityIntermédiaire

    Building Reliable Pipeline Execution Logging in a Data Platform

    Designing a run log for data pipelines: what to record per run and step, where to store it, and how to make it useful for debugging and reporting.

    Sujets prévus

    Built-in run history is often not enough to answer what happened to a specific batch. A deliberate execution log makes that question easy.

    • run and step identifiers
    • what to record
    • where to store logs
    • logging failures reliably
    • row counts and data volumes
    • reports on top of the log
    • Logging
    • Observability
    • Data Pipelines
  • PrévuDevOps & ObservabilityIntermédiaire

    How I Debug Failed Data Pipelines in Production

    A step-by-step approach to failed pipeline runs: triage, finding the failing step, reproducing safely, fixing and rerunning without side effects.

    Sujets prévus

    Pipeline failures need a calm, repeatable process: understand the impact, find the cause, fix it and rerun safely.

    • triage and impact
    • finding the failing step
    • reading error messages
    • reproducing safely
    • rerunning without duplicates
    • post-incident notes
    • Debugging
    • Data Pipelines
    • Observability

Data Platform Monitoring & Observability15 prévus

  • PrévuDevOps & ObservabilityIntermédiaire

    Finding Which Client Is Generating the Most Load

    A planned practical guide to finding which client is generating the most load in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    Monitoring Query Duration, CPU, Reads and Writes

    A planned practical guide to monitoring query duration, cpu, reads and writes in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    Using P50, P95 and P99 for Data Platform Performance

    A planned practical guide to using p50, p95 and p99 for data platform performance in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    Designing Correlation IDs Across APIs, Reports and Data Pipelines

    A planned practical guide to designing correlation ids across apis, reports and data pipelines in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    Detecting Abnormal Request Spikes

    A planned practical guide to detecting abnormal request spikes in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    Monitoring Cache Hit Rate in a Report Engine

    A planned practical guide to monitoring cache hit rate in a report engine in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    Building Operational Dashboards for Data Platforms

    A planned practical guide to building operational dashboards for data platforms in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    Alerting Without Creating Alert Fatigue

    A planned practical guide to alerting without creating alert fatigue in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    Tracking Pipeline Failures, Retries and Recovery

    A planned practical guide to tracking pipeline failures, retries and recovery in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    Monitoring Data Freshness and SLA Compliance

    A planned practical guide to monitoring data freshness and sla compliance in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    Monitoring Microsoft Fabric Capacity and Workload Pressure

    A planned practical guide to monitoring microsoft fabric capacity and workload pressure in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    Monitoring SQL Blocking and Long-Running Transactions

    A planned practical guide to monitoring sql blocking and long-running transactions in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    Designing a Multi-Tenant Monitoring Model

    A planned practical guide to designing a multi-tenant monitoring model in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    Building an Audit Trail for Data Requests

    A planned practical guide to building an audit trail for data requests in production data platforms.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance
  • PrévuDevOps & ObservabilityIntermédiaire

    From Logs to Root Cause: A Production Troubleshooting Workflow

    A planned practical guide from operational logs to root cause through a production troubleshooting workflow.

    • Observability
    • Monitoring
    • Data Architecture
    • Performance

Building a Modern Report Engine26 prévus

  • PrévuCloud & AutomationIntermédiaire

    Querying Fabric from a Report Engine

    A planned production guide to querying fabric from a report engine for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Returning CSV, JSON and Other Report Formats

    A planned production guide to returning csv, json and other report formats for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Caching Report Results in a Lakehouse

    A planned production guide to caching report results in a lakehouse for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Designing Report Cache Keys

    A planned production guide to designing report cache keys for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    When Should a Report Be Cached?

    A planned production guide to when should a report be cached? for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Building Asynchronous Reports for Large Requests

    A planned production guide to building asynchronous reports for large requests for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Report Engine Security and Tenant Isolation

    A planned production guide to report engine security and tenant isolation for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Rate Limiting and Throttling Report Requests

    A planned production guide to rate limiting and throttling report requests for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Designing Report Request Parameters

    A planned production guide to designing report request parameters for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Validating Report Requests Before Querying Data

    A planned production guide to validating report requests before querying data for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    UI Data Access vs Report Downloads

    A planned production guide to ui data access vs report downloads for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Pagination for Interactive Reports

    A planned production guide to pagination for interactive reports for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Streaming Large Report Results

    A planned production guide to streaming large report results for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Monitoring Report Engine Performance

    A planned production guide to monitoring report engine performance for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Designing Custom Reports Safely

    A planned production guide to designing custom reports safely for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Report Engine Metadata and Configuration

    A planned production guide to report engine metadata and configuration for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Versioning Report Definitions

    A planned production guide to versioning report definitions for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Auditing Report Access

    A planned production guide to auditing report access for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Handling Duplicate Report Requests

    A planned production guide to handling duplicate report requests for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Retry and Failure Patterns

    A planned production guide to retry and failure patterns for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Building a Report Engine with .NET or Python

    A planned production guide to building a report engine with .net or python for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    App Service vs Azure Functions for a Report Engine

    A planned production guide to app service vs azure functions for a report engine for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    API Management in Front of a Report Engine

    A planned production guide to api management in front of a report engine for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Using Fabric Warehouse and Lakehouse Together

    A planned production guide to using fabric warehouse and lakehouse together for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    Multi-Tenant Report Engine Architecture

    A planned production guide to multi-tenant report engine architecture for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy
  • PrévuCloud & AutomationIntermédiaire

    End-to-End Report Engine Reference Architecture

    A planned production guide to end-to-end report engine reference architecture for secure analytical data access.

    • Data Architecture
    • API
    • Reporting
    • Multi-Tenancy

Python for Data Platforms1 prévus

  • PrévuCloud & AutomationIntermédiaire

    Using Python and Azure Functions to Build a Data Report Engine

    Architecture and implementation notes for a report engine built with Python and Azure Functions: triggers, data access, rendering, scaling and limits.

    Sujets prévus

    Report generation is a common job that does not need a full application server. Azure Functions fits well if its limits are respected.

    • architecture overview
    • triggers: HTTP, queue and timer
    • data access patterns
    • rendering outputs
    • timeouts, memory and scaling
    • deployment and configuration
    • Python
    • Azure Functions
    • API
    • Azure

Python DataFrames for Data Engineering11 prévus

  • PrévuPythonDébutant

    Selecting Columns in pandas

    Select, rename and organize DataFrame columns as an explicit transformation contract rather than carrying an uncontrolled schema forward.

    Sujets prévus

    Column selection defines the output contract, reduces unnecessary memory, and prevents accidental propagation of sensitive or unstable fields.

    • Selecting one or several columns safely
    • `.loc`, column ordering, renaming and missing-column validation
    • Views, copies and assignment behavior
    • Stable schemas for downstream pipelines
    • Python
    • pandas
    • DataFrames
  • PrévuPythonDébutant

    Sorting DataFrames

    Sort pandas data predictably by one or more columns while handling ties, missing values and stable ordering.

    Sujets prévus

    Sorting affects ranking, deduplication, window-style logic and reproducible output, so tie and null behavior must be intentional.

    • Single and multi-column sorting
    • Ascending directions, null placement and stable algorithms
    • Sorting dates and numeric values after type validation
    • Why display order is not a storage guarantee
    • Python
    • pandas
    • DataFrames
  • PrévuPythonDébutant

    Handling Missing Values

    Detect and handle missing pandas values according to business meaning instead of applying one blanket fill or drop rule.

    Sujets prévus

    Missing can mean unknown, not applicable, anonymous, delayed, or invalid. Each meaning requires different validation and transformation behavior.

    • `isna`, nullable dtypes, counts and percentages
    • `fillna`, `dropna`, propagation and conditional rules
    • Anonymous website events versus damaged customer records
    • Measuring completeness before and after transformation
    • Python
    • pandas
    • DataFrames
    • Dataset
  • PrévuPythonDébutant

    Creating Calculated Columns

    Create typed, testable pandas columns for revenue, margins, flags and date features without fragile chained assignment.

    Sujets prévus

    Calculated columns encode business rules. Clear names, typed inputs, null decisions and reconciliation make those rules reviewable.

    • Vectorized arithmetic and conditional columns
    • Revenue, margin and operational flags
    • `.assign`, `.loc` and avoiding chained assignment
    • Validation and numeric precision choices
    • Python
    • pandas
    • DataFrames
  • PrévuPythonIntermédiaire

    Working with Dates and Timestamps

    Parse, validate and transform dates and timestamps in pandas with explicit timezone and invalid-value behavior.

    Sujets prévus

    Date logic fails quietly when strings, timezones, boundaries and invalid values are left implicit.

    • `to_datetime`, parsing failures and UTC timestamps
    • Date ranges, components, periods and time-based grouping
    • Signup dates, order dates and website events
    • Timezone normalization and inclusive boundary decisions
    • Python
    • pandas
    • DataFrames
    • Dataset
  • PrévuPythonIntermédiaire

    Removing Duplicate Rows

    Distinguish exact duplicates from duplicate business keys and choose deterministic pandas survivorship rules.

    Sujets prévus

    Removing duplicates without defining grain and survivorship can discard valid events or keep an arbitrary conflicting record.

    • Exact rows versus duplicate keys
    • `duplicated`, `drop_duplicates`, subsets and keep rules
    • Sort-before-deduplicate patterns
    • Auditing removed records and conflicting attributes
    • Python
    • pandas
    • DataFrames
    • Dataset
  • PrévuPythonIntermédiaire

    Pivot Tables and Reshaping

    Reshape pandas data between long and wide forms while controlling aggregation, missing combinations and duplicate keys.

    Sujets prévus

    Reshaping changes the grain and often aggregates implicitly, so duplicate key combinations and missing cells need explicit treatment.

    • `pivot`, `pivot_table`, `melt` and stacking concepts
    • Revenue by month and category
    • Duplicate combinations and aggregation functions
    • Returning report-shaped data to analysis-friendly long form
    • Python
    • pandas
    • DataFrames
  • PrévuPythonIntermédiaire

    Reading Large CSV Files Efficiently

    Reduce pandas CSV memory and processing cost through column pruning, explicit dtypes, chunking and better storage formats.

    Sujets prévus

    CSV parsing is CPU- and memory-intensive, and a file that fits on disk can expand substantially in a DataFrame.

    • `usecols`, dtypes, categorical columns and memory measurement
    • Chunked aggregation without concatenating everything again
    • Compression, parsing tradeoffs and error handling
    • When Parquet, a database, or PySpark is the better boundary
    • Python
    • pandas
    • DataFrames
    • CSV
    • Performance
    • Large Scale Data
  • PrévuPythonIntermédiaire

    Writing CSV and Parquet Files

    Write pandas outputs with deliberate schemas, indexes, partitions and validation for downstream consumers.

    Sujets prévus

    Output format decisions affect types, file size, interoperability and whether a downstream reader can reproduce the intended schema.

    • CSV index, quoting, null and encoding options
    • Parquet schemas, compression and timestamps
    • Atomic output and post-write validation
    • File naming, partitioning and consumer contracts
    • Python
    • pandas
    • DataFrames
    • CSV
    • Parquet
  • PrévuPythonIntermédiaire

    pandas vs PySpark DataFrames

    Compare pandas and PySpark DataFrames through execution model, scale, APIs, null semantics and practical migration choices.

    Sujets prévus

    Similar-looking APIs hide very different local and distributed execution models. Dataset size alone does not decide when Spark is justified.

    • Eager versus lazy execution and local versus distributed memory
    • Equivalent filters, groups and joins
    • Schema, null and ordering differences
    • Operational cost, debugging, scale and migration boundaries
    • Python
    • pandas
    • PySpark
    • DataFrames
    • Large Scale Data
  • PrévuPythonIntermédiaire

    Common DataFrame Problems and How to Debug Them

    Diagnose pandas schema drift, dtype surprises, nulls, duplicate keys, bad joins, memory pressure and misleading aggregates.

    Sujets prévus

    Most DataFrame failures become easier when the pipeline is reduced to grain, schema, counts, key quality and one reproducible transformation.

    • Schema and dtype drift
    • Null propagation, chained assignment and index alignment
    • Join multiplication and lost rows
    • Memory diagnostics, minimal reproductions and validation checkpoints
    • Python
    • pandas
    • DataFrames
    • Troubleshooting

Cloud Automation1 prévus

  • PrévuCloud & AutomationIntermédiaire

    Automating Cloud Data Infrastructure with SDKs and Infrastructure as Code

    Using infrastructure as code tools and cloud SDKs together to provision and manage data platform resources repeatably.

    Sujets prévus

    Data platforms have resources that IaC tools cover well and others that need SDK calls. A clear split keeps automation maintainable.

    • declarative IaC vs imperative SDKs
    • choosing a tool
    • resources IaC covers well
    • filling gaps with SDK scripts
    • secrets and identities
    • testing and drift detection
    • Infrastructure as Code
    • SDK
    • Cloud Architecture
    • Python