Microsoft Fabric
Using Session Tags in Fabric Pipelines
Use Microsoft Fabric pipeline session tags to group compatible notebook activities for Spark session reuse while preserving workload boundaries.
On this page
Pipeline notebooks often perform short, related stages: ingest, validate, transform, and publish. Starting separate Spark applications for every activity can make session acquisition a meaningful part of the total run. Fabric high-concurrency mode can pack compatible notebook workloads into a shared Spark application, and a session tag gives the pipeline an explicit grouping signal.
Problem
A pipeline runs several notebooks under the same operational workflow. Each notebook is correct on its own, but repeated startup makes the end-to-end duration noisy. You want reuse without mixing unrelated workloads or assuming that a matching string overrides compatibility and isolation rules.
Why it happens
In standard execution, notebook activities can acquire separate Spark sessions. In high-concurrency mode, Fabric can host multiple notebook workloads in one Spark application, with separate REPL cores for execution state. Reusing an existing application avoids repeating some acquisition work for subsequent compatible notebooks.
Session tags are configured on notebook activities in the pipeline. Activities with the same tag become candidates for packing into the same high-concurrency session. According to Microsoft’s pipeline high-concurrency documentation, sharing remains within a single-user boundary and requires compatible workspace, default Lakehouse, compute configuration, and libraries. If requirements differ or capacity in the session is unavailable, Fabric creates another session. A tag is therefore a grouping hint within platform rules, not a command to merge incompatible execution contexts.
The workspace setting for high concurrency on pipeline notebook runs must be enabled before session tags can provide reuse. Administrators own that setting, while pipeline authors own the activity design and tag values.
Practical recommendations
Choose a tag that represents an execution compatibility group, not merely a business label. Notebooks in one daily sales pipeline may belong together if they use the same Lakehouse, environment, Spark settings, identity boundary, and overlapping schedule. A finance transformation and an experimental machine-learning notebook should not share a tag only because both run in the same workspace.
Keep tag construction deterministic. A stable value such as sales-daily-standard makes related activities candidates for reuse across the intended window. Adding a unique run identifier to every tag prevents cross-activity reuse because every activity presents a different value. Conversely, using one global tag for every pipeline creates an overly broad group and makes behavior harder to reason about.
Separate incompatible workloads deliberately. A notebook with different compute requirements, libraries, or default Lakehouse should use a different group even when it is part of the same business process. Isolation may also be more important than startup for memory-heavy or latency-critical jobs.
Measure at both levels. Record pipeline activity duration and notebook-stage duration. In high concurrency, monitoring provides notebook-related views and separated logs, which are essential when several notebooks share the underlying application. A shared session can improve acquisition while Spark work inside it remains slow or contended.
Example orchestration pattern
Consider a pipeline with four notebook activities:
| Activity | Purpose | Suggested tag |
|---|---|---|
| Validate input | Check schema and arrival | sales-daily-standard |
| Build Silver orders | Clean and conform | sales-daily-standard |
| Build Gold sales | Aggregate output | sales-daily-standard |
| Train forecast | Memory-heavy model training | sales-forecast-isolated |
The first three are candidates for the same high-concurrency group if their Fabric settings match. The training workload uses a separate tag because its resource behavior and isolation needs differ.
Pass a pipeline correlation identifier as a notebook parameter rather than encoding it into the session tag. That preserves a stable reuse group while keeping logs traceable.
# Parameter cell values supplied by the pipeline
pipeline_run_id = ""
processing_date = ""
print({
"pipeline_run_id": pipeline_run_id,
"processing_date": processing_date,
"stage": "build_silver_orders",
})
Within each notebook, avoid relying on variables created by another notebook. High concurrency provides separate REPL cores, and a reliable pipeline should communicate through parameters, durable tables, files, or explicit orchestration outputs. Session reuse is a compute optimization, not an implicit shared-variable contract.
When activities run concurrently, remember they share underlying resources even though their REPL state is separated. Test concurrency with representative data and watch whether one large shuffle or cache-heavy notebook degrades the others. Sequential activities may benefit from warm compute without the same resource competition, but the workflow dependency should be driven by data correctness.
Common mistakes
- Setting a tag without enabling the workspace’s pipeline high-concurrency option.
- Assuming identical tags can override different Lakehouses, compute settings, libraries, or user boundaries.
- Putting a unique run ID in each activity’s tag and accidentally disabling reuse.
- Giving every notebook in the workspace one tag.
- Depending on another notebook’s in-memory variables or cache as a pipeline interface.
- Combining resource-heavy concurrent activities without measuring contention.
- Looking only at end-to-end duration and not separating session acquisition from Spark work.
- Assuming one tag always maps to one physical session; capacity and platform packing rules can create additional sessions.
What to check next
Confirm the high-concurrency workspace setting, then inspect each activity’s default Lakehouse, environment, compute settings, libraries, identity, and tag. In monitoring, verify which sessions and notebook jobs were created and whether the expected activities were associated with the shared run.
If reuse does not occur, treat compatibility as the first hypothesis. If reuse occurs but the pipeline remains slow, examine concurrency, Spark stages, shuffle, I/O, and writes. If reliability declines under shared execution, reduce the group or isolate the conflicting workload. The correct session strategy depends on workload compatibility, resource behavior, and operational clarity—not on maximizing reuse at all costs.