Skip to content

Spark

Enable Spark integration with spark=True or --spark; it defaults to False. LineScope uses the environment's existing PySpark and does not install Spark. Requesting integration without PySpark raises an error at startup. Observation is lazy: LineScope does not import PySpark or inspect an active Spark session at profile startup. It watches Spark modules loaded by the workload and wraps supported methods, including those already loaded before profiling. Lineage is captured when those methods are called; the JVM query listener starts at the first observed action. Profiles without Spark calls leave the listener dormant.

Use profile(spark=False) or --no-spark to disable integration. The "auto" option is no longer accepted. The profiler observes existing action calls; it never inserts a count, collect, or other action to measure a lazy transformation, and it never creates a Spark session.

from linescope import profile

with profile(backend="trace", spark=True, display="none") as session:
    values = spark.range(1000).filter("id > 100")
    total = values.groupBy().sum("id").collect()

session.save("spark.html")

Lazy execution

filter and groupBy build a plan. The final collect triggers work. Driver line timings preserve that behavior: the action usually contains the long wait. Spark references can explain the transformations participating in that action, but they do not assign fake wall-clock time to lazy source lines.

Worker scope and overall time

Python source profiling covers the driver according to the selected backend's thread coverage. Spark executor JVMs and Python UDF worker processes are outside that source profile, including with local[2]. Setting PYSPARK_PYTHON selects a worker interpreter; it does not enable worker profiling. Spark actions are observed on the thread that starts the integration.

Worker execution still affects the result. The driver's collect() or count() line includes its wait for the action. The Spark view can also show cumulative executor/task time and operator metrics supplied by Spark. These counters do not provide line timings or Python allocation measurements inside a UDF worker.

For example, an action with two overlapping tasks that each take five seconds could have about five seconds of driver waiting and ten seconds of cumulative task time, plus scheduling and transfer overhead. Elapsed wall time remains the duration of the profiling session. Adding driver waiting to cumulative task time would count overlapping work twice. Missing task metrics do not imply that workers were idle; they mean the runtime did not expose sufficient information.

Plans and metrics

The Spark view starts with summary cards for maximum action wall time, maximum operator time, peak operator memory, maximum disk spill, and captured Spark jobs. The job count totals the job records associated with captured actions, including child notebook runs. An action can have several job records. The sortable action table follows these summaries. Each row shows action wall time and the largest reported operator or shared pipeline time, peak memory, and disk spill. The longest measured action appears first, and actions from child notebook runs are included. Select a row or its action link to open the action's detail page and main-step plan overview. Source links open the captured trigger line instead. The overview uses Operator time, Peak memory, and Disk spill for the reported plan costs. The All operator costs table below it uses the same column names. Cumulative executor time and executor peak memory remain available on each action's detail page. Select a column heading to order the table. Operator rankings and plan-step tables use the same header arrows. Select the heading again to reverse the direction. Unknown measurements sort after measured values, including zero, in both directions.

Main plan steps follows the data from inputs to the result. Plain-language labels describe reading, filtering, joining, summarizing, sorting, and moving data between workers. Unmeasured projections and internal wrappers are omitted. The From step column keeps separate branches explicit: 2 + 5 means that this operation consumes both inputs, not that those inputs ran sequentially. Separate inputs can run in parallel. Reported time, peak memory, and total Rows after appear beside each step.

The detail page combines action and operator metrics in one overview above the Main plan steps table. Operator time, Peak memory, and Disk spill identify the largest reported operator or shared pipeline costs, with the responsible step shown below each value. Row multiplication at joins appears when both input counts are known. Operator counters have a different scope from action wall time, cumulative executor time, and executor peak memory. They point to measured work to investigate; partial counters cannot establish the complete cause of a slow action. Time bars compare measured individual steps rather than percentages of the action's elapsed time.

Rows use Spark's output-row counters. Sorts and data exchanges can carry a known input count forward because they preserve rows; those values are labeled From input ยท unchanged. Missing output counts after filters, joins, aggregates, limits, and unknown operations remain unavailable. Estimates, partition counts, and shuffle-record counters do not fill missing row counts.

Each action has a Source link showing its captured filename and exact action line. Select it to open and highlight the line that triggered the execution. The same link appears on the execution's detail page. Notebook actions link to their captured cell and line. When the trigger or its source snapshot is unavailable, the report shows Trigger source unavailable.

The captured action line appears beside the link so repeated collect or count calls are easy to distinguish. Select an action to see its cost summary and main-step overview. Additional action metrics appear only when row, read, shuffle, or spill measurements are available. Select a step to open its action's physical operator tree, expand and highlight that exact step, and show its description and metrics. Query plans, extra action metrics, stage details, and the full operator cost ranking are collapsed below the overview. Runtime limitations and plan provenance appear under Plan provenance and collection notes on each action's detail page when available.

Query plans prefer the actual executed physical plan after the action completes, including the final adaptive plan when accessible. Initial and optimized logical plans provide context. Operator details expose available rows, bytes, shuffle, memory, spill, and time metrics.

Operator metrics are supplied by Spark. Operator time shows the largest available Spark timing on each node, converted to seconds. The selected metric's name appears below the main step's time and in its tooltip. Timings can measure cumulative worker work, preparation, or waiting for upstream input; they are not exclusive durations for individual steps and do not add up to action wall time. Fused pipeline timings can overlap their children, so the ranking never sums them into an action total or assigns a fused pipeline's time to its children. Measured fused pipelines appear under Operations measured together, with the main step numbers they cover. A step with no separate timing shows Shared timing when its pipeline has a measurement. Input adapters end pipeline membership, so upstream work is not assigned to a downstream fused group. Peak memory uses an explicit peak-memory counter; shuffle data size and spill do not stand in for memory usage. Peak memory does not represent the total data size processed by the step.

Metric Meaning
Action wall time Elapsed time waiting on the action from the driver
Executor/task time Cumulative work across parallel tasks, when exposed
Operator rows/bytes Spark-reported runtime counters
Peak memory/spill Operator/executor metrics, separate from Python memory

Cumulative executor and operator time can exceed wall time because tasks run in parallel. Missing metrics remain unknown and appear as a dash in the report. Spark Connect, runtime access restrictions, JVM API changes, and Databricks policies can limit plan access; the report keeps collected information and records limitations.

When the runtime's status store permits access, stage details include submission-to-completion wall time, cumulative executor time, I/O/shuffle/spill bytes, and task-duration p50/p95/maximum. Stage wall time includes scheduling. Action executor totals sum unique completed stage attempts whose submission and completion fall inside the observed action. Previously completed or skipped dependencies are excluded. Missing, overlapping, or concurrently observed work leaves the total unknown, and SQL operator timings are never added together to produce the action total.

Scope

Interface Version 0.1.0 support
Classic DataFrame and writers Supported actions and executed plans
Driver Python Source profiling through the selected backend
Executor Python UDF workers Separate worker processes are not collected
RDD, streaming, Spark Connect, ML Outside automatic observation
Deferred iterators and eager SQL Not comprehensively observed

External integrations can use SparkIntegration.record_action to provide an explicit observation boundary around an operation. Source-to-plan links remain explanatory: Catalyst and AQE can reorder, fuse, or eliminate operations, and one lineage can participate in multiple actions. These plans do not imply exact per-line Spark runtimes.

For a real integration test without an external cluster, use local Spark notebook.

Where a runtime permits a query-execution listener, the observer can capture the final action plan, including writes. If listener access is restricted, a writer may only expose the input DataFrame plan; that fallback is labeled instead of presented as the final write execution.