Skip to content

Backends


The backend determines how line time is collected, which metrics are available, and how much instrumentation affects the workload. Choose it with backend=... in the Python API or --backend on the CLI.

Choose a backend

Compare the features of LineScope's built-in adapters below. ✓ means supported; ✗ means unavailable.

Feature Scalene Trace Tachyon
Python versions 3.11–3.14 3.11–3.15 3.15
Collection Sampling Tracing External sampling
Line and function time ✓ ✓ ✓
Recorded line hits ✗ ✓ ✗
Recorded function calls ✗ ✓ ✗
Per-hit average time ✗ ✓ ✗
Observed sample counts ✓ ✗ ✓
Retained Python allocation delta ✓ ✓ ✓
Per-line Python allocation peak ✗ ✗ ✗
Process RAM, changes, observed peak ✓ ✓ ✓
Process RAM timeline ✓ ✓ ✓
GPU time and memory ✓ ✗ ✗
Native/library cost on project lines ✓ ✓ ✓
Multiple Python threads ✓ ✗ ✗
Automatic worker-process profiling ✗ ✗ ✗
Configurable sampling rate ✓ ✗ ✓
Best suited to Bottleneck searches Small runs and tests 3.15 sampling

Enable the shared memory collector with memory=True on any built-in backend. GPU collection uses gpu=True with Scalene and requires a supported platform and device. See memory and GPU for their scope. The thread coverage row describes timing and execution counts; memory covers the profiled process. Custom backend features depend on their declared BackendCapabilities.

Choose Scalene for representative workloads or GPU investigations. Choose Trace for short investigations, recorded execution counts, and deterministic tests. Choose Tachyon for sampling on Python 3.15. Sampling can miss brief operations; Trace's wall intervals include tracing overhead. See performance for the tradeoffs.

Trace is the default on every supported Python version (3.11–3.15). Scalene is available through linescope[scalene] or linescope[full] on Python 3.11–3.14. Collectors must be supported by the runtime; an unavailable collector fails explicitly instead of silently changing backends. See GPU for device collection support.

Scalene

backend="scalene" selects the sampling collector on Python 3.11–3.14. Install it with uv pip install "linescope[scalene]". It collects line time and optional driver memory or GPU metrics. Sampling can miss short lines; hits and averages remain unavailable. Native and library work stays on the nearest project line. See memory and GPU for optional collection.

Trace

backend="trace" uses Python's built-in tracing hooks on Python 3.11–3.15. It records current-thread line events, hit counts and wall intervals, making it useful for small investigations, tests and portable environments. Tracing adds overhead and shares hooks with debuggers. Enable memory=True for process RAM and retained Python allocation changes; no native preload is needed. Python allocation peaks and GPU measurements are unavailable. Requesting gpu=True emits a RuntimeWarning and records it in session.result.warnings while Python profiling continues. Select backend="scalene" to collect supported GPU measurements.

Tachyon

backend="tachyon" uses Python 3.15's profiling.sampling stack sampler in a separate collector process. It estimates line time from samples and attributes third-party frames to their nearest project caller. Python 3.15 defaults to Trace; select Tachyon explicitly for sampling. Tachyon supplies no line hit counts or GPU measurements. memory=True adds the shared memory collector. Operating-system restrictions on reading process memory can prevent attachment; failures provide diagnostics instead of changing backends.

Profiling scope

The source boundary (root, include, and exclude) selects project files for collection and display. It does not attach the profiler to additional worker processes.

Backend Python thread coverage Separate worker processes
Scalene Samples threads in the profiled process Not collected
Trace Thread that starts the session Not collected
Tachyon Thread that starts the session Not collected
Custom Defined by the registered backend Defined by that backend

Scalene attributes native library work to the relevant project caller where available. Its line estimates are not individual worker CPU counters. Trace and Tachyon can show time waiting for a thread or process pool on the calling line, even though those workers' Python lines are outside their collection scope.

Elapsed wall time describes the session's duration. It includes waits inside the session, but does not represent the sum of CPU time consumed by every worker. For example, if four processes run concurrently while the parent waits five seconds for results, the parent can show about five seconds of waiting without reporting each process's work. Worker activity can continue after collection if the application does not wait for it before leaving the session.

In Spark, the profiled Python process is the driver. Executor JVMs and Python UDF workers contribute only the separate Spark runtime metrics that Spark exposes. GPU kernels also have separate measurements; synchronizing within the session keeps queued device work inside the observation period.

Performance

Sampling rate

Set sample_rate to a positive integer to configure the target number of samples per second for Scalene or Tachyon:

from linescope import profile

with profile(backend="scalene", sample_rate=250) as session:
    run_pipeline()

For scripts, use linescope --sample-rate 250 application.py. You can also set sample_rate = 250 in [tool.linescope] or pass sample_rate=250 to profile. When omitted or set to None in Python, both backends default to 1000 samples per second. Trace ignores this setting.

This is a target rate: operating-system timer resolution and scheduling settings, stack capture, and workload behavior affect the actual number of observations. Scalene randomizes intervals on POSIX and uses fixed intervals on Windows. Higher rates can capture more short lines but increase profiling overhead. The report shows observed sample counts and measured samples per second separately from hits or function calls.

Collection overhead

Trace observes line events and records wall intervals while the workload runs. Its per-event bookkeeping adds overhead, particularly in Python loops with many short lines. Hit counts help explain how often a line runs, but the measured intervals include the effects of instrumentation.

Scalene and Tachyon estimate line time from periodic samples. Sampling avoids recording every line event, but short runs or brief operations may receive too few samples to represent their cost. Use a representative workload that runs long enough to expose repeated expensive work. Samples never become hit counts or per-hit averages.

Collector startup, source discovery, notebook hooks, Spark observation, and report generation also contribute to the overall cost of a profiling run. Memory collection adds allocation tracking, RAM observations, and line-boundary tracing, including when timing uses a sampling backend. Live notebook reports with display="cell" repeat rendering after each cell; use display="end" for one final display or display="none" to save explicitly after collection. Use display="cell-summary" for a smaller inline overview of each cell's own measurements. It still takes collector snapshots at cell boundaries; memory allocation snapshots can add substantial work.

Compare results using the same backend, options, inputs, and environment. Check an optimization with separate unprofiled runs as well: profiler timings help locate costs, while instrumentation can change the workload's execution time. LineScope does not promise a fixed overhead percentage for any backend.

Memory

Enable process RAM and retained Python allocation collection with profile(memory=True) or --memory. This works with every timing backend. No native allocator preload or notebook kernel restart is required.

from linescope import profile

with profile(backend="trace", memory=True, display="none") as session:
    values = [bytearray(1024) for _ in range(10_000)]

session.save("memory.html")

Python allocation changes

The shared memory collector uses tracemalloc to compare snapshots at startup, at stop, and on live report requests. Per-line memory.delta_bytes in the collected profile data records the change in tracked bytes still allocated at each allocation site. Library allocations go to the nearest project frame in their captured traceback. Freed bytes belong to the original allocation site; allocations created and freed between snapshots leave no retained change. Per-line Python allocation peaks remain unavailable.

These snapshots include all Python threads, even when timing and hits cover only the starting thread. They exclude untracked native allocations, GPU memory, and separate processes. RAM and allocation deltas have different meanings and must not be added together.

New allocation tracers capture up to 25 frames and are stopped during cleanup. Existing tracers keep their depth, history, and peak and remain running. A shallow existing traceback can prevent library attribution to a project line. If the workload stops allocation tracing, values remain unknown and a warning explains the missing data. The CLI retains script/module globals through the final snapshot so their allocations can be reported.

RAM readings and growth

LineScope observes resident process memory (RSS), the memory currently in RAM. On Windows this is the process working set. It includes Python objects, native library buffers, shared pages, and profiler overhead. It excludes swapped-out pages, separate worker processes, and GPU memory.

The memory observer traces the starting thread's project line boundaries, even when the timing backend uses sampling. It reads RAM after each completed line interval and attributes external calls to their nearest project caller. Project child calls have their own intervals rather than charging the same increase to both caller and child. Other threads can still change process RAM, so these observations identify where growth occurred, not allocation ownership.

The source columns show Mem Change, the accumulated process-memory change over all intervals of that line, and Peak Mem, the highest reading during its intervals. Zero change is a measured value; unexecuted lines and failed readings stay unavailable. Repeated loop iterations keep their history in the Memory view, rather than being reduced to the final line reading alone.

A background observer also samples every 10 ms during long calls when the scheduler and GIL permit. Peaks are observed rather than guaranteed: brief spikes between readings, or native calls holding the GIL, can escape sampling. Tracing and RAM reads add overhead. Python and native allocators can retain freed memory for reuse, so deleting an object need not immediately reduce RSS.

Each run keeps at most 4096 timeline readings. When compression is needed, chronological bucket peaks and troughs, missing readings, and the first and last readings are retained. A notice explains compression in the report. Source-line RAM statistics continue to incorporate every observation.

Driver and executor memory

RAM columns concern the profiled Python process. In Spark that is normally the Python driver, not its JVM or executors. A lazy DataFrame can describe a large distributed dataset while occupying little Python RAM. Conversely, toPandas() may move substantial data into the driver.

Child notebook processes have separate timelines. Their RAM is never added to parent RAM. If a source line has RAM measurements from multiple runs, the shared source page retains the largest observed peak and leaves its combined Mem Change unavailable. Open each run's Memory timeline for its own readings.

Spark operator peak memory and spill, when available, appear with the Spark execution. They have different semantics and must not be added to process RAM.

Custom backends

Implement ProfilerBackend with start(), stop() and result(), declare BackendCapabilities, and register your factory with register_backend. Return RawBackendResult measurements independently of source discovery and rendering. Keep unavailable metrics as None, restore resources after failures, and use sampling counts only to estimate time.