GPU workload¶
This notebook demonstrates preparing signals on the CPU, transferring them to an NVIDIA GPU, and applying matrix projections with ReLU activation. The Scalene backend collects estimated GPU time and peak device memory separately from Python driver time. The outputs below were captured on a real CUDA device.
from time import perf_counter
import torch
from linescope import profile
if not torch.cuda.is_available():
raise RuntimeError("This notebook requires an NVIDIA GPU and CUDA-enabled PyTorch")
# Initialize the device before collecting workload measurements.
torch.cuda.init()
print(f"GPU: {torch.cuda.get_device_name()}")
GPU: NVIDIA GeForce RTX 5060
Prepare and project signals¶
Use reproducible input and synchronize after each projection so device work finishes within the profiling session. Scalene samples Python driver time, estimated GPU time, and peak device memory. Driver memory collection is disabled for this example. Device metrics depend on the available runtime; unavailable measurements remain unknown. NVIDIA WDDM drivers on Windows do not expose process GPU memory, so those values appear as unavailable while GPU time remains visible.
def load_signals() -> tuple[torch.Tensor, torch.Tensor]:
"""Prepare reproducible signals and transfer them to the GPU.
Generate input on the CPU so host preparation and device transfers appear
separately from the repeated GPU calculations.
Returns
-------
tuple[torch.Tensor, torch.Tensor]
Signal and projection-weight tensors stored on the CUDA device.
"""
generator = torch.Generator().manual_seed(42)
signals = torch.randn(2048, 1024, generator=generator)
weights = torch.randn(1024, 1024, generator=generator) / 32
return signals.to("cuda"), weights.to("cuda")
def project_signals(signals: torch.Tensor, weights: torch.Tensor) -> tuple[torch.Tensor, int]:
"""Project signals repeatedly to sample driver and device work.
Synchronize each iteration to finish queued device work inside the
profiling session. Scalene collects driver time and separate estimated
GPU time and peak device memory. The iteration count depends on the device.
Parameters
----------
signals : torch.Tensor
Input signals already transferred to the CUDA device.
weights : torch.Tensor
Projection weights stored on the same device as the signals.
Returns
-------
tuple[torch.Tensor, int]
Final projected signals and the number of completed projections.
"""
deadline = perf_counter() + 3
iterations = 0
while True:
projected = torch.relu(signals @ weights)
torch.cuda.synchronize()
iterations += 1
if perf_counter() >= deadline:
return projected, iterations
session = profile.start(backend="scalene", gpu=True, memory=False, spark=False, display="end")
signals, weights = load_signals()
projected, iterations = project_signals(signals, weights)
mean = projected.mean().item()
print(f"Projections: {iterations}; mean: {mean:.4f}")
Projections: 5131; mean: 0.3987
Inspect the GPU measurements¶
The report opens on Overview with Samples and Measured samples / sec cards. Select GPU in the sidebar to see Attributed GPU time and GPU peak memory cards and compare GPU time with Driver time for each measured source line. Select a line's location to inspect the captured code; the source table also includes the device columns. The printed summary below exposes the additional measurements directly.
Change gpu=True to gpu=False in the profiling call and run the notebook
again to compare collection modes. The same calculations still execute on
CUDA, but the device cards and GPU view disappear and only Python driver
measurements remain. GPU time and driver time can overlap; do not add them.
profile.stop()
result = profile.result
gpu_times = [
line.gpu.time_ns
for line in result.root_run.lines
if line.gpu is not None and line.gpu.time_ns is not None
]
gpu_peaks = [
line.gpu.peak_memory_bytes
for line in result.root_run.lines
if line.gpu is not None and line.gpu.peak_memory_bytes is not None
]
gpu_time = f"{sum(gpu_times) / 1_000_000_000:.3f} s" if gpu_times else "unavailable"
gpu_memory = f"{max(gpu_peaks) / 1024**2:.1f} MiB" if gpu_peaks else "unavailable"
print(f"Attributed GPU time: {gpu_time}; GPU peak memory: {gpu_memory}")
print("Select GPU in the report to compare device estimates with Python driver time.")
for warning in result.warnings:
if "GPU" in warning:
print(f"GPU collection note: {warning}")
report_path = session.save("gpu.html")
NOTE: The GPU is currently running in a mode that can reduce Scalene's accuracy when reporting GPU utilization. If you have sudo privileges, you can run this command (Linux only) to enable per-process GPU accounting: python3 -m scalene.set_nvidia_gpu_modes
Attributed GPU time: 2.425 s; GPU peak memory: unavailable Select GPU in the report to compare device estimates with Python driver time. GPU collection note: Per-process GPU memory is unavailable with NVIDIA WDDM.