stormlog.infer.trace_torch

Bounded in-process PyTorch profiler capture for import with import-trace.

capture_torch_trace profiles the code inside its block and writes a Kineto Chrome trace when the block ends, including when it raises. The block is the bound: there is no step or time limit inside it. Stack, shape, and memory recording are off unless asked for, since each adds CPU work per operator. The profiler adds no synchronization per step; stopping it flushes the CUDA activity buffers once, at the end of the block.

Functions

capture_torch_trace(output, *[, cuda, ...])

Profile the block and write its trace to output.

Exceptions

ProfilerBusyError

Another PyTorch profiler is already running in this process.

exception stormlog.infer.trace_torch.ProfilerBusyError[source]

Bases: InferUsageError

Another PyTorch profiler is already running in this process.

stormlog.infer.trace_torch.capture_torch_trace(output, *, cuda=None, with_stack=False, record_shapes=False, profile_memory=False)[source]

Profile the block and write its trace to output.

cuda defaults to whether CUDA is available. Wrap each iteration in stormlog.infer.trace_ranges.iteration_range so the import can link GPU work to iterations.

Parameters:
  • output (str | Path)

  • cuda (bool | None)

  • with_stack (bool)

  • record_shapes (bool)

  • profile_memory (bool)

Return type:

Iterator[Path]