stormlog.infer.trace_torch
Bounded in-process PyTorch profiler capture for import with import-trace.
capture_torch_trace profiles the code inside its block and writes a Kineto
Chrome trace when the block ends, including when it raises. The block is the
bound: there is no step or time limit inside it. Stack, shape, and
memory recording are off unless asked for, since each adds CPU work per
operator. The profiler adds no synchronization per step; stopping it flushes
the CUDA activity buffers once, at the end of the block.
Functions
|
Profile the block and write its trace to |
Exceptions
Another PyTorch profiler is already running in this process. |
- exception stormlog.infer.trace_torch.ProfilerBusyError[source]
Bases:
InferUsageErrorAnother PyTorch profiler is already running in this process.
- stormlog.infer.trace_torch.capture_torch_trace(output, *, cuda=None, with_stack=False, record_shapes=False, profile_memory=False)[source]
Profile the block and write its trace to
output.cudadefaults to whether CUDA is available. Wrap each iteration instormlog.infer.trace_ranges.iteration_rangeso the import can link GPU work to iterations.- Parameters:
output (str | Path)
cuda (bool | None)
with_stack (bool)
record_shapes (bool)
profile_memory (bool)
- Return type:
Iterator[Path]