stormlog.infer.trace_kineto

Import PyTorch/Kineto profiler traces as inference activity references.

A Kineto trace (PyTorch profiler, or vLLM’s rank*.pt.trace.json.gz) lists GPU kernels, copies, and memsets with device timestamps, and the CPU runtime or driver calls that launched them. Each GPU event carries the CUDA correlation ID of its launch call. A CUDA graph launch is one call and many GPU events.

This importer emits one infer.activity_ref per launch, device, and activity kind (detail="launch", the default) or per GPU event (detail="kernel"). A launch record spans its first event’s start to its last event’s end. When it holds several events, its metadata.intervals lists the exact busy intervals inside that span, so GPU time accounting does not count the idle gaps inside a CUDA graph replay as busy. On vLLM 0.30.0 traces a launch record set is about 6x smaller than a per-event one and gives the same device busy time. A GPU activity is linked to an iteration only when its launch call sits inside exactly one stormlog.iteration/... range on the launching thread. Otherwise it stays unresolved and says why. Launch calls are CPU work and are never emitted as GPU activity.

Functions

capture_trace(trace, path, *, run_id, session_id)

Build the capture for a loaded trace of any supported format.

import_kineto_trace(path, *, run_id, session_id)

Return activity references, capabilities, and a summary for one trace.

index_spans(trace)

Sort each thread's iteration spans and index them for launch lookups.

link_gpu_event(trace, event)

Link a GPU event to the one iteration range around its launch call.

load_kineto_trace(path)

Parse a plain or gzipped Kineto Chrome trace.

merge_intervals(intervals)

Merge intervals into disjoint, sorted ones; touching intervals join.

Classes

GpuEvent(start_ns, end_ns, kind, name, ...)

One kernel, copy, or memset as the trace recorded it.

GpuLink(iteration_ref, reason, launch)

An event's iteration link, or the reason there is none.

IterationSpan(start_us, end_us, iteration_ref)

KinetoTrace(base_ns, host, trace_id, rank, ...)

The parts of a profiler trace the importer uses (Kineto or Nsight).

LaunchCall(pid, tid, ts_us, name)

The CPU runtime or driver call that launched GPU work.

class stormlog.infer.trace_kineto.GpuEvent(start_ns, end_ns, kind, name, device, stream, correlation, graph_id, pid=None, device_uuid=None, device_name=None, graph_node_id=None)[source]

Bases: object

One kernel, copy, or memset as the trace recorded it.

Parameters:
  • start_ns (int)

  • end_ns (int)

  • kind (str)

  • name (str)

  • device (int | None)

  • stream (int | None)

  • correlation (int | None)

  • graph_id (int | None)

  • pid (int | None)

  • device_uuid (str | None)

  • device_name (str | None)

  • graph_node_id (int | None)

start_ns: int
end_ns: int
kind: str
name: str
device: int | None
stream: int | None
correlation: int | None
graph_id: int | None
pid: int | None = None
device_uuid: str | None = None
device_name: str | None = None
graph_node_id: int | None = None
class stormlog.infer.trace_kineto.IterationSpan(start_us: 'float', end_us: 'float', iteration_ref: 'EntityRef')[source]

Bases: object

Parameters:
  • start_us (float)

  • end_us (float)

  • iteration_ref (EntityRef)

start_us: float
end_us: float
iteration_ref: EntityRef
class stormlog.infer.trace_kineto.LaunchCall(pid, tid, ts_us, name)[source]

Bases: object

The CPU runtime or driver call that launched GPU work.

Parameters:
  • pid (int)

  • tid (int)

  • ts_us (float)

  • name (str)

pid: int
tid: int
ts_us: float
name: str
stormlog.infer.trace_kineto.capture_trace(trace, path, *, run_id, session_id, attachment=None, device_uuids=None, detail='launch', device_uuids_by_pid=None)[source]

Build the capture for a loaded trace of any supported format.

device_uuids (--device-uuid) names a CUDA ordinal’s GPU for every process in the trace and is checked against UUIDs the trace names itself. device_uuids_by_pid is what the vLLM execution log’s worker hellos bound: a process’s own ordinals only, so a process no hello matched stays unmeasured. The option wins where both name an ordinal.

Parameters:
  • trace (KinetoTrace)

  • path (str | Path)

  • run_id (str)

  • session_id (str)

  • attachment (TraceAttachment | None)

  • device_uuids (Mapping[int, str] | None)

  • detail (Literal['kernel', 'launch'])

  • device_uuids_by_pid (Mapping[int, Mapping[int, str]] | None)

Return type:

TraceCapture

stormlog.infer.trace_kineto.index_spans(trace)[source]

Sort each thread’s iteration spans and index them for launch lookups.

Parameters:

trace (KinetoTrace)

Return type:

None

Bases: object

An event’s iteration link, or the reason there is none.

Parameters:
iteration_ref: EntityRef | None
reason: str | None
launch: LaunchCall | None
class stormlog.infer.trace_kineto.KinetoTrace(base_ns, host, trace_id, rank, world_size, engine_version, cupti_version, device_names, source='kineto', gpu_events=<factory>, launches=<factory>, spans=<factory>, span_starts=<factory>, longest_span_us=<factory>, notes=<factory>, not_imported=<factory>)[source]

Bases: object

The parts of a profiler trace the importer uses (Kineto or Nsight).

Parameters:
  • base_ns (int)

  • host (str | None)

  • trace_id (str | None)

  • rank (int | None)

  • world_size (int | None)

  • engine_version (str | None)

  • cupti_version (str | None)

  • device_names (dict[int, str])

  • source (str)

  • gpu_events (list[GpuEvent])

  • launches (dict[tuple[int | None, int], LaunchCall])

  • spans (dict[tuple[int, int], list[IterationSpan]])

  • span_starts (dict[tuple[int, int], list[float]])

  • longest_span_us (dict[tuple[int, int], float])

  • notes (list[str])

  • not_imported (dict[str, int])

base_ns: int
host: str | None
trace_id: str | None
rank: int | None
world_size: int | None
engine_version: str | None
cupti_version: str | None
device_names: dict[int, str]
source: str = 'kineto'
gpu_events: list[GpuEvent]
launches: dict[tuple[int | None, int], LaunchCall]
spans: dict[tuple[int, int], list[IterationSpan]]
span_starts: dict[tuple[int, int], list[float]]
longest_span_us: dict[tuple[int, int], float]
notes: list[str]
not_imported: dict[str, int]
stormlog.infer.trace_kineto.import_kineto_trace(path, *, run_id, session_id, attachment=None, device_uuids=None, detail='launch')[source]

Return activity references, capabilities, and a summary for one trace.

Parameters:
  • path (str | Path)

  • run_id (str)

  • session_id (str)

  • attachment (TraceAttachment | None)

  • device_uuids (Mapping[int, str] | None)

  • detail (Literal['kernel', 'launch'])

Return type:

TraceCapture

Link a GPU event to the one iteration range around its launch call.

Parameters:
Return type:

GpuLink

stormlog.infer.trace_kineto.load_kineto_trace(path)[source]

Parse a plain or gzipped Kineto Chrome trace.

Parameters:

path (str | Path)

Return type:

KinetoTrace

stormlog.infer.trace_kineto.merge_intervals(intervals)[source]

Merge intervals into disjoint, sorted ones; touching intervals join.

Parameters:

intervals (Iterable[tuple[int, int]])

Return type:

list[tuple[int, int]]