stormlog.infer.trace_kineto
Import PyTorch/Kineto profiler traces as inference activity references.
A Kineto trace (PyTorch profiler, or vLLM’s rank*.pt.trace.json.gz) lists
GPU kernels, copies, and memsets with device timestamps, and the CPU runtime or
driver calls that launched them. Each GPU event carries the CUDA correlation ID
of its launch call. A CUDA graph launch is one call and many GPU events.
This importer emits one infer.activity_ref per launch, device, and activity
kind (detail="launch", the default) or per GPU event (detail="kernel"). A launch record spans its
first event’s start to its last event’s end. When it holds several events, its
metadata.intervals lists the exact busy intervals inside that span, so
GPU time accounting does not count the idle gaps inside a CUDA graph replay as
busy. On vLLM 0.30.0 traces a launch record set is about 6x smaller than a
per-event one and gives the same device busy time. A GPU
activity is linked to an iteration only when its launch call sits inside exactly
one stormlog.iteration/... range on the launching thread. Otherwise it stays
unresolved and says why. Launch calls are CPU work and are never emitted as GPU
activity.
Functions
|
Build the capture for a loaded trace of any supported format. |
|
Return activity references, capabilities, and a summary for one trace. |
|
Sort each thread's iteration spans and index them for launch lookups. |
|
Link a GPU event to the one iteration range around its launch call. |
|
Parse a plain or gzipped Kineto Chrome trace. |
|
Merge intervals into disjoint, sorted ones; touching intervals join. |
Classes
|
One kernel, copy, or memset as the trace recorded it. |
|
An event's iteration link, or the reason there is none. |
|
|
|
The parts of a profiler trace the importer uses (Kineto or Nsight). |
|
The CPU runtime or driver call that launched GPU work. |
- class stormlog.infer.trace_kineto.GpuEvent(start_ns, end_ns, kind, name, device, stream, correlation, graph_id, pid=None, device_uuid=None, device_name=None, graph_node_id=None)[source]
Bases:
objectOne kernel, copy, or memset as the trace recorded it.
- Parameters:
start_ns (int)
end_ns (int)
kind (str)
name (str)
device (int | None)
stream (int | None)
correlation (int | None)
graph_id (int | None)
pid (int | None)
device_uuid (str | None)
device_name (str | None)
graph_node_id (int | None)
- start_ns: int
- end_ns: int
- kind: str
- name: str
- device: int | None
- stream: int | None
- correlation: int | None
- graph_id: int | None
- pid: int | None = None
- device_uuid: str | None = None
- device_name: str | None = None
- graph_node_id: int | None = None
- class stormlog.infer.trace_kineto.IterationSpan(start_us: 'float', end_us: 'float', iteration_ref: 'EntityRef')[source]
Bases:
object- Parameters:
start_us (float)
end_us (float)
iteration_ref (EntityRef)
- start_us: float
- end_us: float
- class stormlog.infer.trace_kineto.LaunchCall(pid, tid, ts_us, name)[source]
Bases:
objectThe CPU runtime or driver call that launched GPU work.
- Parameters:
pid (int)
tid (int)
ts_us (float)
name (str)
- pid: int
- tid: int
- ts_us: float
- name: str
- stormlog.infer.trace_kineto.capture_trace(trace, path, *, run_id, session_id, attachment=None, device_uuids=None, detail='launch', device_uuids_by_pid=None)[source]
Build the capture for a loaded trace of any supported format.
device_uuids(--device-uuid) names a CUDA ordinal’s GPU for every process in the trace and is checked against UUIDs the trace names itself.device_uuids_by_pidis what the vLLM execution log’s worker hellos bound: a process’s own ordinals only, so a process no hello matched stays unmeasured. The option wins where both name an ordinal.- Parameters:
trace (KinetoTrace)
path (str | Path)
run_id (str)
session_id (str)
attachment (TraceAttachment | None)
device_uuids (Mapping[int, str] | None)
detail (Literal['kernel', 'launch'])
device_uuids_by_pid (Mapping[int, Mapping[int, str]] | None)
- Return type:
- stormlog.infer.trace_kineto.index_spans(trace)[source]
Sort each thread’s iteration spans and index them for launch lookups.
- Parameters:
trace (KinetoTrace)
- Return type:
None
- class stormlog.infer.trace_kineto.GpuLink(iteration_ref, reason, launch)[source]
Bases:
objectAn event’s iteration link, or the reason there is none.
- Parameters:
iteration_ref (EntityRef | None)
reason (str | None)
launch (LaunchCall | None)
- reason: str | None
- launch: LaunchCall | None
- class stormlog.infer.trace_kineto.KinetoTrace(base_ns, host, trace_id, rank, world_size, engine_version, cupti_version, device_names, source='kineto', gpu_events=<factory>, launches=<factory>, spans=<factory>, span_starts=<factory>, longest_span_us=<factory>, notes=<factory>, not_imported=<factory>)[source]
Bases:
objectThe parts of a profiler trace the importer uses (Kineto or Nsight).
- Parameters:
base_ns (int)
host (str | None)
trace_id (str | None)
rank (int | None)
world_size (int | None)
engine_version (str | None)
cupti_version (str | None)
device_names (dict[int, str])
source (str)
gpu_events (list[GpuEvent])
launches (dict[tuple[int | None, int], LaunchCall])
spans (dict[tuple[int, int], list[IterationSpan]])
span_starts (dict[tuple[int, int], list[float]])
longest_span_us (dict[tuple[int, int], float])
notes (list[str])
not_imported (dict[str, int])
- base_ns: int
- host: str | None
- trace_id: str | None
- rank: int | None
- world_size: int | None
- engine_version: str | None
- cupti_version: str | None
- device_names: dict[int, str]
- source: str = 'kineto'
- launches: dict[tuple[int | None, int], LaunchCall]
- spans: dict[tuple[int, int], list[IterationSpan]]
- span_starts: dict[tuple[int, int], list[float]]
- longest_span_us: dict[tuple[int, int], float]
- notes: list[str]
- not_imported: dict[str, int]
- stormlog.infer.trace_kineto.import_kineto_trace(path, *, run_id, session_id, attachment=None, device_uuids=None, detail='launch')[source]
Return activity references, capabilities, and a summary for one trace.
- Parameters:
path (str | Path)
run_id (str)
session_id (str)
attachment (TraceAttachment | None)
device_uuids (Mapping[int, str] | None)
detail (Literal['kernel', 'launch'])
- Return type:
- stormlog.infer.trace_kineto.link_gpu_event(trace, event)[source]
Link a GPU event to the one iteration range around its launch call.
- Parameters:
trace (KinetoTrace)
event (GpuEvent)
- Return type: