stormlog.infer.trace_capture

Bounded vLLM profiler windows during stormlog infer profile.

vLLM’s torch profiler is configured when the server starts (--profiler-config.profiler=torch and torch_profiler_dir). A client can then start and stop it over HTTP (/start_profile and /stop_profile). This module opens one window per profiled phase, closes it at the phase’s end, at a time bound, or on cancellation, and finds the worker traces vLLM wrote. A start the server may have received, whether it answered 2xx or not at all, is always followed by one stop: the engine runs /start_profile before the HTTP reply goes out, so a lost or failed reply can leave it profiling. A stop that fails is not retried; the window’s record says the profiler may still be running. Only a 401, 403, 404, 405 or 407, which come before vLLM’s handler runs, is not stopped. vLLM 0.30.0 answers 200 to a second /start_profile and to /stop_profile with nothing running, so a client cannot tell from HTTP whether another profile was already active; do not run two profilers against one server. The traces are imported after the run.

The profiler adds no synchronization per request; the server writes the trace while handling /stop_profile, so that call can take tens of seconds.

Functions

server_root(endpoint)

The scheme and host of an endpoint URL, where vLLM serves its controls.

start_outcome(result)

acknowledged (2xx), rejected (REJECTING_STATUSES), else unknown.

Classes

ControlResult(status[, error])

What a profiler control call returned; it never raises.

HttpProfilerControl(root, *, api_key, timeout)

POST to vLLM's profiler routes with the profile's API key.

ProfilerControl(*args, **kwargs)

TraceCaptureConfig(control_url[, trace_dir, ...])

What to capture, where vLLM writes it, and the bounds on the capture.

TraceWindow(case_id, phase, control_url, ...)

One profiler window, recorded as an infer.trace_window event.

TraceWindows(config, *[, api_key, control, ...])

Open, bound, and close profiler windows; import their traces afterwards.

class stormlog.infer.trace_capture.ControlResult(status, error=None)[source]

Bases: object

What a profiler control call returned; it never raises.

Parameters:
  • status (int | None)

  • error (str | None)

status: int | None
error: str | None = None
property ok: bool
class stormlog.infer.trace_capture.HttpProfilerControl(root, *, api_key, timeout)[source]

Bases: object

POST to vLLM’s profiler routes with the profile’s API key.

Parameters:
  • root (str)

  • api_key (str | None)

  • timeout (float)

post(route)[source]
Parameters:

route (str)

Return type:

ControlResult

class stormlog.infer.trace_capture.ProfilerControl(*args, **kwargs)[source]

Bases: Protocol

post(route)[source]
Parameters:

route (str)

Return type:

ControlResult

class stormlog.infer.trace_capture.TraceCaptureConfig(control_url, trace_dir=None, mode='vllm-torch', phase='measured', max_seconds=None, max_bytes=None, device_uuids=<factory>, detail='launch', control_timeout_seconds=600.0, flush_timeout_seconds=60.0, missing_grace_seconds=5.0, settle_seconds=2.0)[source]

Bases: object

What to capture, where vLLM writes it, and the bounds on the capture.

Parameters:
  • control_url (str)

  • trace_dir (Path | None)

  • mode (str)

  • phase (str)

  • max_seconds (float | None)

  • max_bytes (int | None)

  • device_uuids (DeviceUuids)

  • detail (Literal['kernel', 'launch'])

  • control_timeout_seconds (float)

  • flush_timeout_seconds (float)

  • missing_grace_seconds (float)

  • settle_seconds (float)

control_url: str
trace_dir: Path | None = None
mode: str = 'vllm-torch'
phase: str = 'measured'
max_seconds: float | None = None
max_bytes: int | None = None
device_uuids: DeviceUuids
detail: Literal['kernel', 'launch'] = 'launch'
control_timeout_seconds: float = 600.0
flush_timeout_seconds: float = 60.0
missing_grace_seconds: float = 5.0
settle_seconds: float = 2.0
class stormlog.infer.trace_capture.TraceWindow(case_id, phase, control_url, requested_at_ns, started=False, started_at_ns=None, start=None, start_outcome=None, start_requested_at_ns=None, start_returned_at_ns=None, stop=None, stop_reason=None, stopped_at_ns=None, files=<factory>, note=None, before=None, stopping=None)[source]

Bases: object

One profiler window, recorded as an infer.trace_window event.

Parameters:
  • case_id (str)

  • phase (str)

  • control_url (str)

  • requested_at_ns (int)

  • started (bool)

  • started_at_ns (int | None)

  • start (ControlResult | None)

  • start_outcome (str | None)

  • start_requested_at_ns (int | None)

  • start_returned_at_ns (int | None)

  • stop (ControlResult | None)

  • stop_reason (str | None)

  • stopped_at_ns (int | None)

  • files (list[Path])

  • note (str | None)

  • before (dict[str, int] | None)

  • stopping (Future[None] | None)

case_id: str
phase: str
control_url: str
requested_at_ns: int
started: bool = False
started_at_ns: int | None = None
start: ControlResult | None = None
start_outcome: str | None = None
start_requested_at_ns: int | None = None
start_returned_at_ns: int | None = None
stop: ControlResult | None = None
stop_reason: str | None = None
stopped_at_ns: int | None = None
files: list[Path]
note: str | None = None
before: dict[str, int] | None = None
stopping: Future[None] | None = None
to_record(*, session_id)[source]
Parameters:

session_id (str)

Return type:

dict[str, Any]

class stormlog.infer.trace_capture.TraceWindows(config, *, api_key=None, control=None, on_warning=None)[source]

Bases: object

Open, bound, and close profiler windows; import their traces afterwards.

Parameters:
window(case_id, phase)[source]

Profile the body if phase is the configured one.

Parameters:
  • case_id (str)

  • phase (str)

Return type:

AsyncIterator[TraceWindow | None]

take_records(*, session_id)[source]

infer.trace_window records for windows closed since the last call.

Parameters:

session_id (str)

Return type:

list[dict[str, Any]]

import_into(artifact, *, run_id, session, worker_index=None)[source]

Append every window’s traces to the artifact; warn instead of failing.

With no trace to import, the trace collector is still recorded, with nothing collected and each window’s reason in its summary. The worker_index is the vLLM execution hook’s, which names each traced process’s GPU where --trace-device-uuid does not.

Parameters:
Return type:

None

stormlog.infer.trace_capture.server_root(endpoint)[source]

The scheme and host of an endpoint URL, where vLLM serves its controls.

Parameters:

endpoint (str)

Return type:

str

stormlog.infer.trace_capture.start_outcome(result)[source]

acknowledged (2xx), rejected (REJECTING_STATUSES), else unknown.

Any other answer can follow a start that reached the engine. vLLM 0.30.0’s handler awaits the engine’s start, and its own profiler’s, before it replies, and maps an exception raised meanwhile to 400 (ValueError, TypeError, OverflowError), 422, 501 or 500 (vllm/entrypoints/serve/profile/api_router.py and serve/exception_handling/error_response.py). An intermediary can answer 408 or 499 after forwarding the call, and a timeout, a reset or a malformed reply says nothing either.

Parameters:

result (ControlResult)

Return type:

str