stormlog.infer.trace_capture
Bounded vLLM profiler windows during stormlog infer profile.
vLLM’s torch profiler is configured when the server starts
(--profiler-config.profiler=torch and torch_profiler_dir). A client can
then start and stop it over HTTP (/start_profile and /stop_profile).
This module opens one window per profiled phase, closes it at the phase’s end,
at a time bound, or on cancellation, and finds the worker traces vLLM wrote.
A start the server may have received, whether it answered 2xx or not at all,
is always followed by one stop: the engine runs /start_profile before the
HTTP reply goes out, so a lost or failed reply can leave it profiling. A stop
that fails is not retried; the window’s record says the profiler may still be
running. Only a 401, 403, 404, 405 or 407, which come before vLLM’s handler
runs, is not stopped. vLLM 0.30.0 answers 200 to a second /start_profile
and to /stop_profile with nothing running, so a client cannot tell from
HTTP whether another profile was already active; do not run two profilers
against one server. The traces are imported after the run.
The profiler adds no synchronization per request; the server writes the trace
while handling /stop_profile, so that call can take tens of seconds.
Functions
|
The scheme and host of an endpoint URL, where vLLM serves its controls. |
|
|
Classes
|
What a profiler control call returned; it never raises. |
|
POST to vLLM's profiler routes with the profile's API key. |
|
|
|
What to capture, where vLLM writes it, and the bounds on the capture. |
|
One profiler window, recorded as an |
|
Open, bound, and close profiler windows; import their traces afterwards. |
- class stormlog.infer.trace_capture.ControlResult(status, error=None)[source]
Bases:
objectWhat a profiler control call returned; it never raises.
- Parameters:
status (int | None)
error (str | None)
- status: int | None
- error: str | None = None
- property ok: bool
- class stormlog.infer.trace_capture.HttpProfilerControl(root, *, api_key, timeout)[source]
Bases:
objectPOST to vLLM’s profiler routes with the profile’s API key.
- Parameters:
root (str)
api_key (str | None)
timeout (float)
- class stormlog.infer.trace_capture.ProfilerControl(*args, **kwargs)[source]
Bases:
Protocol
- class stormlog.infer.trace_capture.TraceCaptureConfig(control_url, trace_dir=None, mode='vllm-torch', phase='measured', max_seconds=None, max_bytes=None, device_uuids=<factory>, detail='launch', control_timeout_seconds=600.0, flush_timeout_seconds=60.0, missing_grace_seconds=5.0, settle_seconds=2.0)[source]
Bases:
objectWhat to capture, where vLLM writes it, and the bounds on the capture.
- Parameters:
control_url (str)
trace_dir (Path | None)
mode (str)
phase (str)
max_seconds (float | None)
max_bytes (int | None)
device_uuids (DeviceUuids)
detail (Literal['kernel', 'launch'])
control_timeout_seconds (float)
flush_timeout_seconds (float)
missing_grace_seconds (float)
settle_seconds (float)
- control_url: str
- trace_dir: Path | None = None
- mode: str = 'vllm-torch'
- phase: str = 'measured'
- max_seconds: float | None = None
- max_bytes: int | None = None
- device_uuids: DeviceUuids
- detail: Literal['kernel', 'launch'] = 'launch'
- control_timeout_seconds: float = 600.0
- flush_timeout_seconds: float = 60.0
- missing_grace_seconds: float = 5.0
- settle_seconds: float = 2.0
- class stormlog.infer.trace_capture.TraceWindow(case_id, phase, control_url, requested_at_ns, started=False, started_at_ns=None, start=None, start_outcome=None, start_requested_at_ns=None, start_returned_at_ns=None, stop=None, stop_reason=None, stopped_at_ns=None, files=<factory>, note=None, before=None, stopping=None)[source]
Bases:
objectOne profiler window, recorded as an
infer.trace_windowevent.- Parameters:
case_id (str)
phase (str)
control_url (str)
requested_at_ns (int)
started (bool)
started_at_ns (int | None)
start (ControlResult | None)
start_outcome (str | None)
start_requested_at_ns (int | None)
start_returned_at_ns (int | None)
stop (ControlResult | None)
stop_reason (str | None)
stopped_at_ns (int | None)
files (list[Path])
note (str | None)
before (dict[str, int] | None)
stopping (Future[None] | None)
- case_id: str
- phase: str
- control_url: str
- requested_at_ns: int
- started: bool = False
- started_at_ns: int | None = None
- start: ControlResult | None = None
- start_outcome: str | None = None
- start_requested_at_ns: int | None = None
- start_returned_at_ns: int | None = None
- stop: ControlResult | None = None
- stop_reason: str | None = None
- stopped_at_ns: int | None = None
- files: list[Path]
- note: str | None = None
- before: dict[str, int] | None = None
- stopping: Future[None] | None = None
- class stormlog.infer.trace_capture.TraceWindows(config, *, api_key=None, control=None, on_warning=None)[source]
Bases:
objectOpen, bound, and close profiler windows; import their traces afterwards.
- Parameters:
config (TraceCaptureConfig)
api_key (str | None)
control (ProfilerControl | None)
on_warning (Callable[[str], None] | None)
- window(case_id, phase)[source]
Profile the body if
phaseis the configured one.- Parameters:
case_id (str)
phase (str)
- Return type:
AsyncIterator[TraceWindow | None]
- take_records(*, session_id)[source]
infer.trace_windowrecords for windows closed since the last call.- Parameters:
session_id (str)
- Return type:
list[dict[str, Any]]
- import_into(artifact, *, run_id, session, worker_index=None)[source]
Append every window’s traces to the artifact; warn instead of failing.
With no trace to import, the trace collector is still recorded, with nothing collected and each window’s reason in its summary. The
worker_indexis the vLLM execution hook’s, which names each traced process’s GPU where--trace-device-uuiddoes not.- Parameters:
artifact (Path)
run_id (str)
session (SessionSummary)
worker_index (WorkerIndex | None)
- Return type:
None
- stormlog.infer.trace_capture.server_root(endpoint)[source]
The scheme and host of an endpoint URL, where vLLM serves its controls.
- Parameters:
endpoint (str)
- Return type:
str
- stormlog.infer.trace_capture.start_outcome(result)[source]
acknowledged(2xx),rejected(REJECTING_STATUSES), elseunknown.Any other answer can follow a start that reached the engine. vLLM 0.30.0’s handler awaits the engine’s start, and its own profiler’s, before it replies, and maps an exception raised meanwhile to 400 (ValueError, TypeError, OverflowError), 422, 501 or 500 (
vllm/entrypoints/serve/profile/api_router.pyandserve/exception_handling/error_response.py). An intermediary can answer 408 or 499 after forwarding the call, and a timeout, a reset or a malformed reply says nothing either.- Parameters:
result (ControlResult)
- Return type:
str