stormlog.infer.vllm_telemetry

Records for vLLM native telemetry inside an inference artifact.

A scrape record holds one /metrics response in the compact form from stormlog.infer.vllm_metrics, stamped on the client’s clock. A span record holds one OpenTelemetry span exactly as vLLM exported it. Both are appended to the client artifact next to the request events, and both are aggregate or engine-side evidence: nothing in them says which request used which GPU time.

Functions

load_vllm_records(records)

Parse the vLLM records of an artifact; an invalid one is an error.

request_id_from_span_id(native_id)

Recover the X-Request-Id vLLM embedded in gen_ai.request.id.

Classes

VllmScrapeRecord(session_id, run_id, ...[, ...])

One /metrics response, kept whole, stamped on the client clock.

VllmSpanRecord(session_id, run_id, source, ...)

One span as exported, with its attributes under their native names.

class stormlog.infer.vllm_telemetry.VllmScrapeRecord(session_id, run_id, observed_at_ns, source_url, marker, interval_ms, status, clock_domain, case_id=None, phase=None, duration_ms=None, http_status=None, error=None, content_digest=None, content_bytes=None, scrape=None, discovery=None)[source]

Bases: object

One /metrics response, kept whole, stamped on the client clock.

Parameters:
  • session_id (str)

  • run_id (str)

  • observed_at_ns (int)

  • source_url (str)

  • marker (str)

  • interval_ms (int)

  • status (str)

  • clock_domain (str)

  • case_id (str | None)

  • phase (str | None)

  • duration_ms (float | None)

  • http_status (int | None)

  • error (str | None)

  • content_digest (str | None)

  • content_bytes (int | None)

  • scrape (CompactScrape | None)

  • discovery (Discovery | None)

session_id: str
run_id: str
observed_at_ns: int
source_url: str
marker: str
interval_ms: int
status: str
clock_domain: str
case_id: str | None = None
phase: str | None = None
duration_ms: float | None = None
http_status: int | None = None
error: str | None = None
content_digest: str | None = None
content_bytes: int | None = None
scrape: CompactScrape | None = None
discovery: Discovery | None = None
to_record()[source]
Return type:

dict[str, Any]

classmethod from_record(record)[source]
Parameters:

record (dict[str, Any])

Return type:

VllmScrapeRecord

class stormlog.infer.vllm_telemetry.VllmSpanRecord(session_id, run_id, source, name, clock_domain, received_at_ns=None, trace_id=None, span_id=None, parent_span_id=None, kind=None, start_unix_ns=None, end_unix_ns=None, attributes=<factory>, resource=<factory>, scope=<factory>, status=None, dropped=<factory>, request_id=None)[source]

Bases: object

One span as exported, with its attributes under their native names.

The timestamps are the exporter’s wall clock, named by clock_domain. request_id is the client-chosen X-Request-Id recovered from gen_ai.request.id when the span carries one; the join to a Stormlog request uses it, never a rebuilt string.

Parameters:
  • session_id (str)

  • run_id (str)

  • source (str)

  • name (str)

  • clock_domain (str)

  • received_at_ns (int | None)

  • trace_id (str | None)

  • span_id (str | None)

  • parent_span_id (str | None)

  • kind (str | None)

  • start_unix_ns (int | None)

  • end_unix_ns (int | None)

  • attributes (dict[str, Any])

  • resource (dict[str, Any])

  • scope (dict[str, Any])

  • status (dict[str, Any] | None)

  • dropped (dict[str, int])

  • request_id (str | None)

session_id: str
run_id: str
source: str
name: str
clock_domain: str
received_at_ns: int | None = None
trace_id: str | None = None
span_id: str | None = None
parent_span_id: str | None = None
kind: str | None = None
start_unix_ns: int | None = None
end_unix_ns: int | None = None
attributes: dict[str, Any]
resource: dict[str, Any]
scope: dict[str, Any]
status: dict[str, Any] | None = None
dropped: dict[str, int]
request_id: str | None = None
property duration_ns: int | None
to_record()[source]
Return type:

dict[str, Any]

classmethod from_record(record)[source]
Parameters:

record (dict[str, Any])

Return type:

VllmSpanRecord

stormlog.infer.vllm_telemetry.request_id_from_span_id(native_id)[source]

Recover the X-Request-Id vLLM embedded in gen_ai.request.id.

vLLM names a chat completion chatcmpl-<X-Request-Id> and a text completion cmpl-<X-Request-Id>-<index>, with the per-prompt index only on the completions path. A trailing -<digits> is stripped only when it is one; the request ids Stormlog sends end in _<n>, never -<n>. Without the client’s header vLLM uses a random id, which no Stormlog request owns.

Parameters:

native_id (object)

Return type:

str | None

stormlog.infer.vllm_telemetry.load_vllm_records(records)[source]

Parse the vLLM records of an artifact; an invalid one is an error.

Parameters:

records (Iterable[dict[str, Any]])

Return type:

tuple[list[VllmScrapeRecord], list[VllmSpanRecord]]