stormlog.infer.vllm_execution

Reduce the vLLM execution hook’s raw log into canonical correlation records.

The hook (docs/vllm_execution.md) writes immutable snapshots: an alias when a request is admitted, a scheduled snapshot when a step is scheduled, a completed one when its output has been processed, and a terminal when the request is freed. This module turns them into infer.iteration, infer.membership, infer.request and infer.clock_alignment records, each written exactly once:

  • an iteration is reduced only once it is final: its completed snapshot arrived, or it never will because the epoch ended (a goodbye, or silence) or a later iteration completed, which makes it incomplete; a pending iteration waits for a later import, with no late corrections;

  • a request record holds admission facts only, which never change; how the request finished goes on the membership of the iteration it was freed in;

  • the per-epoch high-water mark advances only past records that were fully reduced, and entities the artifact already holds are never emitted again.

Requests are bound to the run exactly, by the alias’s external ID against the X-Request-Id values the client recorded; the internal ID’s shape only proposes a candidate for that same match. Other clients’ requests are kept as members under keyed pseudonyms, never by their IDs.

Functions

foreign_scheme(epoch, raw_foreign_ids)

How an import with these options writes an epoch's other clients: raw IDs, keyed pseudonyms, or nothing (withheld without the key).

reduce_execution_log(read, facts[, options])

Reduce every engine epoch of read for the run in facts.

Classes

Binding(ownership[, request, child_index, ...])

How an internal ID relates to the run.

Execution(key, internal, binding[, alias, ...])

One backend execution of an internal ID; a reused ID has several.

Iteration(iteration, scheduled, completed, ...)

Member(iteration, execution, data, outcome)

One execution's part in one kept iteration.

ReduceOptions([raw_foreign_ids])

ReduceResult(events, summary, high_water)

RunFacts(run_id, session_id, ...[, windows, ...])

What the artifact tells the reducer about the run it binds to.

RunRequest(request_id, x_request_id, ...)

One request the client sent, as the artifact recorded it.

Window(kind, start_ns, end_ns[, case_id, phase])

A client-clock window: a run phase or a trace window.

class stormlog.infer.vllm_execution.Binding(ownership, request=None, child_index=None, completion_index=None, external=None, via='none')[source]

Bases: object

How an internal ID relates to the run.

Parameters:
  • ownership (str)

  • request (RunRequest | None)

  • child_index (int | None)

  • completion_index (int | None)

  • external (str | None)

  • via (str)

ownership: str
request: RunRequest | None = None
child_index: int | None = None
completion_index: int | None = None
external: str | None = None
via: str = 'none'
class stormlog.infer.vllm_execution.ReduceOptions(raw_foreign_ids: 'bool' = False)[source]

Bases: object

Parameters:

raw_foreign_ids (bool)

raw_foreign_ids: bool = False
class stormlog.infer.vllm_execution.ReduceResult(events: 'list[CorrelationEvent]', summary: 'dict[str, Any]', high_water: 'dict[str, int]')[source]

Bases: object

Parameters:
  • events (list[CorrelationEvent])

  • summary (dict[str, Any])

  • high_water (dict[str, int])

events: list[CorrelationEvent]
summary: dict[str, Any]
high_water: dict[str, int]
class stormlog.infer.vllm_execution.RunFacts(run_id, session_id, client_clock_domain, requests, windows=(), referenced_iterations=frozenset({}), existing_iterations=frozenset({}), existing_attempts=frozenset({}), existing_alignments=frozenset({}), existing_admissions=<factory>)[source]

Bases: object

What the artifact tells the reducer about the run it binds to.

Parameters:
  • run_id (str)

  • session_id (str)

  • client_clock_domain (str | None)

  • requests (dict[str, RunRequest])

  • windows (tuple[Window, ...])

  • referenced_iterations (frozenset[EntityRef])

  • existing_iterations (frozenset[EntityRef])

  • existing_attempts (frozenset[EntityRef])

  • existing_alignments (frozenset[str])

  • existing_admissions (dict[tuple[str, int], EntityRef])

run_id: str
session_id: str
client_clock_domain: str | None
requests: dict[str, RunRequest]
windows: tuple[Window, ...] = ()
referenced_iterations: frozenset[EntityRef] = frozenset({})
existing_iterations: frozenset[EntityRef] = frozenset({})
existing_attempts: frozenset[EntityRef] = frozenset({})
existing_alignments: frozenset[str] = frozenset({})
existing_admissions: dict[tuple[str, int], EntityRef]
class stormlog.infer.vllm_execution.RunRequest(request_id, x_request_id, case_id, phase)[source]

Bases: object

One request the client sent, as the artifact recorded it.

Parameters:
  • request_id (str)

  • x_request_id (str)

  • case_id (str | None)

  • phase (str | None)

request_id: str
x_request_id: str
case_id: str | None
phase: str | None
class stormlog.infer.vllm_execution.Window(kind, start_ns, end_ns, case_id=None, phase=None)[source]

Bases: object

A client-clock window: a run phase or a trace window.

Parameters:
  • kind (str)

  • start_ns (int)

  • end_ns (int)

  • case_id (str | None)

  • phase (str | None)

kind: str
start_ns: int
end_ns: int
case_id: str | None = None
phase: str | None = None
stormlog.infer.vllm_execution.foreign_scheme(epoch, raw_foreign_ids)[source]

How an import with these options writes an epoch’s other clients: raw IDs, keyed pseudonyms, or nothing (withheld without the key).

Parameters:
Return type:

str

stormlog.infer.vllm_execution.reduce_execution_log(read, facts, options=None)[source]

Reduce every engine epoch of read for the run in facts.

Parameters:
Return type:

ReduceResult