stormlog.infer.vllm_hook.worker

Worker side: wrap each serving step’s launches in its iteration range.

The engine core puts (producer, iteration) on the scheduler output, which reaches the worker as the same object (uniproc) or pickled (multiproc). The model runner’s execute_model is wrapped in that iteration’s range; when it returns None, sampling for the same step comes in a later sample_tokens call, which a FIFO pairs with it. Internal dummy runs are never ranged or paired: the V2 runner’s go through execute_model(dummy_run=True), and the V1 runner’s _dummy_run never calls execute_model. Warm-up and CUDA-graph capture steps before the first serving step are counted as start-up runs.

Functions

worker_identity(worker)

Ranks, process and GPU of a worker, each field best effort.

Classes

RunnerRecorder(writer[, nvtx, ...])

Ranges and counters for one model-runner instance.

class stormlog.infer.vllm_hook.worker.RunnerRecorder(writer, nvtx=False, range_factory=None, pending=<factory>, range_misses=0, startup_unranged=0, serving=False)[source]

Bases: object

Ranges and counters for one model-runner instance.

Parameters:
  • writer (EpochWriter)

  • nvtx (bool)

  • range_factory (Callable[[str, str, bool], Any] | None)

  • pending (deque[tuple[str, str]])

  • range_misses (int)

  • startup_unranged (int)

  • serving (bool)

writer: EpochWriter
nvtx: bool = False
range_factory: Callable[[str, str, bool], Any] | None = None
pending: deque[tuple[str, str]]
range_misses: int = 0
startup_unranged: int = 0
serving: bool = False
status_fields()[source]
Return type:

dict[str, Any]

wrap(runner)[source]
Parameters:

runner (Any)

Return type:

None

stormlog.infer.vllm_hook.worker.worker_identity(worker)[source]

Ranks, process and GPU of a worker, each field best effort.

Parameters:

worker (Any)

Return type:

dict[str, Any]