stormlog.infer.vllm_hook.worker
Worker side: wrap each serving step’s launches in its iteration range.
The engine core puts (producer, iteration) on the scheduler output, which
reaches the worker as the same object (uniproc) or pickled (multiproc). The
model runner’s execute_model is wrapped in that iteration’s range; when it
returns None, sampling for the same step comes in a later sample_tokens
call, which a FIFO pairs with it. Internal dummy runs are never ranged or
paired: the V2 runner’s go through execute_model(dummy_run=True), and the V1
runner’s _dummy_run never calls execute_model. Warm-up and CUDA-graph
capture steps before the first serving step are counted as start-up runs.
Functions
|
Ranks, process and GPU of a worker, each field best effort. |
Classes
|
Ranges and counters for one model-runner instance. |
- class stormlog.infer.vllm_hook.worker.RunnerRecorder(writer, nvtx=False, range_factory=None, pending=<factory>, range_misses=0, startup_unranged=0, serving=False)[source]
Bases:
objectRanges and counters for one model-runner instance.
- Parameters:
writer (EpochWriter)
nvtx (bool)
range_factory (Callable[[str, str, bool], Any] | None)
pending (deque[tuple[str, str]])
range_misses (int)
startup_unranged (int)
serving (bool)
- writer: EpochWriter
- nvtx: bool = False
- range_factory: Callable[[str, str, bool], Any] | None = None
- pending: deque[tuple[str, str]]
- range_misses: int = 0
- startup_unranged: int = 0
- serving: bool = False