stormlog.infer.vllm_hook.engine
Engine-core snapshots: admissions, scheduled steps, their outputs, and exits.
Every value is a scalar copied at the moment it is observed; nothing keeps a
live vLLM object. Per-step context comes from the scheduler output’s own fields
because Scheduler.schedule has already advanced each request’s
num_computed_tokens by the time it returns, and under async scheduling the
next step has been scheduled before this one’s output is processed.
Classes
|
Records for one scheduler instance. |
- class stormlog.infer.vllm_hook.engine.EngineRecorder(writer, producer, next_iteration=0, prompt_tokens=<factory>, committed=<factory>, pending=<factory>)[source]
Bases:
objectRecords for one scheduler instance.
- Parameters:
writer (EpochWriter)
producer (str)
next_iteration (int)
prompt_tokens (dict[str, int])
committed (dict[str, int])
pending (dict[str, _Pending])
- writer: EpochWriter
- producer: str
- next_iteration: int = 0
- prompt_tokens: dict[str, int]
- committed: dict[str, int]
- pending: dict[str, _Pending]
- on_schedule(scheduler, output, start)[source]
- Parameters:
scheduler (Any)
output (Any)
start (tuple[int, int])
- Return type:
None
- before_update(scheduler, output, model_output)[source]
Per member, what vLLM’s own output loop is about to read.
- Parameters:
scheduler (Any)
output (Any)
model_output (Any)
- Return type:
dict[str, dict[str, Any]]