stormlog.infer.diagnosis_loop

Stalls of the vLLM engine loop, from the execution hook’s raw records.

engine_loop_gap reads one engine epoch’s raw hook records (stormlog.vllm_hook/1; see docs/vllm_execution.md) and reports the stretch in which the engine made no progress while it had work it could run that is furthest over its own limit. The same rules serve an online trigger, which tails the raw log, and the offline diagnoser, which feeds imported steps back through the same adapter, so the two cannot disagree about what a stall is.

Where a stall sits decides what it can be blamed on:

  • between_steps: from a step’s completion to the next schedule() entry. Nothing ran on the host’s behalf, so it is the host’s.

  • in_schedule: inside schedule(), which is host work.

  • within_step: a step that took far longer than the steps before it. The host or the GPU may have been slow; without a GPU trace it is host_or_gpu.

Work is ready when a request ran in the step before the stretch and in the step after it: it was running, not waiting for capacity or for streaming input. A stretch covered by a scheduler pause (when the hook records pauses) or by an interval the caller excludes, such as its own profiler stop, has no ready work. A stall still going on at the evaluation time counts from the last completion.

Functions

cadence_baseline(table, before_mono_ns, ...)

The median completion cadence of the busy steps that completed in the window before a stall: those of the stall's own work bucket when there are enough, else those of its bucket or larger when there are enough, else none (the floor applies).

cadence_table(steps)

Each busy step's completion cadence, as (completion, work bucket, cadence), in completion order: the time since the previous completion, for steps whose work continued from that step.

continuing(before, after)

The requests of before that after runs on: not finished in before, and, for a streaming-input request, not starting a new turn in after.

engine_loop_gap(records[, config])

The stall with ready work furthest over its limit in one epoch's records, in seq order: the longest when none is over.

find_stalls(steps, blocked, now_wall_ns[, ...])

Every candidate stall with ready work, in each locus.

observes_pauses(records)

Whether the hook that wrote these records records scheduler pauses.

pause_intervals(records[, state])

Wall-clock spans the scheduler spent in state, from the hook's pause transitions; a pause not yet ended runs to the end of time.

ready(step)

The members still ready to run after step, with nothing scheduled since: not finished in it, and not a streaming-input request, which may be waiting for its client's next input.

record_reasons(records)

Why the records cannot be trusted to show every step: another epoch, a sequence gap, records the writer dropped, or a capped writer.

steps_from_raw(records)

Join each scheduled record to its completed by iteration.

work_bucket(total_tokens)

Steps that schedule within a factor of two of each other's tokens share a bucket: a step running a long prefill is compared with steps of its own size, never with decode-only steps.

Classes

Coverage(spans, last_beat_wall_ns)

Where the records are known to be whole, on the wall clock: spans between two heartbeats (the hello counting as one with nothing lost) whose drop counts and errors did not change, and when the writer was last heard from.

LoopGapConfig([now_wall_ns, exclude_wall, ...])

How to evaluate one window of hook records.

Stall(locus, attribution, duration_ns, ...)

Step(iteration, start_wall_ns, ...[, ...])

One scheduler step, as the raw log or an import describes it.

class stormlog.infer.diagnosis_loop.Coverage(spans, last_beat_wall_ns)[source]

Bases: object

Where the records are known to be whole, on the wall clock: spans between two heartbeats (the hello counting as one with nothing lost) whose drop counts and errors did not change, and when the writer was last heard from. Only inside a span can no record be missing.

Parameters:
  • spans (tuple[tuple[int, int], ...])

  • last_beat_wall_ns (int | None)

spans: tuple[tuple[int, int], ...]
last_beat_wall_ns: int | None
classmethod of(records)[source]
Parameters:

records (Sequence[Mapping[str, Any]])

Return type:

Coverage

covers(stall, now_wall_ns, grace_ns)[source]

A stall lies inside one span. One still going on needs the writer heard from since it began, and within the grace before now.

Parameters:
  • stall (Stall)

  • now_wall_ns (int | None)

  • grace_ns (float)

Return type:

bool

class stormlog.infer.diagnosis_loop.LoopGapConfig(now_wall_ns=None, exclude_wall=(), thresholds=<factory>, status=None, hello=None)[source]

Bases: object

How to evaluate one window of hook records.

now_wall_ns is the evaluation time, for a stall still going on. exclude_wall lists wall-clock intervals with no ready work by the caller’s knowledge, such as its own profiler stop. thresholds overrides entries of the shared table by key. status is the epoch’s status.json, when the caller has it: a capped writer stops writing records and heartbeats alike, and only the status says so. hello is the epoch’s hello for a window that starts later, as a tail does: it says whether the hook records pauses.

Parameters:
  • now_wall_ns (int | None)

  • exclude_wall (Sequence[tuple[int, int]])

  • thresholds (Mapping[str, float])

  • status (Mapping[str, Any] | None)

  • hello (Mapping[str, Any] | None)

now_wall_ns: int | None = None
exclude_wall: Sequence[tuple[int, int]] = ()
thresholds: Mapping[str, float]
status: Mapping[str, Any] | None = None
hello: Mapping[str, Any] | None = None
class stormlog.infer.diagnosis_loop.Stall(locus: 'str', attribution: 'str', duration_ns: 'int', start_mono_ns: 'int', start_wall_ns: 'int', end_wall_ns: 'int', bucket: 'int' = 0, ongoing: 'bool' = False, overlapped: 'bool' = False)[source]

Bases: object

Parameters:
  • locus (str)

  • attribution (str)

  • duration_ns (int)

  • start_mono_ns (int)

  • start_wall_ns (int)

  • end_wall_ns (int)

  • bucket (int)

  • ongoing (bool)

  • overlapped (bool)

locus: str
attribution: str
duration_ns: int
start_mono_ns: int
start_wall_ns: int
end_wall_ns: int
bucket: int = 0
ongoing: bool = False
overlapped: bool = False
class stormlog.infer.diagnosis_loop.Step(iteration, start_wall_ns, start_mono_ns, end_wall_ns, end_mono_ns, completed_wall_ns, completed_mono_ns, members, total_tokens, finished=frozenset({}), streaming=frozenset({}), prompts=())[source]

Bases: object

One scheduler step, as the raw log or an import describes it.

Parameters:
  • iteration (str)

  • start_wall_ns (int)

  • start_mono_ns (int)

  • end_wall_ns (int)

  • end_mono_ns (int)

  • completed_wall_ns (int | None)

  • completed_mono_ns (int | None)

  • members (frozenset[str])

  • total_tokens (int)

  • finished (frozenset[str])

  • streaming (frozenset[str])

  • prompts (tuple[tuple[str, int], ...])

iteration: str
start_wall_ns: int
start_mono_ns: int
end_wall_ns: int
end_mono_ns: int
completed_wall_ns: int | None
completed_mono_ns: int | None
members: frozenset[str]
total_tokens: int
finished: frozenset[str] = frozenset({})
streaming: frozenset[str] = frozenset({})
prompts: tuple[tuple[str, int], ...] = ()
stormlog.infer.diagnosis_loop.cadence_baseline(table, before_mono_ns, config, bucket, ends=None)[source]

The median completion cadence of the busy steps that completed in the window before a stall: those of the stall’s own work bucket when there are enough, else those of its bucket or larger when there are enough, else none (the floor applies). A step is never measured against smaller ones, so a long prefill is not judged by the decode cadence. Only earlier steps count, so the baseline is causal.

Parameters:
  • table (Sequence[tuple[int, int, int]])

  • before_mono_ns (int)

  • config (LoopGapConfig)

  • bucket (int)

  • ends (Sequence[int] | None)

Return type:

tuple[float | None, str]

stormlog.infer.diagnosis_loop.cadence_table(steps)[source]

Each busy step’s completion cadence, as (completion, work bucket, cadence), in completion order: the time since the previous completion, for steps whose work continued from that step.

Parameters:

steps (Sequence[Step])

Return type:

list[tuple[int, int, int]]

stormlog.infer.diagnosis_loop.ready(step)[source]

The members still ready to run after step, with nothing scheduled since: not finished in it, and not a streaming-input request, which may be waiting for its client’s next input.

Parameters:

step (Step)

Return type:

frozenset[str]

stormlog.infer.diagnosis_loop.work_bucket(total_tokens)[source]

Steps that schedule within a factor of two of each other’s tokens share a bucket: a step running a long prefill is compared with steps of its own size, never with decode-only steps.

Parameters:

total_tokens (int)

Return type:

int

stormlog.infer.diagnosis_loop.engine_loop_gap(records, config=None)[source]

The stall with ready work furthest over its limit in one epoch’s records, in seq order: the longest when none is over.

Parameters:
  • records (Sequence[Mapping[str, Any]])

  • config (LoopGapConfig | None)

Return type:

SignalValue

stormlog.infer.diagnosis_loop.find_stalls(steps, blocked, now_wall_ns, terminated=frozenset({}))[source]

Every candidate stall with ready work, in each locus. terminated names requests the engine ended (a terminal record), which no longer hold up a stall going on.

Parameters:
  • steps (Sequence[Step])

  • blocked (Sequence[tuple[int, int]])

  • now_wall_ns (int | None)

  • terminated (frozenset[str])

Return type:

list[Stall]

stormlog.infer.diagnosis_loop.pause_intervals(records, state='PAUSED_ALL')[source]

Wall-clock spans the scheduler spent in state, from the hook’s pause transitions; a pause not yet ended runs to the end of time.

Parameters:
  • records (Sequence[Mapping[str, Any]])

  • state (str)

Return type:

list[tuple[int, int]]

stormlog.infer.diagnosis_loop.record_reasons(records)[source]

Why the records cannot be trusted to show every step: another epoch, a sequence gap, records the writer dropped, or a capped writer.

Parameters:

records (Sequence[Mapping[str, Any]])

Return type:

list[str]

stormlog.infer.diagnosis_loop.steps_from_raw(records)[source]

Join each scheduled record to its completed by iteration.

Parameters:

records (Sequence[Mapping[str, Any]])

Return type:

list[Step]