stormlog.infer.latency_report

Per-case latency quantiles with sufficiency, and chunk-level streaming.

Each latency metric is a criterion key of stormlog.infer.slo, so the report uses the same boundaries as SLO policies: client.* values come from Stormlog’s client, server.* values only from a request’s own joined vLLM span. Every quantile is given twice, over the successful requests and with every request that did not succeed ranked worst, and each says whether the case has enough requests to trust it. What the requests that did not succeed were observed to take is kept apart, by status.

Streamed chunks are reported as chunks. A chunk can carry several tokens, so chunk gaps are never called inter-token latency.

Functions

latency_summary(requests, *[, spans, rule, ...])

Quantiles of each latency metric over a case's measured requests.

streaming_summary(requests)

Chunk counts and gaps of the successful streamed responses.

unsuccessful_summary(requests)

How long each request that did not succeed ran, by status.

stormlog.infer.latency_report.latency_summary(requests, *, spans=None, rule=SufficiencyRule(confidence=0.95, margin=5, tails='symmetric'), levels=(0.5, 0.9, 0.95, 0.99))[source]

Quantiles of each latency metric over a case’s measured requests.

Parameters:
Return type:

dict[str, Any]

stormlog.infer.latency_report.streaming_summary(requests)[source]

Chunk counts and gaps of the successful streamed responses.

Tokens per chunk is a mean over the responses whose output token count the server reported; chunk sizes themselves are not recorded.

Parameters:

requests (Sequence[Mapping[str, Any]])

Return type:

dict[str, Any]

stormlog.infer.latency_report.unsuccessful_summary(requests)[source]

How long each request that did not succeed ran, by status.

This is the time until Stormlog saw the request end, not a latency the request would have had: only for a timeout is it a lower bound on one. A cancelled request’s is its send to its end; a dropped request was never sent and has no time.

Parameters:

requests (Sequence[Mapping[str, Any]])

Return type:

dict[str, dict[str, Any]]