stormlog.infer.slo
SLO policies: latency criteria declared at an explicit boundary.
A criterion is keyed boundary.metric. client criteria are measured by
Stormlog’s own client; server criteria are what vLLM reports, per request
through its span attributes or in aggregate through its histograms. The two
are never merged or relabelled. There is no client inter-token latency:
streamed chunks are not tokens.
A policy is a versioned JSON document (stormlog.infer.slo v1) or a list
of KEY:MS flags. infer profile --slo records it in the artifact as an
infer.slo record, so an artifact can carry its own declared SLO.
Functions
|
The client criteria's values for one |
|
Judge each criterion of |
|
Judge one |
|
Judge an engine-finished span against the policy's server criteria. |
|
Read a policy file; a file that cannot be read or is invalid exits 5. |
|
Build a policy from |
|
A request's trusted span, or None and the reason it has none. |
|
Refuse a sliding policy where each case is judged over its interval. |
|
The per-request server criteria's values from a span's attributes. |
|
True when met, False when missed, None when it cannot be judged. |
|
The policy an artifact declared, or None when it declared none. |
|
Validate a |
|
The |
|
Each measured request's trusted vLLM span attributes, by |
Classes
|
The criteria alone, with no claim about whether the service succeeded. |
|
One latency limit: a value passes when it is at most |
|
One criterion over a case: outcomes among successful requests, and its marginal attainment bounds over every offered request. |
|
What a criterion measures, and where its value comes from. |
|
How one criterion judged one request or span. |
|
A value to judge in milliseconds, or the reason there is none. |
|
A client request judged against a policy. |
|
A case judged against a policy ( |
|
The interval attainment and goodput are judged over. |
|
A named SLO policy. |
|
An engine-finished span judged against a policy's server criteria. |
- class stormlog.infer.slo.CriteriaOutcome(outcome, criteria)[source]
Bases:
objectThe criteria alone, with no claim about whether the service succeeded.
- Parameters:
outcome (Literal['criteria_met', 'criteria_missed', 'unknown'])
criteria (Mapping[str, CriterionOutcome])
- outcome: Literal['criteria_met', 'criteria_missed', 'unknown']
- criteria: Mapping[str, CriterionOutcome]
- class stormlog.infer.slo.Criterion(metric, boundary, max_ms, attainment_target=None)[source]
Bases:
objectOne latency limit: a value passes when it is at most
max_ms.- Parameters:
metric (str)
boundary (Literal['client', 'server'])
max_ms (float)
attainment_target (float | None)
- metric: str
- boundary: Literal['client', 'server']
- max_ms: float
- attainment_target: float | None = None
- property key: str
- property definition: CriterionDef
- class stormlog.infer.slo.CriterionCounts(passed, failed, not_applicable, unknown, attainment_lower, attainment_upper)[source]
Bases:
objectOne criterion over a case: outcomes among successful requests, and its marginal attainment bounds over every offered request.
- Parameters:
passed (int)
failed (int)
not_applicable (int)
unknown (int)
attainment_lower (float | None)
attainment_upper (float | None)
- passed: int
- failed: int
- not_applicable: int
- unknown: int
- attainment_lower: float | None
- attainment_upper: float | None
- class stormlog.infer.slo.CriterionDef(boundary, metric, definition, per_request, aggregate)[source]
Bases:
objectWhat a criterion measures, and where its value comes from.
- Parameters:
boundary (Literal['client', 'server'])
metric (str)
definition (str)
per_request (str | None)
aggregate (str | None)
- boundary: Literal['client', 'server']
- metric: str
- definition: str
- per_request: str | None
- aggregate: str | None
- property key: str
- class stormlog.infer.slo.CriterionOutcome(outcome, value_ms, max_ms, reason=None)[source]
Bases:
objectHow one criterion judged one request or span.
- Parameters:
outcome (Literal['pass', 'fail', 'not_applicable', 'unknown'])
value_ms (float | None)
max_ms (float)
reason (str | None)
- outcome: Literal['pass', 'fail', 'not_applicable', 'unknown']
- value_ms: float | None
- max_ms: float
- reason: str | None = None
- class stormlog.infer.slo.CriterionValue(value_ms, reason=None, not_applicable=False)[source]
Bases:
objectA value to judge in milliseconds, or the reason there is none.
- Parameters:
value_ms (float | None)
reason (str | None)
not_applicable (bool)
- value_ms: float | None
- reason: str | None = None
- not_applicable: bool = False
- class stormlog.infer.slo.RequestSloOutcome(outcome, status, criteria)[source]
Bases:
objectA client request judged against a policy.
metneeds a successful request and every criterion passing; any other status ismissed; a successful request with a criterion that cannot be judged, and none failing, isunknown.- Parameters:
outcome (Literal['met', 'missed', 'unknown'])
status (str)
criteria (Mapping[str, CriterionOutcome])
- outcome: Literal['met', 'missed', 'unknown']
- status: str
- criteria: Mapping[str, CriterionOutcome]
- property met: bool | None
- class stormlog.infer.slo.SpanSloOutcome(outcome, criteria, service_success='unverified')[source]
Bases:
objectAn engine-finished span judged against a policy’s server criteria.
A vLLM span carries no finish reason, and vLLM emits one for aborted and failed requests too, so a span shows whether the criteria were met, never whether the request succeeded.
- Parameters:
outcome (Literal['criteria_met', 'criteria_missed', 'unknown'])
criteria (Mapping[str, CriterionOutcome])
service_success (Literal['unverified'])
- outcome: Literal['criteria_met', 'criteria_missed', 'unknown']
- criteria: Mapping[str, CriterionOutcome]
- service_success: Literal['unverified'] = 'unverified'
- class stormlog.infer.slo.SloEvaluation(slo_name, slo_digest, slo_source, status, reason, population_declared, population_evaluated, offered, met, missed, unknown, attainment_lower, attainment_upper, evidence_coverage, per_criterion, interval, goodput_lower_rps, goodput_upper_rps, goodput_lower_output_tps, cohort_valid=None, cohort_issues=())[source]
Bases:
objectA case judged against a policy (
stormlog.infer.slo_evaluationv1).attainment_lowercounts unknown outcomes as missed andattainment_upperas met, so missing evidence widens the bounds and never moves a single figure. Goodput is SLO goodput at the offered load: good requests per second of the case’s rate interval, judged by vLLM’s per-request rule, not the highest rate that meets a target. For an open loop the interval is the arrival window without the drain, where vLLM’s benchmark divides by its whole duration. An unmeasurable evaluation (a criterion no successful request could be judged on, or one that is aggregate-only) has no attainment or goodput, never zero.cohort_validandcohort_issuesrepeat the case’s cohort checks (None when the caller gave no cohort): the figures are over the records the run holds, which an invalid cohort does not vouch for.- Parameters:
slo_name (str)
slo_digest (str)
slo_source (str)
status (Literal['evaluated', 'unmeasurable'])
reason (str | None)
population_declared (str)
population_evaluated (str)
offered (int)
met (int)
missed (int)
unknown (int)
attainment_lower (float | None)
attainment_upper (float | None)
evidence_coverage (float | None)
per_criterion (Mapping[str, CriterionCounts])
interval (MeasuredInterval | None)
goodput_lower_rps (float | None)
goodput_upper_rps (float | None)
goodput_lower_output_tps (float | None)
cohort_valid (bool | None)
cohort_issues (tuple[str, ...])
- slo_name: str
- slo_digest: str
- slo_source: str
- status: Literal['evaluated', 'unmeasurable']
- reason: str | None
- population_declared: str
- population_evaluated: str
- offered: int
- met: int
- missed: int
- unknown: int
- attainment_lower: float | None
- attainment_upper: float | None
- evidence_coverage: float | None
- per_criterion: Mapping[str, CriterionCounts]
- interval: MeasuredInterval | None
- goodput_lower_rps: float | None
- goodput_upper_rps: float | None
- goodput_lower_output_tps: float | None
- cohort_valid: bool | None = None
- cohort_issues: tuple[str, ...] = ()
- class stormlog.infer.slo.SloInterval(kind='measured_window', seconds=None)[source]
Bases:
objectThe interval attainment and goodput are judged over.
measured_windowis a case’s declared interval, for offline analysis;slidingis a window ofseconds, for an online watcher.- Parameters:
kind (Literal['measured_window', 'sliding'])
seconds (float | None)
- kind: Literal['measured_window', 'sliding'] = 'measured_window'
- seconds: float | None = None
- class stormlog.infer.slo.SloSpec(name, criteria, attainment_target=None, interval=SloInterval(kind='measured_window', seconds=None), population='offered', unknown_policy='bounds')[source]
Bases:
objectA named SLO policy.
attainment_targetis joint: every criterion passes. Each criterion’s own target is marginal. Version 1 has one population,offered, and one unknown policy,bounds: an outcome that cannot be judged widens the reported bounds and is never silently counted as met or missed.- Parameters:
name (str)
criteria (tuple[Criterion, ...])
attainment_target (float | None)
interval (SloInterval)
population (Literal['offered'])
unknown_policy (Literal['bounds'])
- name: str
- attainment_target: float | None = None
- interval: SloInterval = SloInterval(kind='measured_window', seconds=None)
- population: Literal['offered'] = 'offered'
- unknown_policy: Literal['bounds'] = 'bounds'
- stormlog.infer.slo.load_slo(path)[source]
Read a policy file; a file that cannot be read or is invalid exits 5.
- Parameters:
path (str | Path)
- Return type:
- stormlog.infer.slo.parse_slo_flags(items, *, name='cli')[source]
Build a policy from
KEY:MSflags such asttft:500.A key without a boundary is a client criterion;
server.ttft:400names the server one. Flags have no targets and judge the measured window.- Raises:
InferUsageError – for a flag that is not
KEY:MSor names no criterion, a repeated key, or a limit that is not a positive number.- Parameters:
items (Sequence[str])
name (str)
- Return type:
- stormlog.infer.slo.request_span(record, spans)[source]
A request’s trusted span, or None and the reason it has none.
- Parameters:
record (Mapping[str, Any])
spans (JoinedSpans)
- Return type:
tuple[Mapping[str, Any] | None, str]
- stormlog.infer.slo.require_measured_window(spec, label)[source]
Refuse a sliding policy where each case is judged over its interval.
- Raises:
InferInputError – for a sliding interval, which an online watcher judges; offline analysis would judge it over the whole case.
- Parameters:
spec (SloSpec)
label (str)
- Return type:
- stormlog.infer.slo.span_attributes_by_request(records, span_paths=())[source]
Each measured request’s trusted vLLM span attributes, by
x_request_id.Spans come from the artifact and from
span_paths. A request whose span arrived again with different content, or that has several spans, is left out and listed inquarantinedwith the reason; pass that reason toevaluate_requestasmissing_span_reason.- Raises:
InferInputError – for a span file or span record that cannot be read.
- Parameters:
records (Sequence[Mapping[str, Any]])
span_paths (Sequence[str | Path])
- Return type:
- stormlog.infer.slo.slo_from_artifact(records)[source]
The policy an artifact declared, or None when it declared none.
- Raises:
InferInputError – when the artifact holds more than one
infer.slorecord, or one that is not a valid policy.- Parameters:
records (Sequence[Mapping[str, Any]])
- Return type:
SloSpec | None
- stormlog.infer.slo.slo_from_document(payload)[source]
Validate a
stormlog.infer.slov1 document.- Raises:
ValueError – naming the first rule the document breaks.
- Parameters:
payload (Any)
- Return type:
- stormlog.infer.slo.slo_record(spec, *, session_id, source)[source]
The
infer.slorecordinfer profile --slowrites.- Parameters:
spec (SloSpec)
session_id (str)
source (str)
- Return type:
dict[str, Any]
- stormlog.infer.slo.client_values(record)[source]
The client criteria’s values for one
infer.requestrecord.- Parameters:
record (Mapping[str, Any])
- Return type:
dict[str, CriterionValue]
- stormlog.infer.slo.evaluate_criteria(values, spec, *, boundary='both')[source]
Judge each criterion of
specagainstvalueskeyed by criterion.A criterion outside
boundaryis unknown, as is one with no value or a value that is not finite. A failing criterion makes the whole verdictcriteria_missed; otherwise an unknown one makes itunknown.- Parameters:
values (Mapping[str, float | CriterionValue | None])
spec (SloSpec)
boundary (Literal['client', 'server', 'both'])
- Return type:
- stormlog.infer.slo.evaluate_request(record, spec, *, span=None, missing_span_reason='no_joined_span')[source]
Judge one
infer.requestrecord;spanis its joined vLLM span.Client criteria read the record and server criteria read only
span’s attributes, so a missing span leaves the server criteria unknown, withmissing_span_reason, and is never filled from client values. A request that did not succeed ismissedwhatever its criteria say; they are still judged on what was recorded, for diagnosis.- Parameters:
record (Mapping[str, Any])
spec (SloSpec)
span (Mapping[str, Any] | None)
missing_span_reason (str)
- Return type:
- stormlog.infer.slo.evaluate_span(span_attributes, spec)[source]
Judge an engine-finished span against the policy’s server criteria.
Client criteria are unknown on a span. The outcome is
criteria_met,criteria_missedorunknown, nevermet: success is unverified.- Parameters:
span_attributes (Mapping[str, Any])
spec (SloSpec)
- Return type:
- stormlog.infer.slo.server_values(span_attributes, *, missing_span_reason='no_joined_span')[source]
The per-request server criteria’s values from a span’s attributes.
- Parameters:
span_attributes (Mapping[str, Any] | None)
missing_span_reason (str)
- Return type:
dict[str, CriterionValue]