stormlog.infer.slo

SLO policies: latency criteria declared at an explicit boundary.

A criterion is keyed boundary.metric. client criteria are measured by Stormlog’s own client; server criteria are what vLLM reports, per request through its span attributes or in aggregate through its histograms. The two are never merged or relabelled. There is no client inter-token latency: streamed chunks are not tokens.

A policy is a versioned JSON document (stormlog.infer.slo v1) or a list of KEY:MS flags. infer profile --slo records it in the artifact as an infer.slo record, so an artifact can carry its own declared SLO.

Functions

client_values(record)

The client criteria's values for one infer.request record.

evaluate_criteria(values, spec, *[, boundary])

Judge each criterion of spec against values keyed by criterion.

evaluate_request(record, spec, *[, span, ...])

Judge one infer.request record; span is its joined vLLM span.

evaluate_span(span_attributes, spec)

Judge an engine-finished span against the policy's server criteria.

load_slo(path)

Read a policy file; a file that cannot be read or is invalid exits 5.

parse_slo_flags(items, *[, name])

Build a policy from KEY:MS flags such as ttft:500.

request_span(record, spans)

A request's trusted span, or None and the reason it has none.

require_measured_window(spec, label)

Refuse a sliding policy where each case is judged over its interval.

server_values(span_attributes, *[, ...])

The per-request server criteria's values from a span's attributes.

slo_attained(record, spec, *[, span])

True when met, False when missed, None when it cannot be judged.

slo_from_artifact(records)

The policy an artifact declared, or None when it declared none.

slo_from_document(payload)

Validate a stormlog.infer.slo v1 document.

slo_record(spec, *, session_id, source)

The infer.slo record infer profile --slo writes.

span_attributes_by_request(records[, span_paths])

Each measured request's trusted vLLM span attributes, by x_request_id.

Classes

CriteriaOutcome(outcome, criteria)

The criteria alone, with no claim about whether the service succeeded.

Criterion(metric, boundary, max_ms[, ...])

One latency limit: a value passes when it is at most max_ms.

CriterionCounts(passed, failed, ...)

One criterion over a case: outcomes among successful requests, and its marginal attainment bounds over every offered request.

CriterionDef(boundary, metric, definition, ...)

What a criterion measures, and where its value comes from.

CriterionOutcome(outcome, value_ms, max_ms)

How one criterion judged one request or span.

CriterionValue(value_ms[, reason, ...])

A value to judge in milliseconds, or the reason there is none.

RequestSloOutcome(outcome, status, criteria)

A client request judged against a policy.

SloEvaluation(slo_name, slo_digest, ...[, ...])

A case judged against a policy (stormlog.infer.slo_evaluation v1).

SloInterval([kind, seconds])

The interval attainment and goodput are judged over.

SloSpec(name, criteria[, attainment_target, ...])

A named SLO policy.

SpanSloOutcome(outcome, criteria[, ...])

An engine-finished span judged against a policy's server criteria.

class stormlog.infer.slo.CriteriaOutcome(outcome, criteria)[source]

Bases: object

The criteria alone, with no claim about whether the service succeeded.

Parameters:
  • outcome (Literal['criteria_met', 'criteria_missed', 'unknown'])

  • criteria (Mapping[str, CriterionOutcome])

outcome: Literal['criteria_met', 'criteria_missed', 'unknown']
criteria: Mapping[str, CriterionOutcome]
class stormlog.infer.slo.Criterion(metric, boundary, max_ms, attainment_target=None)[source]

Bases: object

One latency limit: a value passes when it is at most max_ms.

Parameters:
  • metric (str)

  • boundary (Literal['client', 'server'])

  • max_ms (float)

  • attainment_target (float | None)

metric: str
boundary: Literal['client', 'server']
max_ms: float
attainment_target: float | None = None
property key: str
property definition: CriterionDef
to_record()[source]
Return type:

dict[str, Any]

class stormlog.infer.slo.CriterionCounts(passed, failed, not_applicable, unknown, attainment_lower, attainment_upper)[source]

Bases: object

One criterion over a case: outcomes among successful requests, and its marginal attainment bounds over every offered request.

Parameters:
  • passed (int)

  • failed (int)

  • not_applicable (int)

  • unknown (int)

  • attainment_lower (float | None)

  • attainment_upper (float | None)

passed: int
failed: int
not_applicable: int
unknown: int
attainment_lower: float | None
attainment_upper: float | None
to_record()[source]
Return type:

dict[str, Any]

class stormlog.infer.slo.CriterionDef(boundary, metric, definition, per_request, aggregate)[source]

Bases: object

What a criterion measures, and where its value comes from.

Parameters:
  • boundary (Literal['client', 'server'])

  • metric (str)

  • definition (str)

  • per_request (str | None)

  • aggregate (str | None)

boundary: Literal['client', 'server']
metric: str
definition: str
per_request: str | None
aggregate: str | None
property key: str
class stormlog.infer.slo.CriterionOutcome(outcome, value_ms, max_ms, reason=None)[source]

Bases: object

How one criterion judged one request or span.

Parameters:
  • outcome (Literal['pass', 'fail', 'not_applicable', 'unknown'])

  • value_ms (float | None)

  • max_ms (float)

  • reason (str | None)

outcome: Literal['pass', 'fail', 'not_applicable', 'unknown']
value_ms: float | None
max_ms: float
reason: str | None = None
class stormlog.infer.slo.CriterionValue(value_ms, reason=None, not_applicable=False)[source]

Bases: object

A value to judge in milliseconds, or the reason there is none.

Parameters:
  • value_ms (float | None)

  • reason (str | None)

  • not_applicable (bool)

value_ms: float | None
reason: str | None = None
not_applicable: bool = False
class stormlog.infer.slo.RequestSloOutcome(outcome, status, criteria)[source]

Bases: object

A client request judged against a policy.

met needs a successful request and every criterion passing; any other status is missed; a successful request with a criterion that cannot be judged, and none failing, is unknown.

Parameters:
  • outcome (Literal['met', 'missed', 'unknown'])

  • status (str)

  • criteria (Mapping[str, CriterionOutcome])

outcome: Literal['met', 'missed', 'unknown']
status: str
criteria: Mapping[str, CriterionOutcome]
property met: bool | None
class stormlog.infer.slo.SpanSloOutcome(outcome, criteria, service_success='unverified')[source]

Bases: object

An engine-finished span judged against a policy’s server criteria.

A vLLM span carries no finish reason, and vLLM emits one for aborted and failed requests too, so a span shows whether the criteria were met, never whether the request succeeded.

Parameters:
  • outcome (Literal['criteria_met', 'criteria_missed', 'unknown'])

  • criteria (Mapping[str, CriterionOutcome])

  • service_success (Literal['unverified'])

outcome: Literal['criteria_met', 'criteria_missed', 'unknown']
criteria: Mapping[str, CriterionOutcome]
service_success: Literal['unverified'] = 'unverified'
class stormlog.infer.slo.SloEvaluation(slo_name, slo_digest, slo_source, status, reason, population_declared, population_evaluated, offered, met, missed, unknown, attainment_lower, attainment_upper, evidence_coverage, per_criterion, interval, goodput_lower_rps, goodput_upper_rps, goodput_lower_output_tps, cohort_valid=None, cohort_issues=())[source]

Bases: object

A case judged against a policy (stormlog.infer.slo_evaluation v1).

attainment_lower counts unknown outcomes as missed and attainment_upper as met, so missing evidence widens the bounds and never moves a single figure. Goodput is SLO goodput at the offered load: good requests per second of the case’s rate interval, judged by vLLM’s per-request rule, not the highest rate that meets a target. For an open loop the interval is the arrival window without the drain, where vLLM’s benchmark divides by its whole duration. An unmeasurable evaluation (a criterion no successful request could be judged on, or one that is aggregate-only) has no attainment or goodput, never zero. cohort_valid and cohort_issues repeat the case’s cohort checks (None when the caller gave no cohort): the figures are over the records the run holds, which an invalid cohort does not vouch for.

Parameters:
  • slo_name (str)

  • slo_digest (str)

  • slo_source (str)

  • status (Literal['evaluated', 'unmeasurable'])

  • reason (str | None)

  • population_declared (str)

  • population_evaluated (str)

  • offered (int)

  • met (int)

  • missed (int)

  • unknown (int)

  • attainment_lower (float | None)

  • attainment_upper (float | None)

  • evidence_coverage (float | None)

  • per_criterion (Mapping[str, CriterionCounts])

  • interval (MeasuredInterval | None)

  • goodput_lower_rps (float | None)

  • goodput_upper_rps (float | None)

  • goodput_lower_output_tps (float | None)

  • cohort_valid (bool | None)

  • cohort_issues (tuple[str, ...])

slo_name: str
slo_digest: str
slo_source: str
status: Literal['evaluated', 'unmeasurable']
reason: str | None
population_declared: str
population_evaluated: str
offered: int
met: int
missed: int
unknown: int
attainment_lower: float | None
attainment_upper: float | None
evidence_coverage: float | None
per_criterion: Mapping[str, CriterionCounts]
interval: MeasuredInterval | None
goodput_lower_rps: float | None
goodput_upper_rps: float | None
goodput_lower_output_tps: float | None
cohort_valid: bool | None = None
cohort_issues: tuple[str, ...] = ()
to_record()[source]
Return type:

dict[str, Any]

class stormlog.infer.slo.SloInterval(kind='measured_window', seconds=None)[source]

Bases: object

The interval attainment and goodput are judged over.

measured_window is a case’s declared interval, for offline analysis; sliding is a window of seconds, for an online watcher.

Parameters:
  • kind (Literal['measured_window', 'sliding'])

  • seconds (float | None)

kind: Literal['measured_window', 'sliding'] = 'measured_window'
seconds: float | None = None
to_record()[source]
Return type:

dict[str, Any]

class stormlog.infer.slo.SloSpec(name, criteria, attainment_target=None, interval=SloInterval(kind='measured_window', seconds=None), population='offered', unknown_policy='bounds')[source]

Bases: object

A named SLO policy.

attainment_target is joint: every criterion passes. Each criterion’s own target is marginal. Version 1 has one population, offered, and one unknown policy, bounds: an outcome that cannot be judged widens the reported bounds and is never silently counted as met or missed.

Parameters:
  • name (str)

  • criteria (tuple[Criterion, ...])

  • attainment_target (float | None)

  • interval (SloInterval)

  • population (Literal['offered'])

  • unknown_policy (Literal['bounds'])

name: str
criteria: tuple[Criterion, ...]
attainment_target: float | None = None
interval: SloInterval = SloInterval(kind='measured_window', seconds=None)
population: Literal['offered'] = 'offered'
unknown_policy: Literal['bounds'] = 'bounds'
to_record()[source]

The policy as its versioned JSON document.

Return type:

dict[str, Any]

digest()[source]

SHA-256 of the canonical document.

500 and 500.0 digest alike, and so do the same criteria in another order: they are joined by AND.

Return type:

str

criterion(key)[source]
Parameters:

key (str)

Return type:

Criterion | None

stormlog.infer.slo.load_slo(path)[source]

Read a policy file; a file that cannot be read or is invalid exits 5.

Parameters:

path (str | Path)

Return type:

SloSpec

stormlog.infer.slo.parse_slo_flags(items, *, name='cli')[source]

Build a policy from KEY:MS flags such as ttft:500.

A key without a boundary is a client criterion; server.ttft:400 names the server one. Flags have no targets and judge the measured window.

Raises:

InferUsageError – for a flag that is not KEY:MS or names no criterion, a repeated key, or a limit that is not a positive number.

Parameters:
  • items (Sequence[str])

  • name (str)

Return type:

SloSpec

stormlog.infer.slo.request_span(record, spans)[source]

A request’s trusted span, or None and the reason it has none.

Parameters:
Return type:

tuple[Mapping[str, Any] | None, str]

stormlog.infer.slo.require_measured_window(spec, label)[source]

Refuse a sliding policy where each case is judged over its interval.

Raises:

InferInputError – for a sliding interval, which an online watcher judges; offline analysis would judge it over the whole case.

Parameters:
Return type:

SloSpec

stormlog.infer.slo.span_attributes_by_request(records, span_paths=())[source]

Each measured request’s trusted vLLM span attributes, by x_request_id.

Spans come from the artifact and from span_paths. A request whose span arrived again with different content, or that has several spans, is left out and listed in quarantined with the reason; pass that reason to evaluate_request as missing_span_reason.

Raises:

InferInputError – for a span file or span record that cannot be read.

Parameters:
  • records (Sequence[Mapping[str, Any]])

  • span_paths (Sequence[str | Path])

Return type:

JoinedSpans

stormlog.infer.slo.slo_from_artifact(records)[source]

The policy an artifact declared, or None when it declared none.

Raises:

InferInputError – when the artifact holds more than one infer.slo record, or one that is not a valid policy.

Parameters:

records (Sequence[Mapping[str, Any]])

Return type:

SloSpec | None

stormlog.infer.slo.slo_from_document(payload)[source]

Validate a stormlog.infer.slo v1 document.

Raises:

ValueError – naming the first rule the document breaks.

Parameters:

payload (Any)

Return type:

SloSpec

stormlog.infer.slo.slo_record(spec, *, session_id, source)[source]

The infer.slo record infer profile --slo writes.

Parameters:
  • spec (SloSpec)

  • session_id (str)

  • source (str)

Return type:

dict[str, Any]

stormlog.infer.slo.client_values(record)[source]

The client criteria’s values for one infer.request record.

Parameters:

record (Mapping[str, Any])

Return type:

dict[str, CriterionValue]

stormlog.infer.slo.evaluate_criteria(values, spec, *, boundary='both')[source]

Judge each criterion of spec against values keyed by criterion.

A criterion outside boundary is unknown, as is one with no value or a value that is not finite. A failing criterion makes the whole verdict criteria_missed; otherwise an unknown one makes it unknown.

Parameters:
  • values (Mapping[str, float | CriterionValue | None])

  • spec (SloSpec)

  • boundary (Literal['client', 'server', 'both'])

Return type:

CriteriaOutcome

stormlog.infer.slo.evaluate_request(record, spec, *, span=None, missing_span_reason='no_joined_span')[source]

Judge one infer.request record; span is its joined vLLM span.

Client criteria read the record and server criteria read only span’s attributes, so a missing span leaves the server criteria unknown, with missing_span_reason, and is never filled from client values. A request that did not succeed is missed whatever its criteria say; they are still judged on what was recorded, for diagnosis.

Parameters:
  • record (Mapping[str, Any])

  • spec (SloSpec)

  • span (Mapping[str, Any] | None)

  • missing_span_reason (str)

Return type:

RequestSloOutcome

stormlog.infer.slo.evaluate_span(span_attributes, spec)[source]

Judge an engine-finished span against the policy’s server criteria.

Client criteria are unknown on a span. The outcome is criteria_met, criteria_missed or unknown, never met: success is unverified.

Parameters:
  • span_attributes (Mapping[str, Any])

  • spec (SloSpec)

Return type:

SpanSloOutcome

stormlog.infer.slo.server_values(span_attributes, *, missing_span_reason='no_joined_span')[source]

The per-request server criteria’s values from a span’s attributes.

Parameters:
  • span_attributes (Mapping[str, Any] | None)

  • missing_span_reason (str)

Return type:

dict[str, CriterionValue]

stormlog.infer.slo.slo_attained(record, spec, *, span=None)[source]

True when met, False when missed, None when it cannot be judged.

Parameters:
  • record (Mapping[str, Any])

  • spec (SloSpec)

  • span (Mapping[str, Any] | None)

Return type:

bool | None