stormlog.infer

Inference profiling helpers for OpenAI-compatible serving endpoints.

Modules

analysis

Analysis helpers for inference profiling artifacts.

arrival_report

Per-case arrival accounting for inference reports.

arrivals

Arrival schedules for inference profiling workloads.

cache_state

Requested and verified prefix-cache state for each workload case.

cli

CLI for Stormlog inference profiling.

config

Configuration models for inference profiling.

correlation_accounting

Resolve inference evidence before computing overlap-aware GPU time.

correlation_capture

Optional server evidence capture using existing inference and run artifacts.

correlation_events

Versioned, engine-neutral inference execution records.

diagnosis_loop

Stalls of the vLLM engine loop, from the execution hook's raw records.

diagnosis_signals

Cheap signals that decide whether a window of scrapes looks like an incident.

diagnosis_thresholds

The versioned threshold table shared by online triggers and the diagnoser.

diagnosis_vocabulary

The closed vocabulary of inference diagnosis findings.

errors

Errors that tell the inference CLI which contract exit code a failure gets.

events

Inference profiling event records and JSONL persistence.

host_clock

Identify one host boot when deciding if wall timestamps share a clock.

latency_report

Per-case latency quantiles with sufficiency, and chunk-level streaming.

open_loop

Send requests at scheduled times, bounded by an in-flight limit.

openai_client

OpenAI-compatible Chat Completions client used by inference profiling.

otlp_wire

Count an OTLP/HTTP protobuf trace export's messages before it is parsed.

populations

Who a case's numbers are about, and the interval its rates divide by.

profile

Active inference profiling runner.

prompts

Deterministic prompts with controlled prefix sharing.

quantiles

Latency quantiles, how much data they need, and what failures do to them.

report_stats

Small statistics helpers shared by the inference reports.

samplers

Best-effort system telemetry samplers for inference profiling.

scrape_window

Aggregate one window of vLLM /metrics scrapes for one engine.

server_clock

Place server telemetry timestamps on the client artifact's clock.

server_collector

Optional on-host process and NVML memory collection for inference runs.

server_group

Declared groups of server processes, such as tensor-parallel GPU workers.

slo

SLO policies: latency criteria declared at an explicit boundary.

telemetry

Engine-neutral, scoped samples from an inference host.

tokens

Token counting helpers for inference profiling.

trace_capture

Bounded vLLM profiler windows during stormlog infer profile.

trace_import

Import profiler traces (Kineto, Nsight Systems) into an inference artifact.

trace_kineto

Import PyTorch/Kineto profiler traces as inference activity references.

trace_nsys

Import Nsight Systems SQLite exports as inference activity references.

trace_ranges

Name iteration ranges so a trace importer can link GPU work to iterations.

trace_torch

Bounded in-process PyTorch profiler capture for import with import-trace.

vllm_analysis

Explain a run's queueing, cache and token signals from vLLM's own telemetry.

vllm_execution

Reduce the vLLM execution hook's raw log into canonical correlation records.

vllm_execution_devices

Map a profiler trace to the GPU it ran on, from the vLLM hook's worker hellos.

vllm_execution_import

Import a vLLM execution-hook log into an inference artifact.

vllm_execution_log

Read the vLLM execution hook's raw log (docs/vllm_execution.md).

vllm_execution_report

Coverage of an imported vLLM execution log, as separate dimensions.

vllm_hook

Install the vLLM execution hook's patches; see docs/vllm_execution.md.

vllm_metrics

Prometheus text parsing and the vLLM metric catalog.

vllm_scraper

Scrape vLLM's /metrics during a profile and keep each response whole.

vllm_spans

Ingest vLLM's OpenTelemetry request spans.

vllm_telemetry

Records for vLLM native telemetry inside an inference artifact.

workload

The workload record: what traffic a run sent, so it can be repeated.

workload_report

Per-case prompt and length summaries for inference reports.