stormlog.infer
Inference profiling helpers for OpenAI-compatible serving endpoints.
Modules
Analysis helpers for inference profiling artifacts. |
|
Per-case arrival accounting for inference reports. |
|
Arrival schedules for inference profiling workloads. |
|
Requested and verified prefix-cache state for each workload case. |
|
CLI for Stormlog inference profiling. |
|
Configuration models for inference profiling. |
|
Resolve inference evidence before computing overlap-aware GPU time. |
|
Optional server evidence capture using existing inference and run artifacts. |
|
Versioned, engine-neutral inference execution records. |
|
Stalls of the vLLM engine loop, from the execution hook's raw records. |
|
Cheap signals that decide whether a window of scrapes looks like an incident. |
|
The versioned threshold table shared by online triggers and the diagnoser. |
|
The closed vocabulary of inference diagnosis findings. |
|
Errors that tell the inference CLI which contract exit code a failure gets. |
|
Inference profiling event records and JSONL persistence. |
|
Identify one host boot when deciding if wall timestamps share a clock. |
|
Per-case latency quantiles with sufficiency, and chunk-level streaming. |
|
Send requests at scheduled times, bounded by an in-flight limit. |
|
OpenAI-compatible Chat Completions client used by inference profiling. |
|
Count an OTLP/HTTP protobuf trace export's messages before it is parsed. |
|
Who a case's numbers are about, and the interval its rates divide by. |
|
Active inference profiling runner. |
|
Deterministic prompts with controlled prefix sharing. |
|
Latency quantiles, how much data they need, and what failures do to them. |
|
Small statistics helpers shared by the inference reports. |
|
Best-effort system telemetry samplers for inference profiling. |
|
Aggregate one window of vLLM |
|
Place server telemetry timestamps on the client artifact's clock. |
|
Optional on-host process and NVML memory collection for inference runs. |
|
Declared groups of server processes, such as tensor-parallel GPU workers. |
|
SLO policies: latency criteria declared at an explicit boundary. |
|
Engine-neutral, scoped samples from an inference host. |
|
Token counting helpers for inference profiling. |
|
Bounded vLLM profiler windows during |
|
Import profiler traces (Kineto, Nsight Systems) into an inference artifact. |
|
Import PyTorch/Kineto profiler traces as inference activity references. |
|
Import Nsight Systems SQLite exports as inference activity references. |
|
Name iteration ranges so a trace importer can link GPU work to iterations. |
|
Bounded in-process PyTorch profiler capture for import with |
|
Explain a run's queueing, cache and token signals from vLLM's own telemetry. |
|
Reduce the vLLM execution hook's raw log into canonical correlation records. |
|
Map a profiler trace to the GPU it ran on, from the vLLM hook's worker hellos. |
|
Import a vLLM execution-hook log into an inference artifact. |
|
Read the vLLM execution hook's raw log ( |
|
Coverage of an imported vLLM execution log, as separate dimensions. |
|
Install the vLLM execution hook's patches; see |
|
Prometheus text parsing and the vLLM metric catalog. |
|
Scrape vLLM's |
|
Ingest vLLM's OpenTelemetry request spans. |
|
Records for vLLM native telemetry inside an inference artifact. |
|
The workload record: what traffic a run sent, so it can be repeated. |
|
Per-case prompt and length summaries for inference reports. |