stormlog.infer.config

Configuration models for inference profiling.

Functions

parse_float_list(value, *, field_name)

Parse a comma-separated list of positive, finite numbers.

parse_int_list(value, *, field_name)

Parse a comma-separated positive integer list.

resolve_endpoint(*, endpoint, base_url)

Resolve either a full chat-completions endpoint or a /v1 base URL.

Classes

ProfileConfig(endpoint, model, concurrency, ...)

Resolved configuration for one inference profiling run.

WorkloadCase(case_id, concurrency, ...[, ...])

One inference profiling workload shape.

stormlog.infer.config.parse_int_list(value, *, field_name)[source]

Parse a comma-separated positive integer list.

Parameters:
  • value (str)

  • field_name (str)

Return type:

list[int]

stormlog.infer.config.resolve_endpoint(*, endpoint, base_url)[source]

Resolve either a full chat-completions endpoint or a /v1 base URL.

Parameters:
  • endpoint (str | None)

  • base_url (str | None)

Return type:

str

stormlog.infer.config.parse_float_list(value, *, field_name)[source]

Parse a comma-separated list of positive, finite numbers.

Parameters:
  • value (str)

  • field_name (str)

Return type:

list[float]

class stormlog.infer.config.WorkloadCase(case_id, concurrency, input_tokens, output_tokens, arrival=<factory>)[source]

Bases: object

One inference profiling workload shape.

For an open-loop case, concurrency is the in-flight limit: the most requests that may be outstanding at once.

Parameters:
  • case_id (str)

  • concurrency (int)

  • input_tokens (int)

  • output_tokens (int)

  • arrival (ArrivalSpec)

case_id: str
concurrency: int
input_tokens: int
output_tokens: int
arrival: ArrivalSpec
class stormlog.infer.config.ProfileConfig(endpoint, model, concurrency, input_tokens, output_tokens, output_path, duration_seconds=None, request_count=1, stream=True, stream_include_usage=True, timeout_seconds=60.0, warmup_requests=0, seed=0, api_key=None, max_tokens_field='max_tokens', tokenizer='auto', tokenizer_model=None, tiktoken_encoding=None, strict_token_counts=False, system_sampler='auto', sample_interval_seconds=1.0, run_id=None, arrival_mode='closed', rates=(), burst_size=None, burst_interval_seconds=None, arrival_trace=None, max_in_flight=128, overflow='wait', drain_timeout_seconds=None, prompt_mode='repeat', shared_prefix_ratio=None, prefix_groups=None, cache_state='unspecified', extra_body=None, cache_reset_url=None, cache_reset_timeout_seconds=10.0, vllm_metrics_url=None, vllm_metrics_interval_seconds=1.0, vllm_spans_listen=None, vllm_spans_drain_seconds=6.0, trace=None, vllm_execution_dir=None, slo=None, slo_source=None)[source]

Bases: object

Resolved configuration for one inference profiling run.

Parameters:
  • endpoint (str)

  • model (str)

  • concurrency (tuple[int, ...])

  • input_tokens (tuple[int, ...])

  • output_tokens (tuple[int, ...])

  • output_path (str)

  • duration_seconds (float | None)

  • request_count (int | None)

  • stream (bool)

  • stream_include_usage (bool)

  • timeout_seconds (float)

  • warmup_requests (int)

  • seed (int)

  • api_key (str | None)

  • max_tokens_field (Literal['max_tokens', 'max_completion_tokens'])

  • tokenizer (str)

  • tokenizer_model (str | None)

  • tiktoken_encoding (str | None)

  • strict_token_counts (bool)

  • system_sampler (str)

  • sample_interval_seconds (float)

  • run_id (str | None)

  • arrival_mode (str)

  • rates (tuple[float, ...])

  • burst_size (int | None)

  • burst_interval_seconds (float | None)

  • arrival_trace (ArrivalTrace | None)

  • max_in_flight (int)

  • overflow (Literal['wait', 'drop'])

  • drain_timeout_seconds (float | None)

  • prompt_mode (str)

  • shared_prefix_ratio (float | None)

  • prefix_groups (int | None)

  • cache_state (str)

  • extra_body (dict[str, Any] | None)

  • cache_reset_url (str | None)

  • cache_reset_timeout_seconds (float)

  • vllm_metrics_url (str | None)

  • vllm_metrics_interval_seconds (float)

  • vllm_spans_listen (str | None)

  • vllm_spans_drain_seconds (float)

  • trace (TraceCaptureConfig | None)

  • vllm_execution_dir (Path | None)

  • slo (SloSpec | None)

  • slo_source (str | None)

endpoint: str
model: str
concurrency: tuple[int, ...]
input_tokens: tuple[int, ...]
output_tokens: tuple[int, ...]
output_path: str
duration_seconds: float | None = None
request_count: int | None = 1
stream: bool = True
stream_include_usage: bool = True
timeout_seconds: float = 60.0
warmup_requests: int = 0
seed: int = 0
api_key: str | None = None
max_tokens_field: Literal['max_tokens', 'max_completion_tokens'] = 'max_tokens'
tokenizer: str = 'auto'
tokenizer_model: str | None = None
tiktoken_encoding: str | None = None
strict_token_counts: bool = False
system_sampler: str = 'auto'
sample_interval_seconds: float = 1.0
run_id: str | None = None
arrival_mode: str = 'closed'
rates: tuple[float, ...] = ()
burst_size: int | None = None
burst_interval_seconds: float | None = None
arrival_trace: ArrivalTrace | None = None
max_in_flight: int = 128
overflow: Literal['wait', 'drop'] = 'wait'
drain_timeout_seconds: float | None = None
prompt_mode: str = 'repeat'
shared_prefix_ratio: float | None = None
prefix_groups: int | None = None
cache_state: str = 'unspecified'
extra_body: dict[str, Any] | None = None
cache_reset_url: str | None = None
cache_reset_timeout_seconds: float = 10.0
vllm_metrics_url: str | None = None
vllm_metrics_interval_seconds: float = 1.0
vllm_spans_listen: str | None = None
vllm_spans_drain_seconds: float = 6.0
trace: TraceCaptureConfig | None = None
vllm_execution_dir: Path | None = None
slo: SloSpec | None = None
slo_source: str | None = None
prompt_spec()[source]
Return type:

PromptSpec

arrival_specs()[source]

One arrival shape per case group: per rate, or a single shape.

Return type:

list[ArrivalSpec]

cases()[source]
Return type:

list[WorkloadCase]