stormlog.infer.workload

The workload record: what traffic a run sent, so it can be repeated.

The digest covers what decides the requests a run sends: the cases, arrivals, prompts, warmup, decoding settings, seed, tokenizer and requested cache state. It leaves out the endpoint, model, timeouts, reset URL and where a replay trace was read from, so the same workload sent to two engine configurations has the same digest. The API key is never recorded, and the reset URL is recorded without its credentials or query string.

Functions

chat_template_identity(counter)

The server applies the chat template; Stormlog cannot see which one.

tokenizer_identity(counter)

Which tokenizer sized the prompts and counted tokens without usage.

workload_record(config, *, session_id, ...)

The infer.workload record written at the start of a run.

workload_spec(config, *, counter, prompt_spec)

stormlog.infer.workload.workload_record(config, *, session_id, counter, prompt_spec)[source]

The infer.workload record written at the start of a run.

Parameters:
Return type:

dict[str, Any]

stormlog.infer.workload.workload_spec(config, *, counter, prompt_spec)[source]
Parameters:
Return type:

dict[str, Any]

stormlog.infer.workload.tokenizer_identity(counter)[source]

Which tokenizer sized the prompts and counted tokens without usage.

Parameters:

counter (TokenCounter)

Return type:

dict[str, Any]

stormlog.infer.workload.chat_template_identity(counter)[source]

The server applies the chat template; Stormlog cannot see which one.

When a local transformers tokenizer is available, the digest of its template is recorded as a hint, not as what the server used.

Parameters:

counter (TokenCounter)

Return type:

dict[str, Any]