stormlog.infer.prompts

Deterministic prompts with controlled prefix sharing.

Serving engines cache the key/value state of prompt prefixes they have seen. Whether two requests share a prefix therefore changes how much work the second one needs, so a benchmark has to choose it on purpose:

  • repeat sends one prompt for every request of a case, warmup included, as Stormlog always has. After the first request, most of each prompt can come from the cache.

  • unique starts every request with its own nonce, so no two requests share a prefix beyond whatever the server’s chat template adds.

  • shared-prefix gives each request one of N seeded group prefixes that covers a set share of its tokens, then a request nonce and filler.

Nonces depend on the seed, the case and the phase, so a run can be repeated exactly, and neither another case nor the warmup shares a prefix with the measured requests.

Classes

Prompt(text, prompt_id, counter[, planned, ...])

One request's prompt and what it shares with others.

PromptSource(spec, *, counter, seed, ...[, ...])

The prompts of one case phase, generated on demand and cached.

PromptSpec([mode, shared_prefix_ratio, ...])

How the prompts of a run share their prefixes.

class stormlog.infer.prompts.PromptSpec(mode='repeat', shared_prefix_ratio=None, prefix_groups=None)[source]

Bases: object

How the prompts of a run share their prefixes.

Parameters:
  • mode (str)

  • shared_prefix_ratio (float | None)

  • prefix_groups (int | None)

mode: str = 'repeat'
shared_prefix_ratio: float | None = None
prefix_groups: int | None = None
to_record()[source]
Return type:

dict[str, Any]

property groups: int
class stormlog.infer.prompts.Prompt(text, prompt_id, counter, planned=None, prefix_group=None, shared_prefix_tokens=None)[source]

Bases: object

One request’s prompt and what it shares with others.

planned is the size the generator aimed for, known without tokenizing the prompt. The exact count is only worked out when something reads it: most servers report the prompt’s tokens themselves.

Parameters:
  • text (str)

  • prompt_id (str)

  • counter (TokenCounter)

  • planned (TokenCount | None)

  • prefix_group (int | None)

  • shared_prefix_tokens (int | None)

text: str
prompt_id: str
counter: TokenCounter
planned: TokenCount | None = None
prefix_group: int | None = None
shared_prefix_tokens: int | None = None
property count: TokenCount
property planned_count: TokenCount
property digest: str
class stormlog.infer.prompts.PromptSource(spec, *, counter, seed, case_id, phase, input_tokens, repeated=None)[source]

Bases: object

The prompts of one case phase, generated on demand and cached.

Parameters:
  • spec (PromptSpec)

  • counter (TokenCounter)

  • seed (int)

  • case_id (str)

  • phase (str)

  • input_tokens (int)

  • repeated (dict[tuple[int, int], Prompt] | None)

prompt(index)[source]

Build, or return the built, prompt for one request.

Parameters:

index (int)

Return type:

Prompt

take(index)[source]

The prompt a request is about to use; forget it once it is done.

Parameters:

index (int)

Return type:

Prompt

forget(index)[source]

Drop a used prompt’s text; its digest stays for the phase digest.

Parameters:

index (int)

Return type:

None

warm(sample=64)[source]

Do a phase’s one-off prompt work before its clock starts.

That is the repeated prompt, every group prefix, and the filler for each length the first sample nonces leave room for. Each later prompt then only tokenizes its short nonce. The sample prompts are not handed out, so they are not part of the phase digest.

Parameters:

sample (int)

Return type:

None

prepare(indices)[source]

Build prompts ahead of use.

Parameters:

indices (Iterable[int])

Return type:

None

digest()[source]

One digest over every prompt handed out, in index order.

Return type:

str | None