stormlog.infer.prompts
Deterministic prompts with controlled prefix sharing.
Serving engines cache the key/value state of prompt prefixes they have seen. Whether two requests share a prefix therefore changes how much work the second one needs, so a benchmark has to choose it on purpose:
repeatsends one prompt for every request of a case, warmup included, as Stormlog always has. After the first request, most of each prompt can come from the cache.uniquestarts every request with its own nonce, so no two requests share a prefix beyond whatever the server’s chat template adds.shared-prefixgives each request one of N seeded group prefixes that covers a set share of its tokens, then a request nonce and filler.
Nonces depend on the seed, the case and the phase, so a run can be repeated exactly, and neither another case nor the warmup shares a prefix with the measured requests.
Classes
|
One request's prompt and what it shares with others. |
|
The prompts of one case phase, generated on demand and cached. |
|
How the prompts of a run share their prefixes. |
- class stormlog.infer.prompts.PromptSpec(mode='repeat', shared_prefix_ratio=None, prefix_groups=None)[source]
Bases:
objectHow the prompts of a run share their prefixes.
- Parameters:
mode (str)
shared_prefix_ratio (float | None)
prefix_groups (int | None)
- mode: str = 'repeat'
- prefix_groups: int | None = None
- property groups: int
- class stormlog.infer.prompts.Prompt(text, prompt_id, counter, planned=None, prefix_group=None, shared_prefix_tokens=None)[source]
Bases:
objectOne request’s prompt and what it shares with others.
plannedis the size the generator aimed for, known without tokenizing the prompt. The exactcountis only worked out when something reads it: most servers report the prompt’s tokens themselves.- Parameters:
text (str)
prompt_id (str)
counter (TokenCounter)
planned (TokenCount | None)
prefix_group (int | None)
shared_prefix_tokens (int | None)
- text: str
- prompt_id: str
- counter: TokenCounter
- planned: TokenCount | None = None
- prefix_group: int | None = None
- property count: TokenCount
- property planned_count: TokenCount
- property digest: str
- class stormlog.infer.prompts.PromptSource(spec, *, counter, seed, case_id, phase, input_tokens, repeated=None)[source]
Bases:
objectThe prompts of one case phase, generated on demand and cached.
- Parameters:
spec (PromptSpec)
counter (TokenCounter)
seed (int)
case_id (str)
phase (str)
input_tokens (int)
repeated (dict[tuple[int, int], Prompt] | None)
- prompt(index)[source]
Build, or return the built, prompt for one request.
- Parameters:
index (int)
- Return type:
- take(index)[source]
The prompt a request is about to use; forget it once it is done.
- Parameters:
index (int)
- Return type:
- forget(index)[source]
Drop a used prompt’s text; its digest stays for the phase digest.
- Parameters:
index (int)
- Return type:
None
- warm(sample=64)[source]
Do a phase’s one-off prompt work before its clock starts.
That is the repeated prompt, every group prefix, and the filler for each length the first
samplenonces leave room for. Each later prompt then only tokenizes its short nonce. The sample prompts are not handed out, so they are not part of the phase digest.- Parameters:
sample (int)
- Return type:
None