stormlog.infer.openai_client
OpenAI-compatible Chat Completions client used by inference profiling.
Functions
What the environment asked of proxies, which the opener ignores. |
|
An opener that marks where a request failed and goes straight there. |
|
|
Extra request fields may add settings but not replace the ones Stormlog sets. |
Classes
|
Parsed response and client-observed timing metadata. |
|
Minimal OpenAI-compatible Chat Completions HTTP client. |
Exceptions
|
|
|
The server answered with a status line, then the connection failed. |
|
The endpoint answered with an HTTP error status. |
|
The connection failed after the request was sent, before any response. |
- exception stormlog.infer.openai_client.EndpointHTTPError(status, message)[source]
Bases:
RuntimeErrorThe endpoint answered with an HTTP error status.
- Parameters:
status (int)
message (str)
- Return type:
None
- exception stormlog.infer.openai_client.ConnectError(cause)[source]
Bases:
OSErrorconnect()failed, so no byte of the request was sent.That covers a refused or timed-out connection, a TLS handshake that did not finish, and a socket the peer reset before
connect()returned. urllib wraps this inURLError, as it does every error from sending, and the reason tells the two apart: a request that failed here never reached the server.- Parameters:
cause (BaseException)
- Return type:
None
- exception stormlog.infer.openai_client.NoResponseError(cause)[source]
Bases:
ConnectionErrorThe connection failed after the request was sent, before any response.
A small request fits in the socket buffer, so it counts as sent before the server has read a byte of it. A reset or close while waiting for the status line leaves delivery as unknown as a failure while sending.
- Parameters:
cause (BaseException)
- Return type:
None
- exception stormlog.infer.openai_client.CutResponseError(status, cause)[source]
Bases:
ConnectionErrorThe server answered with a status line, then the connection failed.
The server took the request: delivery is known, and the status is the answer it gave, cut short.
- Parameters:
status (int)
cause (BaseException)
- Return type:
None
- stormlog.infer.openai_client.inference_opener()[source]
An opener that marks where a request failed and goes straight there.
It follows no redirects and ignores proxies from the environment. Through a proxy,
connect()reaches the proxy, so a server that cannot be reached reads as the proxy’s HTTP 502 rather than asunreachable.- Return type:
OpenerDirector
- stormlog.infer.openai_client.ignored_proxies()[source]
What the environment asked of proxies, which the opener ignores.
Only the schemes: a proxy’s URL can carry its credentials.
- Return type:
dict[str, Any]
- class stormlog.infer.openai_client.ChatCompletionResult(text, started_at_ns, ended_at_ns, e2e_latency_ms, ttft_ms, first_chunk_latency_ms, chunk_interarrival_ms=<factory>, usage=None, finish_reason=None)[source]
Bases:
objectParsed response and client-observed timing metadata.
- Parameters:
text (str)
started_at_ns (int)
ended_at_ns (int)
e2e_latency_ms (float)
ttft_ms (float | None)
first_chunk_latency_ms (float | None)
chunk_interarrival_ms (list[float])
usage (dict[str, Any] | None)
finish_reason (str | None)
- text: str
- started_at_ns: int
- ended_at_ns: int
- e2e_latency_ms: float
- ttft_ms: float | None
- first_chunk_latency_ms: float | None
- chunk_interarrival_ms: list[float]
- usage: dict[str, Any] | None = None
- finish_reason: str | None = None
- class stormlog.infer.openai_client.OpenAIChatCompletionsClient(*, endpoint, model, timeout_seconds, api_key=None, max_tokens_field='max_tokens', extra_body=None)[source]
Bases:
objectMinimal OpenAI-compatible Chat Completions HTTP client.
- Parameters:
endpoint (str)
model (str)
timeout_seconds (float)
api_key (str | None)
max_tokens_field (Literal['max_tokens', 'max_completion_tokens'])
extra_body (dict[str, Any] | None)
- complete(*, prompt, output_tokens, stream, stream_include_usage, request_id=None)[source]
Send one chat completion.
request_idgoes out asX-Request-Id, which vLLM embeds in its own request id and in thegen_ai.request.idof the request span.- Parameters:
prompt (str)
output_tokens (int)
stream (bool)
stream_include_usage (bool)
request_id (str | None)
- Return type: