stormlog.infer.openai_client

OpenAI-compatible Chat Completions client used by inference profiling.

Functions

ignored_proxies()

What the environment asked of proxies, which the opener ignores.

inference_opener()

An opener that marks where a request failed and goes straight there.

validate_extra_body(extra_body, max_tokens_field)

Extra request fields may add settings but not replace the ones Stormlog sets.

Classes

ChatCompletionResult(text, started_at_ns, ...)

Parsed response and client-observed timing metadata.

OpenAIChatCompletionsClient(*, endpoint, ...)

Minimal OpenAI-compatible Chat Completions HTTP client.

Exceptions

ConnectError(cause)

connect() failed, so no byte of the request was sent.

CutResponseError(status, cause)

The server answered with a status line, then the connection failed.

EndpointHTTPError(status, message)

The endpoint answered with an HTTP error status.

NoResponseError(cause)

The connection failed after the request was sent, before any response.

exception stormlog.infer.openai_client.EndpointHTTPError(status, message)[source]

Bases: RuntimeError

The endpoint answered with an HTTP error status.

Parameters:
  • status (int)

  • message (str)

Return type:

None

exception stormlog.infer.openai_client.ConnectError(cause)[source]

Bases: OSError

connect() failed, so no byte of the request was sent.

That covers a refused or timed-out connection, a TLS handshake that did not finish, and a socket the peer reset before connect() returned. urllib wraps this in URLError, as it does every error from sending, and the reason tells the two apart: a request that failed here never reached the server.

Parameters:

cause (BaseException)

Return type:

None

exception stormlog.infer.openai_client.NoResponseError(cause)[source]

Bases: ConnectionError

The connection failed after the request was sent, before any response.

A small request fits in the socket buffer, so it counts as sent before the server has read a byte of it. A reset or close while waiting for the status line leaves delivery as unknown as a failure while sending.

Parameters:

cause (BaseException)

Return type:

None

exception stormlog.infer.openai_client.CutResponseError(status, cause)[source]

Bases: ConnectionError

The server answered with a status line, then the connection failed.

The server took the request: delivery is known, and the status is the answer it gave, cut short.

Parameters:
  • status (int)

  • cause (BaseException)

Return type:

None

stormlog.infer.openai_client.inference_opener()[source]

An opener that marks where a request failed and goes straight there.

It follows no redirects and ignores proxies from the environment. Through a proxy, connect() reaches the proxy, so a server that cannot be reached reads as the proxy’s HTTP 502 rather than as unreachable.

Return type:

OpenerDirector

stormlog.infer.openai_client.ignored_proxies()[source]

What the environment asked of proxies, which the opener ignores.

Only the schemes: a proxy’s URL can carry its credentials.

Return type:

dict[str, Any]

class stormlog.infer.openai_client.ChatCompletionResult(text, started_at_ns, ended_at_ns, e2e_latency_ms, ttft_ms, first_chunk_latency_ms, chunk_interarrival_ms=<factory>, usage=None, finish_reason=None)[source]

Bases: object

Parsed response and client-observed timing metadata.

Parameters:
  • text (str)

  • started_at_ns (int)

  • ended_at_ns (int)

  • e2e_latency_ms (float)

  • ttft_ms (float | None)

  • first_chunk_latency_ms (float | None)

  • chunk_interarrival_ms (list[float])

  • usage (dict[str, Any] | None)

  • finish_reason (str | None)

text: str
started_at_ns: int
ended_at_ns: int
e2e_latency_ms: float
ttft_ms: float | None
first_chunk_latency_ms: float | None
chunk_interarrival_ms: list[float]
usage: dict[str, Any] | None = None
finish_reason: str | None = None
class stormlog.infer.openai_client.OpenAIChatCompletionsClient(*, endpoint, model, timeout_seconds, api_key=None, max_tokens_field='max_tokens', extra_body=None)[source]

Bases: object

Minimal OpenAI-compatible Chat Completions HTTP client.

Parameters:
  • endpoint (str)

  • model (str)

  • timeout_seconds (float)

  • api_key (str | None)

  • max_tokens_field (Literal['max_tokens', 'max_completion_tokens'])

  • extra_body (dict[str, Any] | None)

complete(*, prompt, output_tokens, stream, stream_include_usage, request_id=None)[source]

Send one chat completion.

request_id goes out as X-Request-Id, which vLLM embeds in its own request id and in the gen_ai.request.id of the request span.

Parameters:
  • prompt (str)

  • output_tokens (int)

  • stream (bool)

  • stream_include_usage (bool)

  • request_id (str | None)

Return type:

ChatCompletionResult

stormlog.infer.openai_client.validate_extra_body(extra_body, max_tokens_field)[source]

Extra request fields may add settings but not replace the ones Stormlog sets.

Parameters:
  • extra_body (dict[str, Any] | None)

  • max_tokens_field (str)

Return type:

dict[str, Any]