stormlog.infer.server_collector

Optional on-host process and NVML memory collection for inference runs.

Functions

collect_server_telemetry(*, run_id, pid, ...)

Sample a live server process until the duration, a stop, or a change.

describe_gpu_process_match(pid, device_uuid, ...)

Explain how the server PID relates to the processes NVML sees on a GPU.

next_poll_time(previous, interval, now)

Keep a steady cadence, but skip missed polls instead of bursting.

read_process_rss(process)

Return (state, rss, detail); only a gone or replaced process is invalid.

Classes

CollectionResult(polls, stop_reason[, ...])

How many polls a collection wrote and why it stopped.

GpuMemoryReading(used_bytes, reserved_bytes, ...)

GpuMemorySource(*args, **kwargs)

NvmlMemorySource(*[, device_index, ...])

Read NVML v2 memory counters from a verified GPU or MIG handle.

Exceptions

NvmlUnavailableError

The NVML library cannot be loaded on this host.

class stormlog.infer.server_collector.GpuMemorySource(*args, **kwargs)[source]

Bases: Protocol

device_uuid: str
gpu_instance_id: str | None
read()[source]

Return memory values and their availability state.

Return type:

GpuMemoryReading

close()[source]
Return type:

None

class stormlog.infer.server_collector.GpuMemoryReading(used_bytes: 'int | None', reserved_bytes: 'int | None', state: 'str', detail: 'str | None' = None)[source]

Bases: object

Parameters:
  • used_bytes (int | None)

  • reserved_bytes (int | None)

  • state (str)

  • detail (str | None)

used_bytes: int | None
reserved_bytes: int | None
state: str
detail: str | None = None
class stormlog.infer.server_collector.CollectionResult(polls, stop_reason, detail=None, warnings=())[source]

Bases: object

How many polls a collection wrote and why it stopped.

Parameters:
  • polls (int)

  • stop_reason (str)

  • detail (str | None)

  • warnings (tuple[str, ...])

polls: int
stop_reason: str
detail: str | None = None
warnings: tuple[str, ...] = ()
exception stormlog.infer.server_collector.NvmlUnavailableError[source]

Bases: RuntimeError

The NVML library cannot be loaded on this host.

class stormlog.infer.server_collector.NvmlMemorySource(*, device_index=0, expected_uuid=None)[source]

Bases: object

Read NVML v2 memory counters from a verified GPU or MIG handle.

Parameters:
  • device_index (int)

  • expected_uuid (str | None)

read()[source]
Return type:

GpuMemoryReading

compute_pids()[source]

Return PIDs NVML reports on this GPU, or None when NVML cannot say.

Return type:

set[int] | None

close()[source]
Return type:

None

stormlog.infer.server_collector.collect_server_telemetry(*, run_id, pid, output_path, interval_seconds=0.1, duration_seconds=None, device_index=0, device_uuid=None, no_gpu=False, replica_id=None, rank=None, group_id=None, world_size=None, gpu_source=None, stop_event=None, on_warning=None)[source]

Sample a live server process until the duration, a stop, or a change.

Collection stops when stop_event is set or the caller is interrupted, when the duration elapses, when the process ends, or when the GPU identity changes. The result says which, so callers can tell a clean stop from one that leaves later case windows unobserved. group_id, rank and world_size declare this process as one member of a server group, such as one tensor-parallel worker.

Parameters:
  • run_id (str)

  • pid (int)

  • output_path (str | Path)

  • interval_seconds (float)

  • duration_seconds (float | None)

  • device_index (int)

  • device_uuid (str | None)

  • no_gpu (bool)

  • replica_id (str | None)

  • rank (int | None)

  • group_id (str | None)

  • world_size (int | None)

  • gpu_source (GpuMemorySource | None)

  • stop_event (Event | None)

  • on_warning (Callable[[str], None] | None)

Return type:

CollectionResult

stormlog.infer.server_collector.describe_gpu_process_match(pid, device_uuid, gpu_pids, descendant_pids)[source]

Explain how the server PID relates to the processes NVML sees on a GPU.

Parameters:
  • pid (int)

  • device_uuid (str)

  • gpu_pids (set[int] | None)

  • descendant_pids (set[int])

Return type:

list[str]

stormlog.infer.server_collector.next_poll_time(previous, interval, now)[source]

Keep a steady cadence, but skip missed polls instead of bursting.

Parameters:
  • previous (float)

  • interval (float)

  • now (float)

Return type:

float

stormlog.infer.server_collector.read_process_rss(process)[source]

Return (state, rss, detail); only a gone or replaced process is invalid.

is_running compares the process creation time with the original, so it also detects a PID that the OS reused for another process.

Parameters:

process (psutil.Process)

Return type:

tuple[str, int | None, str | None]