stormlog.infer.server_collector
Optional on-host process and NVML memory collection for inference runs.
Functions
|
Sample a live server process until the duration, a stop, or a change. |
|
Explain how the server PID relates to the processes NVML sees on a GPU. |
|
Keep a steady cadence, but skip missed polls instead of bursting. |
|
Return |
Classes
|
How many polls a collection wrote and why it stopped. |
|
|
|
|
|
Read NVML v2 memory counters from a verified GPU or MIG handle. |
Exceptions
The NVML library cannot be loaded on this host. |
- class stormlog.infer.server_collector.GpuMemorySource(*args, **kwargs)[source]
Bases:
Protocol- device_uuid: str
- gpu_instance_id: str | None
- class stormlog.infer.server_collector.GpuMemoryReading(used_bytes: 'int | None', reserved_bytes: 'int | None', state: 'str', detail: 'str | None' = None)[source]
Bases:
object- Parameters:
used_bytes (int | None)
reserved_bytes (int | None)
state (str)
detail (str | None)
- used_bytes: int | None
- reserved_bytes: int | None
- state: str
- detail: str | None = None
- class stormlog.infer.server_collector.CollectionResult(polls, stop_reason, detail=None, warnings=())[source]
Bases:
objectHow many polls a collection wrote and why it stopped.
- Parameters:
polls (int)
stop_reason (str)
detail (str | None)
warnings (tuple[str, ...])
- polls: int
- stop_reason: str
- detail: str | None = None
- warnings: tuple[str, ...] = ()
Bases:
RuntimeErrorThe NVML library cannot be loaded on this host.
- class stormlog.infer.server_collector.NvmlMemorySource(*, device_index=0, expected_uuid=None)[source]
Bases:
objectRead NVML v2 memory counters from a verified GPU or MIG handle.
- Parameters:
device_index (int)
expected_uuid (str | None)
- stormlog.infer.server_collector.collect_server_telemetry(*, run_id, pid, output_path, interval_seconds=0.1, duration_seconds=None, device_index=0, device_uuid=None, no_gpu=False, replica_id=None, rank=None, group_id=None, world_size=None, gpu_source=None, stop_event=None, on_warning=None)[source]
Sample a live server process until the duration, a stop, or a change.
Collection stops when
stop_eventis set or the caller is interrupted, when the duration elapses, when the process ends, or when the GPU identity changes. The result says which, so callers can tell a clean stop from one that leaves later case windows unobserved.group_id,rankandworld_sizedeclare this process as one member of a server group, such as one tensor-parallel worker.- Parameters:
run_id (str)
pid (int)
output_path (str | Path)
interval_seconds (float)
duration_seconds (float | None)
device_index (int)
device_uuid (str | None)
no_gpu (bool)
replica_id (str | None)
rank (int | None)
group_id (str | None)
world_size (int | None)
gpu_source (GpuMemorySource | None)
stop_event (Event | None)
on_warning (Callable[[str], None] | None)
- Return type:
- stormlog.infer.server_collector.describe_gpu_process_match(pid, device_uuid, gpu_pids, descendant_pids)[source]
Explain how the server PID relates to the processes NVML sees on a GPU.
- Parameters:
pid (int)
device_uuid (str)
gpu_pids (set[int] | None)
descendant_pids (set[int])
- Return type:
list[str]
- stormlog.infer.server_collector.next_poll_time(previous, interval, now)[source]
Keep a steady cadence, but skip missed polls instead of bursting.
- Parameters:
previous (float)
interval (float)
now (float)
- Return type:
float
- stormlog.infer.server_collector.read_process_rss(process)[source]
Return
(state, rss, detail); only a gone or replaced process is invalid.is_runningcompares the process creation time with the original, so it also detects a PID that the OS reused for another process.- Parameters:
process (psutil.Process)
- Return type:
tuple[str, int | None, str | None]