Skip to content

Monitor Devices, Nodes, and Workloads ​

Target Outcome ​

All four NPU cards have traceable health and utilization, and each abnormal signal can be linked to a node and workload.

Applicable Roles ​

  • Platform Operator
  • End User reviewing authorized workload monitoring

Before You Start ​

  • Record the cluster, node, device, and workload names used during onboarding and deployment.
  • Use one common time range for device, node, and workload comparisons.

Entry ​

  • Role: Operator
  • Menu: AI Infrastructure > On-Prem > Monitoring > Device / Node / Workload Monitoring
  • Routes: /powerone/monitor/device, /powerone/monitor/node, /powerone/monitor/work

Steps ​

  1. In Device Monitoring, verify that all four NPU cards are visible and compare utilization, memory, temperature, and health for each card.

Check individual NPU devices

  1. In Node Monitoring, verify that accelerator nodes are Ready and not resource constrained.

Check accelerator nodes

  1. In Workload Monitoring, locate the deployment or training job using each card.

Locate workloads using the cards

  1. Correlate an abnormal device with its node and workload before choosing a hardware, driver, quota, or application fix.

Use the same region, availability-zone, object, and time-range scope on all three pages. If the pages show different totals, clear filters and restore them one at a time before treating the difference as a fault.

Four-NPU Inspection Table ​

CheckExpected Result
Device countAll four cards are visible
Health stateNo offline card, missing card, or persistent alert
Device usageMatches the card count requested by running workloads
Node stateReady, with metrics updating continuously
Workload stateNo abnormal queueing or repeated failure

Completion Checklist ​

Purpose: These are the exit criteria for the current feature task. Use them to decide whether the result is observable and reviewable and whether you can continue to the next step in the scenario. They do not repeat the procedure; if any item fails, follow the troubleshooting section below.

CheckPass Criteria
1Every device maps to a node and occupying workload.
2Requested card count, device usage, and tenant quota agree.
3A single unhealthy card can be isolated without treating the entire cluster as unavailable.

Troubleshooting ​

SymptomCheck First
Device metrics are emptyMonitoring agent, device plug-in, time range, cluster state, and device mapping
A card is idle while jobs waitRequested specification, scheduler events, node labels, quota, and card health
Device, node, and workload totals differRegion/zone scope, time range, update time, aggregation level, and collection delay

User Manual ​