gke-ai-troubleshooting-tpu-metrics-monitoring
SkillAbout
Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory, node readiness, multi-host TPU node pool availability, host maintenance or preemption interruptions, and calculating MTTR or MTBI metrics for GKE TPUs. Don't use for general non-TPU GKE workload monitoring or non-metric TPU debugging.
Capabilities
The crawler did not record capability metadata for this resource. Inspect the endpoint
directly to see what it exposes.
Provenance
Discovered Relayed by agntcy
URN authority urn:air:outshift.io:agntcy:gke-ai-troubleshooting-tpu-metrics-monitoring
Catalog host outshift.io
Anchor check Not anchored
Last crawled seen 5h ago