gke-ai-troubleshooting-handle-disruption-gpu-tpu
SkillAbout
Diagnoses, predicts, and mitigates node disruptions during Compute Engine host maintenance and hardware or software maintenance events for GPU and TPU workloads on GKE. Use when diagnosing node disruptions, predicting host maintenance events on GPU/TPU nodepools, inspecting node interruption PromQL metrics, auditing node taints, or configuring workload protection strategies (graceful termination, opportunistic maintenance, PodDisruptionBudgets). Don't use for general GKE cluster creation, network policy configuration, or non-disruption workload deployment.
Capabilities
The crawler did not record capability metadata for this resource. Inspect the endpoint
directly to see what it exposes.
Provenance
Discovered Relayed by agntcy
URN authority urn:air:outshift.io:agntcy:gke-ai-troubleshooting-handle-disruption-gpu-tpu
Catalog host outshift.io
Anchor check Not anchored
Last crawled seen 5h ago