stealthstack.ai
Back to results

gke-ai-troubleshooting-tpu-metrics-monitoring

Skill
outshift.io · via agntcy registry Unverified — relayed by outshift.io seen 5h ago

About

Monitors and troubleshoots GKE TPU workloads, nodes, and node pools using GKE system metrics and PromQL. Use when monitoring TensorCore duty cycle, TPU memory, node readiness, multi-host TPU node pool availability, host maintenance or preemption interruptions, and calculating MTTR or MTBI metrics for GKE TPUs. Don't use for general non-TPU GKE workload monitoring or non-metric TPU debugging.

Capabilities

The crawler did not record capability metadata for this resource. Inspect the endpoint directly to see what it exposes.

Provenance

Discovered Relayed by agntcy
URN authority urn:air:outshift.io:agntcy:gke-ai-troubleshooting-tpu-metrics-monitoring
Catalog host outshift.io
Anchor check Not anchored
Last crawled seen 5h ago

Tags

environmental monitoringgcpsystem administrationmulti agent coordinationsystem prompt supportllm observabilityproduction llm monitoringtrace analysiscontainer debuggingmonitoring alertingperformance monitoringroot cause debugging