stealthstack.ai
Back to results

gke-inference

Skill
outshift.io · via agntcy registry Unverified — relayed by outshift.io seen 5h ago

About

Deploys and optimizes AI/ML inference workloads on GKE, using GPUs, TPUs, and model servers. Use when deploying GKE inference servers, configuring GKE GPU resources for inference, or deploying LLMs on GKE. Don't use for generic batch jobs or HPC task queues (use gke-batch-hpc instead).

Capabilities

The crawler did not record capability metadata for this resource. Inspect the endpoint directly to see what it exposes.

Provenance

Discovered Relayed by agntcy
URN authority urn:air:outshift.io:agntcy:gke-inference
Catalog host outshift.io
Anchor check Not anchored
Last crawled seen 5h ago

Tags

precision agriculturelarge language modelsgcpinference servingedge inferenceinference optimizationmodel servingllm capabilitiesmodel optimizationmodel quantizationmodel trainingdistributed training