Catalog Details
CATEGORY
deploymentUUID
fb8a31f7-0adb-40e7-ac26-bffb0a0f4547CREATED BY
CREATED AT
March 07, 2024VERSION
0.0.1-
MODELS
What This Design Does:
Deploy torchserve inference server with prepared T5 model and Client Application. Manifests were tested against GKE Autopilot Kubernetes cluster.
Caveats and Consideration:
To configure HPA base on metrics from torchserve you need to: Enable Google Manager Prometheus or install OSS Prometheus. Install Custom Metrics Adapter. Apply pod-monitoring.yaml and hpa.yaml