Catalog Details
CATEGORY
deploymentCREATED BY
CREATED AT
March 07, 2024VERSION
0.0.1-
MODELS
What This Design Does:
Deploy torchserve inference server with prepared T5 model and Client Application. Manifests were tested against GKE Autopilot Kubernetes cluster.
Caveats and Consideration:
To configure HPA base on metrics from torchserve you need to: Enable Google Manager Prometheus or install OSS Prometheus. Install Custom Metrics Adapter. Apply pod-monitoring.yaml and hpa.yaml