{"repo":"uzunenes/triton-server-hpa","free":true,"listed":false,"github":"https://github.com/uzunenes/triton-server-hpa","clone":"git clone https://github.com/uzunenes/triton-server-hpa.git","description":"Horizontal Pod Autoscaler (HPA) project on Kubernetes using NVIDIA Triton Inference Server with an AI model","language":"Python","stars":16,"topics":["ai","hpa","kubernetes","triton-inference-server","cloud-ai-service","gpu","nvidia","prometheus"],"license":"MIT","category":"deployment-docker-iac","readme_excerpt":"Triton Server HPA GPU-based Horizontal Pod Autoscaling for NVIDIA Triton Inference Server In this guide, you'll learn how to build a scalable AI inference system that dynamically handles fluctuating workloads. Using tools like Docker, Kubernetes, and Nvidia Triton Inference Server, this step-by-step tutorial covers everything from installation to horizontal scaling. Table of Contents - Mastering AI Request Volumes: Scalable Solutions for High and Low Demands - 1. Create simple Vision based AI Model Application - 1.1 Installation - 1.1.1 NVIDIA Container Toolkit - 1.1.2 K8s - Minikube - 1.1.3 Kubectl - 1.1.4 Helm and GPU Operator - 1.2 Preparing the YOLOv7 AI Model - 1.3 Deploying Triton Inference Server - Deployment Configuration - Triton Service Configuration - Verify Model Deployment - 1.4 Create High GPU Usage and Check Results - 2. Manage Demands with Horizontal Pod Autoscale - 2.1 Install DCGM on Host - 2.2 Deploy DCGM Exporter - 2.3 Set Up Prometheus and Prometheus Adapter - 2.4 Configure Horizontal Pod Autoscaler (HPA) - Acknowledgements - References --- 1. Create simple Vision based AI Model Application 1.1 Installation 1.1.1 NVIDIA Container Toolkit Install the NVIDIA Container Toolkit to enable GPU support for Docker containers. Verify Installation: Expected Output: The nvidia-smi command should display GPU details. --- 1.1.2 K8s - Minikube Install and configure Minikube with GPU support. Verify Installation: --- 1.1.3 Kubectl Install kubectl , the Kubernetes comman","default_branch":null,"files":null,"tree":[],"storefront":"/r/uzunenes","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/uzunenes/triton-server-hpa/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}