{"repo":"aws-samples/aws-parallelcluster-monitoring","free":true,"listed":false,"github":"https://github.com/aws-samples/aws-parallelcluster-monitoring","clone":"git clone https://github.com/aws-samples/aws-parallelcluster-monitoring.git","description":"Monitoring Dashboard for AWS ParallelCluster AWS ParallelCluster & AWS PCS + Amazon RES virtual desktops","language":"Shell","stars":41,"topics":["aws-parallelcluster","grafana-dashboard","metrics","slurm","monitoring","hpc","gpu","aws","parallelcluster","res"],"license":"MIT-0","category":"analytics","readme_excerpt":"HPC Cluster Monitoring Dashboard for AWS ParallelCluster & AWS PCS + Amazon RES virtual desktops A zero-setup monitoring solution for HPC clusters built with AWS ParallelCluster or AWS Parallel Computing Service (PCS). Deploys Prometheus, Grafana, node exporter, NVIDIA DCGM exporter, and Slurm metrics as containers — no manual configuration required. It can also monitor Amazon Research and Engineering Studio (RES) virtual desktops alongside the cluster for rightsizing. Features - Dual-platform : works on both ParallelCluster and AWS PCS - Zero-setup : add a few lines to your config, create the cluster, done - Slurm 25.11 native metrics (PCS): scrapes OpenMetrics directly from slurmctld — no extra exporter - Per-user / per-account / per-partition visibility : see who's using what, queue health, scheduler RPC stats - GPU profiling : SM activity, tensor-core utilization, FP64/32/16 pipe activity (Volta+) - GPU health monitoring : XID errors, throttle reasons, ECC counters, NVLink errors, retired pages - Secure by default : per-cluster random password in SSM, optional Cognito SSO - GPU-ready : NVIDIA DCGM exporter with custom counters auto-deploys on GPU instances - EFA fabric metrics : bandwidth, packet rate, RDMA read/write throughput, SRD retransmits and work-request errors — collected automatically on EFA hardware - Amazon RES desktop monitoring : monitor Research and Engineering Studio VDI desktops (CPU, RAM, GPU) for rightsizing, on the same stack — opt-in, discovered by re","default_branch":null,"files":null,"tree":[],"storefront":"/r/aws-samples","claimed":false,"request_supported":{"post":"https://gitbuyer.com/r/aws-samples/aws-parallelcluster-monitoring/request-supported","requests":0},"note":"indexed from public GitHub; nothing is for sale on this page. Clone it from GitHub. Paid listings live at /search."}