Working remotely, the full-time NVIDIA GPU AI Engineer will build and operate the GPU and inference serving stack NVIDIA AI Enterprise on EKS NIM microservices while managing GPU sharing scheduling integration and serving performance.
Key responsibilities
Deploy and operate NVIDIA AI Enterprise on EKS, including GPU Operator drivers and CUDA runtime
Configure GPU sharing and integrate RunAI for GPU scheduling and autoscaling
Tune GPU serving performance and troubleshoot CUDA driver and scheduling issues
Required qualifications
8 years of experience in infrastructure ML engineering with hands-on NVIDIA GPU operations
Proficiency with NVIDIA GPU stack drivers, CUDA, and DCGM
Experience with Kubernetes GPU workloads and device plugins
Familiarity with GPU scheduling tools like RunAI and GPU partitioning techniques
Knowledge of autoscaling methods and OpenAI-compatible inference API patterns