To contribute to groundbreaking technologies in supercomputing and AI, the full-time Senior HPC AI Engineer will design, implement, and maintain large-scale HPC/AI clusters while collaborating with researchers and developers in a remote or onsite capacity.
Key responsibilities:
Design, implement, and maintain large-scale HPC/AI clusters with monitoring and alerting systems
Manage Linux job/workload schedules and develop continuous integration and delivery pipelines
Automate deployment and management of infrastructure environments and support R&D activities
Required qualifications:
A degree in Computer Science, Engineering, or a related field with 8+ years of experience
Knowledge of HPC and AI solution technologies, including CPU, GPU, and high-speed interconnects
Experience with job scheduling tools such as Slurm and Kubernetes
Proficiency in Linux networking and internals, including security protocols
Python programming and bash scripting experience