Skip to main content
T

Artificial Intelligence Engineer

Tiger Advisory
2 hours ago
Full-time
On-site
AI Infrastructure Engineer Location:

Remote — must be available during Eastern Time business hours Engagement:

6–12-month contract Industry:

Banking; Financial Services

If you are interested in applying for this job, please make sure you meet the following requirements as listed below.

About the Role We are seeking a hands-on

AI Infrastructure Engineer

to help stand up and operate a new enterprise GPU compute environment. The infrastructure includes an

18-node NVIDIA HGX cluster with 144 GPUs , supported by NVIDIA InfiniBand and high-speed Ethernet switching. This is a hybrid infrastructure role spanning three areas: GPU fabric and high-speed networking Broader network operations and datacenter connectivity GPU platform, orchestration, and workload enablement The ideal candidate can operate across the full stack—from physical switch ports and fabric telemetry through GPU drivers, workload schedulers, containers, and production AI workloads. This is not a pure network engineering or pure platform engineering position.

Key Responsibilities GPU Fabric and High-Speed Networking Manage and operate an

InfiniBand XDR / Quantum-3400 rail-aligned fabric topology

supporting an 18-node, 144-GPU NVIDIA HGX cluster. Administer

NVIDIA UFM and NetQ

for subnet management, telemetry, monitoring, and firmware management. Support

800G optics , troubleshoot link-flap issues, and manage congestion isolation. Run

NCCL and HPL validation and performance regression testing

to confirm cluster health and throughput following infrastructure changes. Operate

Cumulus SN5600 and Arista border switches

across the Ethernet side of the environment. Diagnose performance and connectivity issues across multi-node GPU workloads. Network Operations Manage and support large-scale IP networking technologies and infrastructure. Monitor network health across on-premises and cloud environments. Work with peering and datacenter interconnect technologies, including: Private Network Interconnects (PNIs) Transit Internet Exchanges Passive DWDM Wave circuits Improve change-management processes and day-to-day network operations. Document technical standards, operating procedures, and infrastructure best practices. Develop workflow enhancements that improve reliability and operational efficiency. GPU Platform and Workload Enablement Manage the lifecycle of

NVIDIA AI Enterprise and the NVIDIA GPU Operator . Maintain driver, firmware, CUDA, and platform compatibility across the cluster. Administer

Run , including quotas, project structures, resource allocation, and tenant onboarding. Configure and support

Slurm

across dedicated worker-node pools for scheduled batch workloads. Manage container images, runtime dependencies, and model-serving components across production and pre-production environments. Partner with the DNA MLOps team to onboard the first production workloads onto the new cluster. Support engineering and data science teams as adoption of the GPU platform expands.

Required Qualifications Hands-on experience supporting

NVIDIA InfiniBand , preferably the Quantum series, within an HPC or GPU cluster environment. Experience with high-speed Ethernet switching in a multi-node GPU environment. Familiarity with

NVIDIA HGX architecture

and rail-aligned GPU networking topologies. Experience using

NCCL, HPL, or comparable tools

to validate GPU cluster health, connectivity, and performance. Working knowledge of the

NVIDIA GPU Operator , GPU driver and firmware lifecycle management, and compatibility planning. Experience deploying and supporting container-based workloads in an HPC or AI environment. Experience with at least one workload orchestration platform, such as

Slurm, Run, or an equivalent scheduler . Strong knowledge of large-scale IP networking and datacenter peering/interconnect technologies. Experience working within structured change-management and production operations processes. Ability to troubleshoot across physical infrastructure, network fabric, operating systems, containers, orchestration platforms, and application workloads.

Preferred Qualifications Experience administering

NVIDIA UFM, NetQ, or equivalent InfiniBand fabric-management tools . Experience with new GPU cluster bring-up or greenfield AI infrastructure deployments. Exposure to MLOps pipelines, production AI workloads, and model-serving infrastructure. Familiarity with

DWDM and wave circuit technologies

in a datacenter interconnect environment. Experience supporting enterprise GPU infrastructure in a regulated or highly controlled environment.

Ideal Candidate Profile The strongest candidate will combine deep GPU-cluster networking knowledge with practical platform engineering experience. xsgimln They should be comfortable moving between InfiniBand troubleshooting, Ethernet and datacenter operations, GPU software lifecycle management, workload scheduling, and production AI enablement.