T
Artificial Intelligence Engineer
Tiger Advisory
2 hours ago
Full-time
On-site
AI Infrastructure Engineer
Location:
Remote — must be available during Eastern Time business hours Engagement:
6–12-month contract Industry:
Banking; Financial Services
If you are interested in applying for this job, please make sure you meet the following requirements as listed below.
About the Role We are seeking a hands-on
AI Infrastructure Engineer
to help stand up and operate a new enterprise GPU compute environment. The infrastructure includes an
18-node NVIDIA HGX cluster with 144 GPUs , supported by NVIDIA InfiniBand and high-speed Ethernet switching. This is a hybrid infrastructure role spanning three areas: GPU fabric and high-speed networking Broader network operations and datacenter connectivity GPU platform, orchestration, and workload enablement The ideal candidate can operate across the full stack—from physical switch ports and fabric telemetry through GPU drivers, workload schedulers, containers, and production AI workloads. This is not a pure network engineering or pure platform engineering position.
Key Responsibilities GPU Fabric and High-Speed Networking Manage and operate an
InfiniBand XDR / Quantum-3400 rail-aligned fabric topology
supporting an 18-node, 144-GPU NVIDIA HGX cluster. Administer
NVIDIA UFM and NetQ
for subnet management, telemetry, monitoring, and firmware management. Support
800G optics , troubleshoot link-flap issues, and manage congestion isolation. Run
NCCL and HPL validation and performance regression testing
to confirm cluster health and throughput following infrastructure changes. Operate
Cumulus SN5600 and Arista border switches
across the Ethernet side of the environment. Diagnose performance and connectivity issues across multi-node GPU workloads. Network Operations Manage and support large-scale IP networking technologies and infrastructure. Monitor network health across on-premises and cloud environments. Work with peering and datacenter interconnect technologies, including: Private Network Interconnects (PNIs) Transit Internet Exchanges Passive DWDM Wave circuits Improve change-management processes and day-to-day network operations. Document technical standards, operating procedures, and infrastructure best practices. Develop workflow enhancements that improve reliability and operational efficiency. GPU Platform and Workload Enablement Manage the lifecycle of
NVIDIA AI Enterprise and the NVIDIA GPU Operator . Maintain driver, firmware, CUDA, and platform compatibility across the cluster. Administer
Run , including quotas, project structures, resource allocation, and tenant onboarding. Configure and support
Slurm
across dedicated worker-node pools for scheduled batch workloads. Manage container images, runtime dependencies, and model-serving components across production and pre-production environments. Partner with the DNA MLOps team to onboard the first production workloads onto the new cluster. Support engineering and data science teams as adoption of the GPU platform expands.
Required Qualifications Hands-on experience supporting
NVIDIA InfiniBand , preferably the Quantum series, within an HPC or GPU cluster environment. Experience with high-speed Ethernet switching in a multi-node GPU environment. Familiarity with
NVIDIA HGX architecture
and rail-aligned GPU networking topologies. Experience using
NCCL, HPL, or comparable tools
to validate GPU cluster health, connectivity, and performance. Working knowledge of the
NVIDIA GPU Operator , GPU driver and firmware lifecycle management, and compatibility planning. Experience deploying and supporting container-based workloads in an HPC or AI environment. Experience with at least one workload orchestration platform, such as
Slurm, Run, or an equivalent scheduler . Strong knowledge of large-scale IP networking and datacenter peering/interconnect technologies. Experience working within structured change-management and production operations processes. Ability to troubleshoot across physical infrastructure, network fabric, operating systems, containers, orchestration platforms, and application workloads.
Preferred Qualifications Experience administering
NVIDIA UFM, NetQ, or equivalent InfiniBand fabric-management tools . Experience with new GPU cluster bring-up or greenfield AI infrastructure deployments. Exposure to MLOps pipelines, production AI workloads, and model-serving infrastructure. Familiarity with
DWDM and wave circuit technologies
in a datacenter interconnect environment. Experience supporting enterprise GPU infrastructure in a regulated or highly controlled environment.
Ideal Candidate Profile The strongest candidate will combine deep GPU-cluster networking knowledge with practical platform engineering experience. xsgimln They should be comfortable moving between InfiniBand troubleshooting, Ethernet and datacenter operations, GPU software lifecycle management, workload scheduling, and production AI enablement.
Remote — must be available during Eastern Time business hours Engagement:
6–12-month contract Industry:
Banking; Financial Services
If you are interested in applying for this job, please make sure you meet the following requirements as listed below.
About the Role We are seeking a hands-on
AI Infrastructure Engineer
to help stand up and operate a new enterprise GPU compute environment. The infrastructure includes an
18-node NVIDIA HGX cluster with 144 GPUs , supported by NVIDIA InfiniBand and high-speed Ethernet switching. This is a hybrid infrastructure role spanning three areas: GPU fabric and high-speed networking Broader network operations and datacenter connectivity GPU platform, orchestration, and workload enablement The ideal candidate can operate across the full stack—from physical switch ports and fabric telemetry through GPU drivers, workload schedulers, containers, and production AI workloads. This is not a pure network engineering or pure platform engineering position.
Key Responsibilities GPU Fabric and High-Speed Networking Manage and operate an
InfiniBand XDR / Quantum-3400 rail-aligned fabric topology
supporting an 18-node, 144-GPU NVIDIA HGX cluster. Administer
NVIDIA UFM and NetQ
for subnet management, telemetry, monitoring, and firmware management. Support
800G optics , troubleshoot link-flap issues, and manage congestion isolation. Run
NCCL and HPL validation and performance regression testing
to confirm cluster health and throughput following infrastructure changes. Operate
Cumulus SN5600 and Arista border switches
across the Ethernet side of the environment. Diagnose performance and connectivity issues across multi-node GPU workloads. Network Operations Manage and support large-scale IP networking technologies and infrastructure. Monitor network health across on-premises and cloud environments. Work with peering and datacenter interconnect technologies, including: Private Network Interconnects (PNIs) Transit Internet Exchanges Passive DWDM Wave circuits Improve change-management processes and day-to-day network operations. Document technical standards, operating procedures, and infrastructure best practices. Develop workflow enhancements that improve reliability and operational efficiency. GPU Platform and Workload Enablement Manage the lifecycle of
NVIDIA AI Enterprise and the NVIDIA GPU Operator . Maintain driver, firmware, CUDA, and platform compatibility across the cluster. Administer
Run , including quotas, project structures, resource allocation, and tenant onboarding. Configure and support
Slurm
across dedicated worker-node pools for scheduled batch workloads. Manage container images, runtime dependencies, and model-serving components across production and pre-production environments. Partner with the DNA MLOps team to onboard the first production workloads onto the new cluster. Support engineering and data science teams as adoption of the GPU platform expands.
Required Qualifications Hands-on experience supporting
NVIDIA InfiniBand , preferably the Quantum series, within an HPC or GPU cluster environment. Experience with high-speed Ethernet switching in a multi-node GPU environment. Familiarity with
NVIDIA HGX architecture
and rail-aligned GPU networking topologies. Experience using
NCCL, HPL, or comparable tools
to validate GPU cluster health, connectivity, and performance. Working knowledge of the
NVIDIA GPU Operator , GPU driver and firmware lifecycle management, and compatibility planning. Experience deploying and supporting container-based workloads in an HPC or AI environment. Experience with at least one workload orchestration platform, such as
Slurm, Run, or an equivalent scheduler . Strong knowledge of large-scale IP networking and datacenter peering/interconnect technologies. Experience working within structured change-management and production operations processes. Ability to troubleshoot across physical infrastructure, network fabric, operating systems, containers, orchestration platforms, and application workloads.
Preferred Qualifications Experience administering
NVIDIA UFM, NetQ, or equivalent InfiniBand fabric-management tools . Experience with new GPU cluster bring-up or greenfield AI infrastructure deployments. Exposure to MLOps pipelines, production AI workloads, and model-serving infrastructure. Familiarity with
DWDM and wave circuit technologies
in a datacenter interconnect environment. Experience supporting enterprise GPU infrastructure in a regulated or highly controlled environment.
Ideal Candidate Profile The strongest candidate will combine deep GPU-cluster networking knowledge with practical platform engineering experience. xsgimln They should be comfortable moving between InfiniBand troubleshooting, Ethernet and datacenter operations, GPU software lifecycle management, workload scheduling, and production AI enablement.