C
Director AI Engineering (Remote Eligible)
Capital One
2 hours ago
Full-time
On-site
Cambridge, Massachusetts, United States
Director AI Engineering (Remote Eligible)
At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent
along with our deep experience in machine learning
position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Generative AI Training team builds the platform Capital One's scientists and engineers use to train, fine-tune, and experiment with foundation models at scale. We run the GPU clusters and distributed training infrastructure behind the company's generative AI: fair-share scheduling across teams, large-scale fine-tuning and reinforcement learning, resilience for long-running jobs, and the utilization controls that keep that expensive hardware working efficiently. We also provide the self-service environments where teams deploy, serve, and evaluate models during experimentation, built on tools like KServe and vLLM. The platform we build is a central leverage point for AI across Capital One, shaping how fast and how affordable the rest of the company can build with foundation models. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Oversee the design, development, testing, deployment, and operation of the platform's core systems: distributed training and fine-tuning, reinforcement learning workflows, fair-share GPU scheduling, job resilience and fault tolerance, GPU utilization and efficiency, and self-service environments for model experimentation and evaluation. Make high judgment build-vs-buy decisions across a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art techniques to improve the scalability, cost, throughput, and reliability of large-scale distributed training and fine-tuning. Own GPU capacity planning and cost governance: right-size clusters, instance types, and quotas to the needs of training and experimentation workloads across teams. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Attract and retain top talent in the AI industry and nurture personal and professional development for your team. Foster a culture of learning and staying abreast of the state-of-the-art in AI. Translate the enterprise AI strategy into portfolio-level execution plans across multiple product areas, balancing innovation with delivery discipline Scale AI engineering practices across teams through shared infrastructure, reusable components, and unified observability and governance frameworks Establish enterprise standard for Responsible AI, including fairness metrics, model evaluation protocols, documentation requirements, and audit readiness Partner with research, compliance, and enterprise risk teams to ensure deployed systems meet emerging ethical and regulatory standards The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production You get fulfillment from empowering others to achieve their potential and you actively drive professional development through mentoring and coaching. You are hands-on when necessary and lead by example You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enables you to see and exploit optimization opportunities that others miss You are a resilient trailblazer who can forge new paths to achieve business goals when the route is unknown Capital One is open to hiring a Remote Employee for this opportunity. Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 6 years of experience developing AI and ML algorithms or technologies At least 3 years of people leadership experience Preferred Qualifications: 5+ years of experience managing and leading an engineering team 7+ years of experience building and operating large-scale ML or GPU training infrastructure on cloud platforms (e.g., AWS, Google Cloud, Azure, or equivalent private cloud) Hands-on experience with distributed training at scale: multi-node, multi-GPU jobs and parallelism strategies (e.g., PyTorch FSDP, DeepSpeed, Megatron), with proficiency in Python, Go, C++, or CUDA Experience with the ML orchestration and scheduling stack (e.g., Kubernetes, Kubeflow, Kueue, Slurm, Ray, KServe, vLLM) Experience operating large GPU fleets with a focus on reliability, fault tolerance, utilization, and cost efficiency Experience right-sizing GPU clusters, instance types, interconnect, and quotas to training and experimentation workload requirements (e.g., model size, parallelism strategy, throughput targets) Passion for staying abreast of the latest AI and ML-systems research, and judiciously applying novel training and optimization techniques Excellent communication and presentation skills, with the ability to articulate complex AI and infrastructure concepts to peers Experience building and leading a multi-team AI organization delivering multiple enterprise capabilities concurrently Proven ability to expand and execute long-term AI platform strategies aligned to enterprise priorities and regulatory frameworks Experience establishing cross-functional operating rhythms and review cadences (OKRs, AI governance councils, quarterly reviews)
At Capital One, we are creating responsible and reliable AI systems, changing banking for good. For years, Capital One has been an industry leader in using machine learning to create real-time, personalized customer experiences. Our investments in technology infrastructure and world-class talent
along with our deep experience in machine learning
position us to be at the forefront of enterprises leveraging AI. From informing customers about unusual charges to answering their questions in real time, our applications of AI & ML are bringing humanity and simplicity to banking. We are committed to continuing to build world-class applied science and engineering teams to deliver our industry leading capabilities with breakthrough product experiences and scalable, high-performance AI infrastructure. At Capital One, you will help bring the transformative power of emerging AI capabilities to reimagine how we serve our customers and businesses who have come to love the products and services we build. Team Description: The Generative AI Training team builds the platform Capital One's scientists and engineers use to train, fine-tune, and experiment with foundation models at scale. We run the GPU clusters and distributed training infrastructure behind the company's generative AI: fair-share scheduling across teams, large-scale fine-tuning and reinforcement learning, resilience for long-running jobs, and the utilization controls that keep that expensive hardware working efficiently. We also provide the self-service environments where teams deploy, serve, and evaluate models during experimentation, built on tools like KServe and vLLM. The platform we build is a central leverage point for AI across Capital One, shaping how fast and how affordable the rest of the company can build with foundation models. What You'll Do: Partner with a cross-functional team of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how our associates work and how our customers interact with Capital One. Oversee the design, development, testing, deployment, and operation of the platform's core systems: distributed training and fine-tuning, reinforcement learning workflows, fair-share GPU scheduling, job resilience and fault tolerance, GPU utilization and efficiency, and self-service environments for model experimentation and evaluation. Make high judgment build-vs-buy decisions across a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more. Invent and introduce state-of-the-art techniques to improve the scalability, cost, throughput, and reliability of large-scale distributed training and fine-tuning. Own GPU capacity planning and cost governance: right-size clusters, instance types, and quotas to the needs of training and experimentation workloads across teams. Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One. Attract and retain top talent in the AI industry and nurture personal and professional development for your team. Foster a culture of learning and staying abreast of the state-of-the-art in AI. Translate the enterprise AI strategy into portfolio-level execution plans across multiple product areas, balancing innovation with delivery discipline Scale AI engineering practices across teams through shared infrastructure, reusable components, and unified observability and governance frameworks Establish enterprise standard for Responsible AI, including fairness metrics, model evaluation protocols, documentation requirements, and audit readiness Partner with research, compliance, and enterprise risk teams to ensure deployed systems meet emerging ethical and regulatory standards The Ideal Candidate: You love to build systems, take pride in the quality of your work, and also share our passion to do the right thing. You want to work on problems that will help change banking for good Passion for staying abreast of the latest research, and an ability to intuitively understand scientific publications and judiciously apply novel techniques in production You get fulfillment from empowering others to achieve their potential and you actively drive professional development through mentoring and coaching. You are hands-on when necessary and lead by example You adapt quickly and thrive on bringing clarity to big, undefined problems. You love asking questions and digging deep to uncover the root of problems and can articulate your findings concisely with clarity. You have the courage to share new ideas even when they are unproven You are deeply Technical. You possess a strong foundation in engineering and mathematics, and your expertise in hardware, software, and AI enables you to see and exploit optimization opportunities that others miss You are a resilient trailblazer who can forge new paths to achieve business goals when the route is unknown Capital One is open to hiring a Remote Employee for this opportunity. Basic Qualifications: Bachelor's Degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 8 years of experience developing AI and ML algorithms or technologies, or a Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 6 years of experience developing AI and ML algorithms or technologies At least 3 years of people leadership experience Preferred Qualifications: 5+ years of experience managing and leading an engineering team 7+ years of experience building and operating large-scale ML or GPU training infrastructure on cloud platforms (e.g., AWS, Google Cloud, Azure, or equivalent private cloud) Hands-on experience with distributed training at scale: multi-node, multi-GPU jobs and parallelism strategies (e.g., PyTorch FSDP, DeepSpeed, Megatron), with proficiency in Python, Go, C++, or CUDA Experience with the ML orchestration and scheduling stack (e.g., Kubernetes, Kubeflow, Kueue, Slurm, Ray, KServe, vLLM) Experience operating large GPU fleets with a focus on reliability, fault tolerance, utilization, and cost efficiency Experience right-sizing GPU clusters, instance types, interconnect, and quotas to training and experimentation workload requirements (e.g., model size, parallelism strategy, throughput targets) Passion for staying abreast of the latest AI and ML-systems research, and judiciously applying novel training and optimization techniques Excellent communication and presentation skills, with the ability to articulate complex AI and infrastructure concepts to peers Experience building and leading a multi-team AI organization delivering multiple enterprise capabilities concurrently Proven ability to expand and execute long-term AI platform strategies aligned to enterprise priorities and regulatory frameworks Experience establishing cross-functional operating rhythms and review cadences (OKRs, AI governance councils, quarterly reviews)