M
Manager, AI Engineering (Tester )
Mastercard
2 hours ago
Full-time
On-site
Missouri, United States
Manager, AI Engineering (Tester)
Mastercard's Business & Market Insights (B&MI) group delivers unparalleled data-driven intelligence and frontier AI solutions that help organizations make smarter, faster, and more impactful decisions. We are currently looking for an AI Tester for the Operational Intelligence Program within B&MI. This is a highly specialized, hands-on AI testing leadership position dedicated to ensuring our Generative AI, LLM, and agentic systems are accurate, safe, reliable, and enterprise-ready. This role will lead AI quality engineering efforts
defining evaluation frameworks, red-teaming strategies, and LLMOps quality gates
while fostering a culture of rigorous, first-class AI testing across the program. Roles and Responsibilities: Design and own end-to-end LLM evaluation frameworks
including automated prompt regression pipelines, output scoring, semantic benchmarking, and hallucination detection across model versions and prompt variations. Build comprehensive test suites for agentic AI systems
validating tool selection, inter-agent coordination, task decomposition, goal completion, and failure handling across multi-step reasoning workflows. Develop RAG pipeline evaluation frameworks assessing retrieval precision, chunk relevance, context faithfulness, answer grounding, and hallucination rates using tools like RAGAS, TruLens, and DeepEval. Lead structured red-teaming and adversarial testing exercises targeting prompt injection, jailbreaks, data leakage, context poisoning, and model manipulation
building and maintaining an evolving adversarial test library. Execute fairness, bias, and Responsible AI audits
testing for demographic bias, sentiment skew, representation gaps, and validating explainability mechanisms, citations, and confidence score accuracy. Design and run inference performance benchmarks
measuring latency, throughput, token efficiency, and degradation under peak load
and enforce LLM quality gates within CI/CD pipelines on Databricks (AWS). Build production monitoring and drift detection pipelines tracking semantic output drift, embedding shifts, retrieval degradation, and anomalous agent behaviors using observability tooling (Grafana, Datadog, CloudWatch). Define the AI testing roadmap and quality standards for the program
establishing evaluation metrics, tooling choices, and documentation practices across all Gen AI workstreams. Partner with Gen AI engineers, ML engineers, and product stakeholders to embed quality from day one
reviewing prompt architectures, agent designs, and system workflows for testability and risk. Continuously research and adopt frontier evaluation benchmarks (RAGAS, MMLU, TruthfulQA, MT-Bench) and emerging AI testing methodologies to keep quality practices at the cutting edge. All About You: Master's/Bachelor's degree in Computer Science, AI/ML, or Software Engineering, with considerable hands-on experience leading AI/ML quality engineering or LLM testing programs in production environments. Demonstrated expertise testing LLM and Gen AI systems
including prompt testing, output evaluation, hallucination detection, RAG pipeline assessment, and agentic workflow validation in real production settings. Deep hands-on knowledge of AI evaluation frameworks and tooling: RAGAS, DeepEval, TruLens, LangSmith, PromptFlow, Weights & Biases Evals, or equivalent platforms. Strong understanding of Gen AI failure modes
hallucination, prompt injection, retrieval grounding failures, context drift, agent loop failures
and proven methods to surface and document them systematically. Strong Python programming skills with the ability to independently build test automation scripts, evaluation pipelines, and API-level integration tests; SQL proficiency required. Working knowledge of LLM ecosystems
OpenAI, Anthropic, Hugging Face, LangChain/LangGraph
sufficient to understand model behavior, prompt structure, and agent architecture deeply enough to test them rigorously. Familiarity with MLOps/LLMOps pipelines (MLflow, Databricks, SageMaker) and experience integrating automated quality gates into CI/CD workflows for AI systems. Experience with cloud AI infrastructure (AWS, Azure, or GCP) and observability tooling for monitoring live AI system behavior and output quality in production. Strong analytical, communication, and stakeholder management skills
with the ability to translate complex AI failure patterns into clear risk assessments and remediation recommendations for both technical and business audiences. Mastercard is a merit-based, inclusive, equal opportunity employer that considers applicants without regard to gender, gender identity, sexual orientation, race, ethnicity, disabled or veteran status, or any other characteristic protected by law. We hire the most qualified candidate for the role. In the US or Canada, if you require accommodations or assistance to complete the online application process or during the recruitment process, please contact reasonable_accommodation@mastercard.com and identify the type of accommodation or assistance you are requesting. Do not include any medical or health information in this email. The Reasonable Accommodations team will respond to your email promptly.
Mastercard's Business & Market Insights (B&MI) group delivers unparalleled data-driven intelligence and frontier AI solutions that help organizations make smarter, faster, and more impactful decisions. We are currently looking for an AI Tester for the Operational Intelligence Program within B&MI. This is a highly specialized, hands-on AI testing leadership position dedicated to ensuring our Generative AI, LLM, and agentic systems are accurate, safe, reliable, and enterprise-ready. This role will lead AI quality engineering efforts
defining evaluation frameworks, red-teaming strategies, and LLMOps quality gates
while fostering a culture of rigorous, first-class AI testing across the program. Roles and Responsibilities: Design and own end-to-end LLM evaluation frameworks
including automated prompt regression pipelines, output scoring, semantic benchmarking, and hallucination detection across model versions and prompt variations. Build comprehensive test suites for agentic AI systems
validating tool selection, inter-agent coordination, task decomposition, goal completion, and failure handling across multi-step reasoning workflows. Develop RAG pipeline evaluation frameworks assessing retrieval precision, chunk relevance, context faithfulness, answer grounding, and hallucination rates using tools like RAGAS, TruLens, and DeepEval. Lead structured red-teaming and adversarial testing exercises targeting prompt injection, jailbreaks, data leakage, context poisoning, and model manipulation
building and maintaining an evolving adversarial test library. Execute fairness, bias, and Responsible AI audits
testing for demographic bias, sentiment skew, representation gaps, and validating explainability mechanisms, citations, and confidence score accuracy. Design and run inference performance benchmarks
measuring latency, throughput, token efficiency, and degradation under peak load
and enforce LLM quality gates within CI/CD pipelines on Databricks (AWS). Build production monitoring and drift detection pipelines tracking semantic output drift, embedding shifts, retrieval degradation, and anomalous agent behaviors using observability tooling (Grafana, Datadog, CloudWatch). Define the AI testing roadmap and quality standards for the program
establishing evaluation metrics, tooling choices, and documentation practices across all Gen AI workstreams. Partner with Gen AI engineers, ML engineers, and product stakeholders to embed quality from day one
reviewing prompt architectures, agent designs, and system workflows for testability and risk. Continuously research and adopt frontier evaluation benchmarks (RAGAS, MMLU, TruthfulQA, MT-Bench) and emerging AI testing methodologies to keep quality practices at the cutting edge. All About You: Master's/Bachelor's degree in Computer Science, AI/ML, or Software Engineering, with considerable hands-on experience leading AI/ML quality engineering or LLM testing programs in production environments. Demonstrated expertise testing LLM and Gen AI systems
including prompt testing, output evaluation, hallucination detection, RAG pipeline assessment, and agentic workflow validation in real production settings. Deep hands-on knowledge of AI evaluation frameworks and tooling: RAGAS, DeepEval, TruLens, LangSmith, PromptFlow, Weights & Biases Evals, or equivalent platforms. Strong understanding of Gen AI failure modes
hallucination, prompt injection, retrieval grounding failures, context drift, agent loop failures
and proven methods to surface and document them systematically. Strong Python programming skills with the ability to independently build test automation scripts, evaluation pipelines, and API-level integration tests; SQL proficiency required. Working knowledge of LLM ecosystems
OpenAI, Anthropic, Hugging Face, LangChain/LangGraph
sufficient to understand model behavior, prompt structure, and agent architecture deeply enough to test them rigorously. Familiarity with MLOps/LLMOps pipelines (MLflow, Databricks, SageMaker) and experience integrating automated quality gates into CI/CD workflows for AI systems. Experience with cloud AI infrastructure (AWS, Azure, or GCP) and observability tooling for monitoring live AI system behavior and output quality in production. Strong analytical, communication, and stakeholder management skills
with the ability to translate complex AI failure patterns into clear risk assessments and remediation recommendations for both technical and business audiences. Mastercard is a merit-based, inclusive, equal opportunity employer that considers applicants without regard to gender, gender identity, sexual orientation, race, ethnicity, disabled or veteran status, or any other characteristic protected by law. We hire the most qualified candidate for the role. In the US or Canada, if you require accommodations or assistance to complete the online application process or during the recruitment process, please contact reasonable_accommodation@mastercard.com and identify the type of accommodation or assistance you are requesting. Do not include any medical or health information in this email. The Reasonable Accommodations team will respond to your email promptly.