D
Senior AI Engineer II
DataJobs
2 hours ago
Full-time
On-site
Atlanta, Georgia, United States
Design, build, and operate
an agentic AI framework and production solutions on Microsoft Azure for end-to-end business impact.
Responsibilities
Design, build, and evolve a shared agentic framework for teams to build, deploy, and operate agentic solutions across the organization
Provide reusable platform building blocks including agent scaffolding, tool and integration interfaces, context and memory management, evaluation, guardrails, and observability
Convert one-off agent work into reusable components such as service templates, harness components, MCP server patterns, evaluation harnesses, and deployment pipelines
Design and maintain MCP integrations with enterprise systems (inventory, merchandising, and downstream systems) so agents can use live operational data without rebuilding integrations
Define and document standards, contracts, and reference architectures used for agentic workloads
Make architectural decisions for reliability, scalability, security, and observability for AI services deployed in Azure
Deliver production agentic solutions end to end, from problem framing through deployment and ongoing operations, using the work to harden the platform
Apply the right approach per use case: tool calling, multi-step planning, retrieval, memory, human-in-the-loop review, or deterministic services when agents are not appropriate
Build and tune retrieval when warranted, including chunking strategies, vector indexing, retrieval ranking, and context engineering using Azure AI Search
Contribute across the stack, including an Angular front end and a Python-based service and LLMOps layer
Support partner teams transitioning mature agentic products with documentation, runbooks, and knowledge transfer
Establish evaluation and regression testing as a first-class platform capability using eval sets, LLM-as-judge scoring, task-level success metrics, and CI regression gates
Own evaluation pipelines for retrieval-based components using Ragas to track faithfulness, answer relevance, and context precision across releases
Instrument agentic systems with Langfuse alongside Azure Monitor to enable tracing, latency, token and cost attribution, tool-call success rates, quality signals, and failure mode analysis
Manage AI cost and performance through token budgeting, caching, model routing, right-sizing, and latency optimization
Own production services including on-call participation, incident response, and follow-through after incidents
Troubleshoot complex issues across agent behavior, retrieval quality, hallucination, latency, and integration reliability
Use spec-driven development by turning ambiguous requests into clear specifications and acceptance criteria, keeping specs and implementation synchronized
Set and model standards for AI-augmented development, including review discipline for agentic coding tools
Conduct code reviews and mentor peers through technical influence and continuous learning
Partner with product managers and business stakeholders to identify automatable workflows, translate them into technical requirements, and help prioritize the backlog
Act as a technical consultant to teams adopting the platform so solutions are built effectively on top of it
Participate in agile ceremonies including sprint planning, retrospectives, and daily stand-ups
Communicate technical trade-offs and architectural decisions to both technical and non-technical audiences
Partner with security, data, and platform teams on data governance, PII handling, prompt injection defense, and responsible use
Evaluate emerging agent capabilities, tooling, protocols, and Azure OpenAI / AI Foundry updates, and recommend and prototype improvements to keep the platform current
Establish and maintain engineering best practices including CI/CD pipelines, infrastructure as code, code quality standards, and security practices for AI workloads
Continuously reduce time and cost to bring the next agentic solution to production
Requirements
7-10 years of professional software engineering experience with increasing scope and ownership
Deep backend engineering proficiency in at least one of: C#/.NET, Java, Node.js/TypeScript, or Python
Experience with service-based and microservice architectures, RESTful API design, and asynchronous service communication including API versioning, contract design, error semantics, and backward compatibility
Production experience with cloud-based serverless microservices (Azure Functions, Container Apps, or equivalent)
Event-driven architecture experience including queues, pub/sub, idempotency, retries, and dead-letter handling
Solid data fundamentals including relational and NoSQL data modeling, query performance, and transactional correctness
Demonstrated code quality ownership with automated testing, code review, and CI/CD as normal practice
Shipped and operated agentic solutions in production (not prototypes or POCs only)
Hands-on harness engineering for production model scaffolding: context construction and management, tool/function-call interfaces, multi-step planning and control flow, memory and state, structured output, guardrails and validation, human-in-the-loop checkpoints, and graceful failure/fallback behavior
Azure OpenAI services experience including LLM APIs and Azure AI Search, plus associated Azure infrastructure (Azure AI Foundry or AWS Bedrock experience also applies)
Agent and LLM orchestration frameworks such as LangChain, LangGraph, or similar
MCP (Model Context Protocol) or other tool-calling/function-calling patterns for LLM-to-system integrations
Knowledge of AI-specific failure and risk modes and mitigation techniques, including hallucination, prompt injection, data leakage, non-determinism, and runaway tool loops
Judgment on when an LLM/agent is appropriate and how to bound and validate its output
Daily working fluency with agentic coding tools (Claude Code is the standard) including Claude Code, Cursor, GitHub Copilot, Devin, or equivalent
Ability to articulate where agentic coding tools accelerate delivery, where they do not, and how to review and test model-generated code responsibly
Ability to describe measurable impact on throughput and quality
Cloud environments experience (Azure preferred) including compute, storage, networking, and IAM fundamentals
Cloud security fundamentals including identity and access management, secrets management, and network boundaries
Agile/scrum experience including sprint ceremonies, story estimation, and backlog grooming
Angular (or comparable modern front-end framework) working knowledge
Clear written and verbal communication with technical and non-technical partners, comfortable with ambiguity and able to move from business need to scoped, specified, shippable increments
Technologies
C#/.NET, Java, Node.js/TypeScript, Python
REST APIs
Azure Functions, Container Apps
Azure OpenAI, Azure AI Search, Azure AI Foundry, AWS Bedrock
LangChain, LangGraph
MCP (Model Context Protocol)
Claude Code, Cursor, GitHub Copilot, Devin
Angular
Ragas, Langfuse
Azure Monitor, Application Insights
CI/CD pipelines
Infrastructure as code
Bicep, Terraform, ARM
RAG where warranted
Benefits
Bonus opportunities
Career advancement opportunities at every level
401k with company match
Employee Stock Purchase Plan
Referral Bonus Program
Medical, Dental, Vision, Life, and other Insurance Plans (subject to eligibility criteria)
Paid vacation and sick time for eligible associates
Paid holidays plus a personal holiday
Paid Volunteer Time Off that starts on Day 1
Education
Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent professional experience
Demonstrated capability is weighted above credentials
Work Environment
Hybrid position based at Atlanta, GA headquarters (onsite)
Standard office schedule Monday through Friday during core business hours
Collaborative, open-plan office environment within the IT department with dedicated space for focused engineering work
Regular in-person collaboration with product managers, business stakeholders, and the engineering team
Occasional visits to retail store locations may be required to gather associate feedback and observe how the product is used in context
Some extended hours may be needed around major releases or on-call rotations for production incidents
Standard physical requirements of a professional office environment apply (prolonged sitting, use of a computer workstation, and participation in in-person and video meetings)
Working Conditions (Travel & Environment)
Travel may be required including air and car travel
Noise level typically quiet to moderate
Physical/Sensory Requirements
Sedentary work involving sitting most of the time with brief periods of walking or standing
Ability to exert 10-20 pounds of force occasionally, and/or negligible amount of force frequently to lift, carry, push, pull, or otherwise move objects
Nice to Have
Spec-driven development, including using specs to drive AI-assisted implementation
Building internal developer platforms, frameworks, or SDKs consumed by other engineering teams
Building or publishing MCP servers, not just consuming them
LLM and agent evaluation frameworks such as Ragas, G-Eval, LLM-as-judge, or agent trajectory evaluation
AI observability tooling such as Langfuse (our platform) or equivalent tracing and evaluation systems
Retrieval infrastructure beyond Azure AI Search (pgvector, Pinecone, Elastic, or hybrid search design)
Multi-agent orchestration, agent-to-agent protocols, or durable-long-running workflow engines
Infrastructure as code (Bicep, Terraform, ARM)
Fine-tuning and model adaptation (LoRA/PEFT, distillation, or evaluating fine-tuning against prompting and retrieval alternatives)
LLMOps tooling such as MLflow, Weights & Biases, or Azure ML
Retail systems familiarity (POS, OMS, inventory/merchandising platforms)
Center of Excellence or innovation-team experience within a larger enterprise
Open source contributions, technical writing, or speaking in the AI engineering space
#J-18808-Ljbffr
an agentic AI framework and production solutions on Microsoft Azure for end-to-end business impact.
Responsibilities
Design, build, and evolve a shared agentic framework for teams to build, deploy, and operate agentic solutions across the organization
Provide reusable platform building blocks including agent scaffolding, tool and integration interfaces, context and memory management, evaluation, guardrails, and observability
Convert one-off agent work into reusable components such as service templates, harness components, MCP server patterns, evaluation harnesses, and deployment pipelines
Design and maintain MCP integrations with enterprise systems (inventory, merchandising, and downstream systems) so agents can use live operational data without rebuilding integrations
Define and document standards, contracts, and reference architectures used for agentic workloads
Make architectural decisions for reliability, scalability, security, and observability for AI services deployed in Azure
Deliver production agentic solutions end to end, from problem framing through deployment and ongoing operations, using the work to harden the platform
Apply the right approach per use case: tool calling, multi-step planning, retrieval, memory, human-in-the-loop review, or deterministic services when agents are not appropriate
Build and tune retrieval when warranted, including chunking strategies, vector indexing, retrieval ranking, and context engineering using Azure AI Search
Contribute across the stack, including an Angular front end and a Python-based service and LLMOps layer
Support partner teams transitioning mature agentic products with documentation, runbooks, and knowledge transfer
Establish evaluation and regression testing as a first-class platform capability using eval sets, LLM-as-judge scoring, task-level success metrics, and CI regression gates
Own evaluation pipelines for retrieval-based components using Ragas to track faithfulness, answer relevance, and context precision across releases
Instrument agentic systems with Langfuse alongside Azure Monitor to enable tracing, latency, token and cost attribution, tool-call success rates, quality signals, and failure mode analysis
Manage AI cost and performance through token budgeting, caching, model routing, right-sizing, and latency optimization
Own production services including on-call participation, incident response, and follow-through after incidents
Troubleshoot complex issues across agent behavior, retrieval quality, hallucination, latency, and integration reliability
Use spec-driven development by turning ambiguous requests into clear specifications and acceptance criteria, keeping specs and implementation synchronized
Set and model standards for AI-augmented development, including review discipline for agentic coding tools
Conduct code reviews and mentor peers through technical influence and continuous learning
Partner with product managers and business stakeholders to identify automatable workflows, translate them into technical requirements, and help prioritize the backlog
Act as a technical consultant to teams adopting the platform so solutions are built effectively on top of it
Participate in agile ceremonies including sprint planning, retrospectives, and daily stand-ups
Communicate technical trade-offs and architectural decisions to both technical and non-technical audiences
Partner with security, data, and platform teams on data governance, PII handling, prompt injection defense, and responsible use
Evaluate emerging agent capabilities, tooling, protocols, and Azure OpenAI / AI Foundry updates, and recommend and prototype improvements to keep the platform current
Establish and maintain engineering best practices including CI/CD pipelines, infrastructure as code, code quality standards, and security practices for AI workloads
Continuously reduce time and cost to bring the next agentic solution to production
Requirements
7-10 years of professional software engineering experience with increasing scope and ownership
Deep backend engineering proficiency in at least one of: C#/.NET, Java, Node.js/TypeScript, or Python
Experience with service-based and microservice architectures, RESTful API design, and asynchronous service communication including API versioning, contract design, error semantics, and backward compatibility
Production experience with cloud-based serverless microservices (Azure Functions, Container Apps, or equivalent)
Event-driven architecture experience including queues, pub/sub, idempotency, retries, and dead-letter handling
Solid data fundamentals including relational and NoSQL data modeling, query performance, and transactional correctness
Demonstrated code quality ownership with automated testing, code review, and CI/CD as normal practice
Shipped and operated agentic solutions in production (not prototypes or POCs only)
Hands-on harness engineering for production model scaffolding: context construction and management, tool/function-call interfaces, multi-step planning and control flow, memory and state, structured output, guardrails and validation, human-in-the-loop checkpoints, and graceful failure/fallback behavior
Azure OpenAI services experience including LLM APIs and Azure AI Search, plus associated Azure infrastructure (Azure AI Foundry or AWS Bedrock experience also applies)
Agent and LLM orchestration frameworks such as LangChain, LangGraph, or similar
MCP (Model Context Protocol) or other tool-calling/function-calling patterns for LLM-to-system integrations
Knowledge of AI-specific failure and risk modes and mitigation techniques, including hallucination, prompt injection, data leakage, non-determinism, and runaway tool loops
Judgment on when an LLM/agent is appropriate and how to bound and validate its output
Daily working fluency with agentic coding tools (Claude Code is the standard) including Claude Code, Cursor, GitHub Copilot, Devin, or equivalent
Ability to articulate where agentic coding tools accelerate delivery, where they do not, and how to review and test model-generated code responsibly
Ability to describe measurable impact on throughput and quality
Cloud environments experience (Azure preferred) including compute, storage, networking, and IAM fundamentals
Cloud security fundamentals including identity and access management, secrets management, and network boundaries
Agile/scrum experience including sprint ceremonies, story estimation, and backlog grooming
Angular (or comparable modern front-end framework) working knowledge
Clear written and verbal communication with technical and non-technical partners, comfortable with ambiguity and able to move from business need to scoped, specified, shippable increments
Technologies
C#/.NET, Java, Node.js/TypeScript, Python
REST APIs
Azure Functions, Container Apps
Azure OpenAI, Azure AI Search, Azure AI Foundry, AWS Bedrock
LangChain, LangGraph
MCP (Model Context Protocol)
Claude Code, Cursor, GitHub Copilot, Devin
Angular
Ragas, Langfuse
Azure Monitor, Application Insights
CI/CD pipelines
Infrastructure as code
Bicep, Terraform, ARM
RAG where warranted
Benefits
Bonus opportunities
Career advancement opportunities at every level
401k with company match
Employee Stock Purchase Plan
Referral Bonus Program
Medical, Dental, Vision, Life, and other Insurance Plans (subject to eligibility criteria)
Paid vacation and sick time for eligible associates
Paid holidays plus a personal holiday
Paid Volunteer Time Off that starts on Day 1
Education
Bachelor's degree in Computer Science, Engineering, or a related field, or equivalent professional experience
Demonstrated capability is weighted above credentials
Work Environment
Hybrid position based at Atlanta, GA headquarters (onsite)
Standard office schedule Monday through Friday during core business hours
Collaborative, open-plan office environment within the IT department with dedicated space for focused engineering work
Regular in-person collaboration with product managers, business stakeholders, and the engineering team
Occasional visits to retail store locations may be required to gather associate feedback and observe how the product is used in context
Some extended hours may be needed around major releases or on-call rotations for production incidents
Standard physical requirements of a professional office environment apply (prolonged sitting, use of a computer workstation, and participation in in-person and video meetings)
Working Conditions (Travel & Environment)
Travel may be required including air and car travel
Noise level typically quiet to moderate
Physical/Sensory Requirements
Sedentary work involving sitting most of the time with brief periods of walking or standing
Ability to exert 10-20 pounds of force occasionally, and/or negligible amount of force frequently to lift, carry, push, pull, or otherwise move objects
Nice to Have
Spec-driven development, including using specs to drive AI-assisted implementation
Building internal developer platforms, frameworks, or SDKs consumed by other engineering teams
Building or publishing MCP servers, not just consuming them
LLM and agent evaluation frameworks such as Ragas, G-Eval, LLM-as-judge, or agent trajectory evaluation
AI observability tooling such as Langfuse (our platform) or equivalent tracing and evaluation systems
Retrieval infrastructure beyond Azure AI Search (pgvector, Pinecone, Elastic, or hybrid search design)
Multi-agent orchestration, agent-to-agent protocols, or durable-long-running workflow engines
Infrastructure as code (Bicep, Terraform, ARM)
Fine-tuning and model adaptation (LoRA/PEFT, distillation, or evaluating fine-tuning against prompting and retrieval alternatives)
LLMOps tooling such as MLflow, Weights & Biases, or Azure ML
Retail systems familiarity (POS, OMS, inventory/merchandising platforms)
Center of Excellence or innovation-team experience within a larger enterprise
Open source contributions, technical writing, or speaking in the AI engineering space
#J-18808-Ljbffr