Skip to main content
E

Principal AI Engineer (3 days onsite) in New York

Energy Jobline ZR
2 hours ago
Full-time
On-site
Job DescriptionJob Description

What We're Looking For

Engineering foundation

8–14 years of software engineering experience, with strong hands-on large-scale Python

Working depth in at least one systems or backend — Go, Rust, Java, or C/C++ — and the judgment to know when to reach for it

Strong data structures and algorithms.

Strong understanding of APIs, microservices, and system design

Hands-on experience building and operating data pipelines and production-grade distributed systems.

Agentic AI and LLMs

2+ years of hands-on LLM engineering, with at least couple agentic system you designed and took to production

Production experience with agent frameworks — LangGraph, Google ADK, CrewAI, Claude Agent SDK, or equivalent — and the fluency to move between them as the ecosystem evolves

Experience building MCP (Model Context Protocol) servers and tool-calling interfaces

RAG from first principles: chunking strategy, embeddings, vector and hybrid retrieval, reranking, and response validation

Strong experience with vector databases (Milvus, Pinecone, Weaviate, FAISS, etc. or cloud equivalents)

Design of guardrails and reliability patterns — validators, policy checks, self-correction loops, deterministic fallbacks, circuit breakers, and rollback paths

Roles & Responsibilities

Design and build agentic systems:

Lead the architecture and implementation of tool-calling agents that combine retrieval, structured reasoning, and secure action execution with least-privilege access.

Productionize LLM applications:

Build retrieval pipelines, prompt synthesis, response validation, and self-correction loops, backed by rigorous evaluation.

Own the full stack:

Deliver the data pipelines, backend services, distributed compute, and orchestration layer that agentic systems depend on — not only the model invocation.

Engineer for reliability and governance:

Build validator models, adversarial test suites, and policy checks; enforce deterministic fallbacks and rollback strategies; instrument continuous evaluation.

Optimize for cost and latency:

Drive measurable improvements in token efficiency, response time, and unit economics against defined SLOs.

Codebase ownership:

Build, maintain, and review high-quality Python and SQL, with an emphasis on reusable components, scalability, and performance.

Cloud integration:

Deploy AI applications on AWS, Azure, or GCP with optimized resource usage and robust CI/CD.

Cross-functional collaboration:

Partner with product owners, data scientists, and business SMEs to define requirements and deliver impactful AI products.

Mentoring and technical leadership:

Set engineering standards and share knowledge across the team, raising the bar on AI and software engineering practice.