A
AI Engineer, Agents & Retrieval
AtomMatrix
3 hours ago
Full-time
On-site
San Francisco, California, United States
An agent that sounds confident but makes something up is worse than no agent at all. We build agents that act
— call tools, change records, resolve tickets — and stay grounded in the customer's own knowledge. We're looking for an engineer who can push that quality forward and prove it with numbers.
Agents that act, grounded in real knowledge.
You'll work on the agent runtime — the loop that understands a message, retrieves relevant knowledge, calls the tools that do the work, and decides when to hand off to a person — and on the retrieval system that keeps answers grounded. This is applied AI engineering: building and improving retrieval-augmented generation, tool calling, and the evaluation harness that tells you whether a change actually helped. You'll care about latency and cost as much as accuracy, because these agents run on live customer traffic across chat, SMS, voice, and email. What you'll do
Build and improve the agent runtime: prompting, tool calling, multi-step runs, guardrails, and handoff. Own the retrieval pipeline — ingestion, chunking, embeddings, and ranking — so agents answer from the customer's approved sources with citations. Design evaluation: datasets, metrics, and offline/online tests that measure quality, grounding, and regressions before they ship. Tune for latency and cost — caching, model selection, and retrieval budgets — without giving up quality. Work with product and design on how confidence, citations, and handoff show up to operators and customers. What we're looking for
4+ years of software engineering, with recent hands-on work building LLM-powered features in production. Practical experience with retrieval-augmented generation, embeddings, and vector search. Comfortable building evaluation harnesses and reasoning about model quality with data, not vibes. Strong general engineering — Python and/or TypeScript, APIs, and production systems. Healthy skepticism about model output, and a habit of designing for failure and human review. Nice to have
Experience with agent frameworks, tool/function calling, or orchestration at scale. Familiarity with conversational systems across voice or messaging. Background in information retrieval or applied ML.
#J-18808-Ljbffr
You'll work on the agent runtime — the loop that understands a message, retrieves relevant knowledge, calls the tools that do the work, and decides when to hand off to a person — and on the retrieval system that keeps answers grounded. This is applied AI engineering: building and improving retrieval-augmented generation, tool calling, and the evaluation harness that tells you whether a change actually helped. You'll care about latency and cost as much as accuracy, because these agents run on live customer traffic across chat, SMS, voice, and email. What you'll do
Build and improve the agent runtime: prompting, tool calling, multi-step runs, guardrails, and handoff. Own the retrieval pipeline — ingestion, chunking, embeddings, and ranking — so agents answer from the customer's approved sources with citations. Design evaluation: datasets, metrics, and offline/online tests that measure quality, grounding, and regressions before they ship. Tune for latency and cost — caching, model selection, and retrieval budgets — without giving up quality. Work with product and design on how confidence, citations, and handoff show up to operators and customers. What we're looking for
4+ years of software engineering, with recent hands-on work building LLM-powered features in production. Practical experience with retrieval-augmented generation, embeddings, and vector search. Comfortable building evaluation harnesses and reasoning about model quality with data, not vibes. Strong general engineering — Python and/or TypeScript, APIs, and production systems. Healthy skepticism about model output, and a habit of designing for failure and human review. Nice to have
Experience with agent frameworks, tool/function calling, or orchestration at scale. Familiarity with conversational systems across voice or messaging. Background in information retrieval or applied ML.
#J-18808-Ljbffr