D
Applied AI Engineer
Day One Partners
2 hours ago
Full-time
On-site
We’re partnering with an early-stage, venture-backed AI startup building the collaboration layer for humans and AI agents.
Ensure you read the information regarding this opportunity thoroughly before making an application. The company is developing systems that allow agents to maintain context, work across long-running tasks, collaborate with people and other agents, and become more effective over time. This is a small, highly technical team tackling problems at the edge of what production agent systems can reliably do today. They’re hiring
1–2 Applied AI Engineers
to take significant ownership of the intelligence and infrastructure behind the product. About the Role You’ll build the systems surrounding frontier models that determine
how agents remember, act, coordinate, and improve. The core scope spans
agent memory, task optimization, agent orchestration, harnesses, evals, and benchmarks , alongside the infrastructure required to run these systems securely and reliably at scale. This is a hands-on engineering role for someone who wants to work below the application layer. You should be excited by the hard parts of making agents actually work in production, not just integrating an LLM into an existing product. What You’ll Do Design and improve
agent memory, context management, task execution, and multi-step orchestration Build and optimize
agent harnesses , including tool use, prompting, planning, feedback loops, and failure recovery Develop
evals and benchmarks
that rigorously measure agent quality, regressions, and improvements Build infrastructure for running and scaling sandboxed agents, including isolation, observability, reliability, and cost/performance Own difficult technical problems end-to-end, from experimentation and architecture through production deployment What We’re Looking For Demonstrated experience
building agent systems, harnesses, evals, benchmarks, or closely related AI infrastructure Strong software engineering fundamentals with depth in
backend, infrastructure, distributed systems, or production AI systems Ability to reason rigorously about agent performance: define success, design experiments, diagnose failures, and prove whether a change actually improved the system Experience thinking through
scalability, security, reliability, and isolation
for complex production workloads High technical ownership and comfort operating on problems where established abstractions or best practices may not exist yet Strong Signals We care considerably more about the
depth and quality of what you’ve built
than a particular number of years of experience. The strongest candidates will have done things like: Built agents or agent infrastructure that
real users depend on in production Designed an
evaluation harness or benchmark
and used it to drive measurable improvements Worked deeply on memory, orchestration, tool use, long-running execution, or agent reliability rather than only prompt/API integration Built technically demanding systems where
latency, concurrency, security, reliability, or scale
meaningfully affected the architecture Taken ambiguous technical problems from first principles through experimentation and into production Why This Role You’ll work on problems where the underlying models are improving rapidly, but the infrastructure around them is still being invented. How should an agent decide what to remember? How do you optimize performance across a task that may run for hours? How do you benchmark something nondeterministic? xsgimln How do you safely run large numbers of agents capable of taking real actions? If you’ve already spent meaningful time wrestling with problems like these and want substantially more ownership over them, we’d love to hear from you.
Ensure you read the information regarding this opportunity thoroughly before making an application. The company is developing systems that allow agents to maintain context, work across long-running tasks, collaborate with people and other agents, and become more effective over time. This is a small, highly technical team tackling problems at the edge of what production agent systems can reliably do today. They’re hiring
1–2 Applied AI Engineers
to take significant ownership of the intelligence and infrastructure behind the product. About the Role You’ll build the systems surrounding frontier models that determine
how agents remember, act, coordinate, and improve. The core scope spans
agent memory, task optimization, agent orchestration, harnesses, evals, and benchmarks , alongside the infrastructure required to run these systems securely and reliably at scale. This is a hands-on engineering role for someone who wants to work below the application layer. You should be excited by the hard parts of making agents actually work in production, not just integrating an LLM into an existing product. What You’ll Do Design and improve
agent memory, context management, task execution, and multi-step orchestration Build and optimize
agent harnesses , including tool use, prompting, planning, feedback loops, and failure recovery Develop
evals and benchmarks
that rigorously measure agent quality, regressions, and improvements Build infrastructure for running and scaling sandboxed agents, including isolation, observability, reliability, and cost/performance Own difficult technical problems end-to-end, from experimentation and architecture through production deployment What We’re Looking For Demonstrated experience
building agent systems, harnesses, evals, benchmarks, or closely related AI infrastructure Strong software engineering fundamentals with depth in
backend, infrastructure, distributed systems, or production AI systems Ability to reason rigorously about agent performance: define success, design experiments, diagnose failures, and prove whether a change actually improved the system Experience thinking through
scalability, security, reliability, and isolation
for complex production workloads High technical ownership and comfort operating on problems where established abstractions or best practices may not exist yet Strong Signals We care considerably more about the
depth and quality of what you’ve built
than a particular number of years of experience. The strongest candidates will have done things like: Built agents or agent infrastructure that
real users depend on in production Designed an
evaluation harness or benchmark
and used it to drive measurable improvements Worked deeply on memory, orchestration, tool use, long-running execution, or agent reliability rather than only prompt/API integration Built technically demanding systems where
latency, concurrency, security, reliability, or scale
meaningfully affected the architecture Taken ambiguous technical problems from first principles through experimentation and into production Why This Role You’ll work on problems where the underlying models are improving rapidly, but the infrastructure around them is still being invented. How should an agent decide what to remember? How do you optimize performance across a task that may run for hours? How do you benchmark something nondeterministic? xsgimln How do you safely run large numbers of agents capable of taking real actions? If you’ve already spent meaningful time wrestling with problems like these and want substantially more ownership over them, we’d love to hear from you.