F
Senior Applied AI Engineer
Fast-growing software company
2 hours ago
Full-time
On-site
San Francisco, California, United States
The Role
Raydar is recruiting for this role on behalf of our client. Design and build the core software that lets large language models carry out multi-step work reliably in production. You will shape how models are prompted and combined with conventional code, measure quality through rigorous testing, and diagnose real-world failures to make systems more robust. What You'll Do
Develop the core execution framework for AI agents , including orchestration loops, tool usage and context handling. Run experiments
with models and iterate on agent behavior across realistic, long-running workflows. Define where model decisions end and deterministic, type-checked code begins. Build evaluation suites on
realistic test environments that are dependable enough to gate releases. Analyze production failures and trace each one to the right layer, then deliver systematic fixes. Extend browser and computer-use agents to systems that lack APIs, with strong safety, security and auditability guarantees. Create feedback loops
and data systems that bring higher-quality real-task data into evaluation and training. What We're Looking For
Experience with
agent frameworks or LLM systems that use tools . Strong Python or TypeScript skills and comfort with modern AI tooling. Hands-on background measuring or improving model quality, whether through testing, tuning or crafting instructions for models. Ability to own systems end to end and debug across the whole stack. A systems mindset focused on user outcomes as well as model metrics. Satisfaction in tracking down unpredictable production problems and converting what you learn into lasting fixes. At least 4 years of relevant experience. Ability to work in person five days a week at a startup pace. Bonus Points
Experience building computer-use or browser-automation agents. Familiarity with virtualization and sandboxed execution environments, including scaling them. AI research experience with publications at leading conferences. Experience delivering systems where correctness must hold through partial failure, such as idempotency and resumability. Experience integrating enterprise SaaS APIs. Comfort working with sizable, unstructured data sources such as application logs. Early-stage engineering experience that included working directly with customers. Benefits
Health insurance Unlimited paid time off Location and Work Model
San Francisco, CA, United States On-site, five days a week in office Full-time
Raydar is recruiting for this role on behalf of our client. Design and build the core software that lets large language models carry out multi-step work reliably in production. You will shape how models are prompted and combined with conventional code, measure quality through rigorous testing, and diagnose real-world failures to make systems more robust. What You'll Do
Develop the core execution framework for AI agents , including orchestration loops, tool usage and context handling. Run experiments
with models and iterate on agent behavior across realistic, long-running workflows. Define where model decisions end and deterministic, type-checked code begins. Build evaluation suites on
realistic test environments that are dependable enough to gate releases. Analyze production failures and trace each one to the right layer, then deliver systematic fixes. Extend browser and computer-use agents to systems that lack APIs, with strong safety, security and auditability guarantees. Create feedback loops
and data systems that bring higher-quality real-task data into evaluation and training. What We're Looking For
Experience with
agent frameworks or LLM systems that use tools . Strong Python or TypeScript skills and comfort with modern AI tooling. Hands-on background measuring or improving model quality, whether through testing, tuning or crafting instructions for models. Ability to own systems end to end and debug across the whole stack. A systems mindset focused on user outcomes as well as model metrics. Satisfaction in tracking down unpredictable production problems and converting what you learn into lasting fixes. At least 4 years of relevant experience. Ability to work in person five days a week at a startup pace. Bonus Points
Experience building computer-use or browser-automation agents. Familiarity with virtualization and sandboxed execution environments, including scaling them. AI research experience with publications at leading conferences. Experience delivering systems where correctness must hold through partial failure, such as idempotency and resumability. Experience integrating enterprise SaaS APIs. Comfort working with sizable, unstructured data sources such as application logs. Early-stage engineering experience that included working directly with customers. Benefits
Health insurance Unlimited paid time off Location and Work Model
San Francisco, CA, United States On-site, five days a week in office Full-time