C
Volunteer AI Engineer (Evaluations & Benchmarking) AI Apps for Human Communication
CROSSING PARTY LINES INC
3 hours ago
Full-time
On-site
Help build the rigorous evaluation pipelines and benchmarks that ensure conversational AI agents remain objective, depolarizing, emotionally safe, and constructive.
Crossing Party Lines is seeking a detail-oriented
AI Engineer specializing in Evaluations (Evals)
to join our Apps and Agents team. Evaluating conversational agents designed for politically charged, high-stakes dialogue requires moving beyond basic accuracy checks. You will define, implement, and automate evaluation suites that measure nuanced qualities such as neutrality, tone calibration, reframing effectiveness, cognitive empathy, and resistance to bias or sycophancy. About the Apps and Agents Team Our team blends generative AI models and multi-agent workflows with Crossing Party Lines’ proprietary depolarization frameworks. As we prepare applications for Y Combinator and work toward broader public deployment, reliable evaluation is central to building tools that users and communities can trust. What You’ll Do Design and maintain systematic evaluation datasets (golden sets) reflecting diverse, realistic conversational conflict and political viewpoints. Implement multi-dimensional evaluation rubrics—combining deterministic checks, reference-based metrics, and LLM-as-a-judge pipelines. Benchmark agent outputs against specific communication standards: active listening, cognitive perspective-taking, de-escalation, and non-judgmental feedback. Run regression tests, track performance across prompt and model revisions, and integrate automated eval gates into CI/CD workflows. Conduct adversarial red-teaming to uncover subtle bias, conversational failure modes, hallucinations, or toxic drift. Collaborate with depolarization experts, product designers, and backend developers to calibrate evaluation criteria against real-world human judgment. You May Be a Good Fit If You Have experience with Python and LLM application frameworks (e.g., LangChain/LangGraph, LlamaIndex, or raw API tool orchestration). Understand modern LLM evaluation concepts and tools (e.g., promptfoo, DeepEval, Ragas, Phoenix, LangSmith, or custom LLM-as-a-judge scorers). Have a keen analytical mindset and can translate qualitative human communication principles into clear quantitative metrics. Are comfortable analyzing edge cases across a wide spectrum of social and political perspectives without imposing personal viewpoints. Value transparency, scientific rigor, and reproducibility in AI benchmarking. What You’ll Gain Direct experience solving one of the most critical challenges in production AI: behavioral evaluation and alignment for multi-turn conversational agents. A key engineering role on early-stage AI products being prepped for high-profile accelerators and widespread adoption. Collaboration with communication experts, prompt engineers, and product builders. Professional references recognizing your technical rigor and contributions to production-grade AI reliability. Time, location, and support This is a remote volunteer position. You may complete most of the work on your own schedule. We hope volunteers can contribute approximately
8 or more hours per week , although hours are flexible. Team members should also be available for
one or two online team meetings per week . You will report to Lisa Swallow, co-founder of Crossing Party Lines and lead of the Apps and Agents team. You will work closely with other volunteers and developers and will receive background information, product specifications, and access to the CPL frameworks underlying the applications. We welcome applicants from different backgrounds, professions, political viewpoints, and lived experiences. Crossing Party Lines does not ask volunteers to agree politically. We ask them to approach differences with curiosity, respect, and a sincere desire to understand.
AI Engineer specializing in Evaluations (Evals)
to join our Apps and Agents team. Evaluating conversational agents designed for politically charged, high-stakes dialogue requires moving beyond basic accuracy checks. You will define, implement, and automate evaluation suites that measure nuanced qualities such as neutrality, tone calibration, reframing effectiveness, cognitive empathy, and resistance to bias or sycophancy. About the Apps and Agents Team Our team blends generative AI models and multi-agent workflows with Crossing Party Lines’ proprietary depolarization frameworks. As we prepare applications for Y Combinator and work toward broader public deployment, reliable evaluation is central to building tools that users and communities can trust. What You’ll Do Design and maintain systematic evaluation datasets (golden sets) reflecting diverse, realistic conversational conflict and political viewpoints. Implement multi-dimensional evaluation rubrics—combining deterministic checks, reference-based metrics, and LLM-as-a-judge pipelines. Benchmark agent outputs against specific communication standards: active listening, cognitive perspective-taking, de-escalation, and non-judgmental feedback. Run regression tests, track performance across prompt and model revisions, and integrate automated eval gates into CI/CD workflows. Conduct adversarial red-teaming to uncover subtle bias, conversational failure modes, hallucinations, or toxic drift. Collaborate with depolarization experts, product designers, and backend developers to calibrate evaluation criteria against real-world human judgment. You May Be a Good Fit If You Have experience with Python and LLM application frameworks (e.g., LangChain/LangGraph, LlamaIndex, or raw API tool orchestration). Understand modern LLM evaluation concepts and tools (e.g., promptfoo, DeepEval, Ragas, Phoenix, LangSmith, or custom LLM-as-a-judge scorers). Have a keen analytical mindset and can translate qualitative human communication principles into clear quantitative metrics. Are comfortable analyzing edge cases across a wide spectrum of social and political perspectives without imposing personal viewpoints. Value transparency, scientific rigor, and reproducibility in AI benchmarking. What You’ll Gain Direct experience solving one of the most critical challenges in production AI: behavioral evaluation and alignment for multi-turn conversational agents. A key engineering role on early-stage AI products being prepped for high-profile accelerators and widespread adoption. Collaboration with communication experts, prompt engineers, and product builders. Professional references recognizing your technical rigor and contributions to production-grade AI reliability. Time, location, and support This is a remote volunteer position. You may complete most of the work on your own schedule. We hope volunteers can contribute approximately
8 or more hours per week , although hours are flexible. Team members should also be available for
one or two online team meetings per week . You will report to Lisa Swallow, co-founder of Crossing Party Lines and lead of the Apps and Agents team. You will work closely with other volunteers and developers and will receive background information, product specifications, and access to the CPL frameworks underlying the applications. We welcome applicants from different backgrounds, professions, political viewpoints, and lived experiences. Crossing Party Lines does not ask volunteers to agree politically. We ask them to approach differences with curiosity, respect, and a sincere desire to understand.