C

Senior AI Engineer - Agentic Systems & Data Pipelines

Collaboration.Ai
2 hours ago
Full-time
On-site
Minneapolis, Minnesota, United States
Job TitleAgentic Systems EngineerJob DescriptionYou'll build the agentic systems and data pipelines behind NetworkOS's AI capabilities: production agent workflows built on industry-leading agent SDKs and harnesses, MCP servers, and Agent Skills standards; the eval and observability layer that keeps LLM quality measurable; and the ingestion pipelines that turn messy, diverse data sources into queryable knowledge.This is an execution seat, not an ivory tower. You'll commit code every week, ship agents as product capability rather than demos, and help shape a roadmap that's heading deep into graph + agents territory — for customers in defense, healthcare, and regulated enterprise.Agents in production. Pipelines that hold. Evals that keep everyone honest. This opportunity is remote with a preference for candidates in the Twin Cities area (Minneapolis, Saint Paul); however all candidates are encouraged to apply!What You'll DoShip production agent systems — design, build, and operate agentic workflows (agent SDKs, MCP servers, Agent Skills standards) powering AI-driven matching, analysis, and data intelligenceOperationalize LLM quality — build the eval and observability layer with Langfuse, golden datasets, LLM-as-judge patterns, and FinOps-style tracking so every workflow has measurable quality, cost, and latencyEngineer data pipelines — robust ingestion of documents, structured data, and external sources into searchable knowledge bases with quality validation, deduplication, and incremental updatesOwn retrieval quality — hybrid search combining vector, keyword, and metadata retrieval, continuously improved through reranking, query expansion, and contextual compressionAccelerate with AI — build custom MCP tools and Agent Skills that make the whole engineering team measurably fasterExecute alongside the team — pair with full-stack engineers on AI integration points, contribute to incident response for AI services, and keep your hands in the codeOur Tech StackLanguages: Python (primary); Kotlin (core platform language at CAI); TypeScript/Node.js and other modern languages (secondary)AI/ML: FastAPI, Pydantic; multi-provider LLM SDKs (Anthropic, OpenAI, and others)Agentic Tooling: Claude Code/Codex/etc.; industry-leading agent SDKs and harnesses; MCP servers; Agent Skills standardsLLM Operations: Langfuse + evals (golden datasets, LLM-as-judge); in-house FinOps tracking (token usage, latency, cost); multi-provider orchestration including AWS BedrockSearch & Retrieval: Vector databases, OpenSearch, embedding modelsData: PostgreSQL, Amazon S3; streaming pipelines (Kafka/Kinesis) where neededInfrastructure: Docker, Kubernetes (AWS EKS); DataDog + OpenTelemetry observabilityWhat We're Looking ForMust Haves7+ years of professional software engineering experience, with 3+ years focused on AI/ML or data engineeringProduction agentic/LLM application experience — built and operated systems around LLM APIs (Anthropic, OpenAI) serving real users: agents, tool-use, or orchestrated LLM workflowsData engineering background — robust, scalable pipelines for AI/ML workloadsLLM operations experience — evals and observability for production LLM systems (quality, cost, latency)Production retrieval experience — vector databases and/or search engines (OpenSearch, Elasticsearch)Modern Python stack proficiency — FastAPI, Pydantic, async/await, modern dependency managementAI-native workflows — demonstrated ability to leverage Claude Code/Codex or similar agentic coding tools to accelerate developmentExperience with Docker, Kubernetes, and AWSUS citizenship required (DoD contracting — IL4/IL5 environments — and FedRAMP compliance)Nice-to-HavesDeep agentic ecosystem experience — Agent Skills standards, custom MCP servers, agent SDKs across major vendorsAdvanced RAG expertise — GraphRAG, agentic RAG, contextual retrieval, reranking strategiesGraph data experience — knowledge graphs, graph databases, or graph-based retrievalModel selection & rightsizing — matching models to domain-specific use cases across quality, cost, and latency tradeoffsStreaming data experience (Kafka, Kinesis) for real-time knowledge base updatesResearch background, open-source contributions, or an advanced degree in ML/IR/NLPWhy Join Collaboration AI?Real AI engineering, not a wrapper shop. Production agents, hybrid retrieval, continuous evals, and a roadmap heading into graph + agents — with the autonomy to shape how it's built.AI-native by default. We build with AI, not just for AI. Agentic coding tools (Claude Code/Codex/etc.), agent SDKs and harnesses, MCP servers, and Agent Skills standards are how we work daily — you'll both use and build them.Work that matters. Defense, healthcare, and regulated industries — SOC 2 and NIST compliance, FedRAMP readiness, and customers whose missions demand AI they can trust.Small, senior team. Early-stage impact with your work visible from week one. You'll help set the bar for how AI engineering is done here.