D

Senior/Staff AI Engineer

DataDirect Networks
2 hours ago
Full-time
On-site
What you’ll doBuild and optimize LLM serving and inference systems for production environmentsImprove performance across GPU and CPU pathwaysWork on KV cache, memory, storage, and throughput bottlenecksDesign and scale systems that support RAG and retrieval-heavy AI workloadsContribute to infrastructure where storage architecture and systems efficiency materially affect AI performanceSolve engineering problems at the intersection of AI, high-performance systems, and distributed infrastructureWhat we’re looking forAn engineer who has spent meaningful time building or optimizing production AI systems, not just experimenting with modelsSomeone who understands how inference performance is shaped by the interaction between compute, memory, storage, and serving architectureDeep hands-on experience working close to the systems layer — for example, improving how workloads run across GPU and CPU resources, reducing bottlenecks, or tuning infrastructure for better throughput and latencyEvidence of real ownership in areas like model serving, retrieval, caching, storage, or distributed performance, rather than purely application-layer AI workThe ability to move comfortably between architecture decisions and hands-on implementation, especially in environments where efficiency and scale matterA background that suggests you can operate in technically demanding environments, whether that comes from AI infrastructure, high-performance systems, storage platforms, or adjacent distributed systems workPhD preferred, but far less important than having built serious systems in the real worldWhy this role is compellingThis is not a “prompt engineering” job.This is not an “AI wrapper” job.This is not a generic backend role with AI sprinkled on top.This is a chance to work on the infrastructure that determines whether modern AI systems are fast, scalable, efficient, and commercially viable.If you want to work on the real mechanics of AI performance — serving, retrieval, compute efficiency, memory behavior, storage architecture, and inference at scale — this is where that work happens.Who will love this roleEngineers who enjoy deep systems problemsBuilders who care about performance, scale, and architecturePeople who want to work where AI meets infrastructureCandidates who would rather solve hard technical bottlenecks than ship surface-level AI featuresWho should not applyThis role is not for:Purely academic researchers without meaningful production ownershipGeneric software engineers without clear AI systems or inference depthCandidates focused mainly on prompt engineering or lightweight application integrationsMLOps generalists who have not worked deeply on serving, storage, or performance-critical AI systemsLocationSanta ClaraEmployment TypeFull timeLocation TypeHybridDepartmentDepartmentEngineeringEngineeringInfinia Engineering