Skip to main content
C

Senior Generative AI Engineer

Cogniify
1 hour ago
Full-time
On-site
San Francisco, California, United States
Cogniify is building production-grade Generative AI applications and is hiring a hands-on engineer to design, build, and deploy systems that power LLM-powered products. This role in the San Francisco Bay Area (hybrid) focuses on core capabilities like RAG pipelines, intelligent agent workflows, and reliable integration with internal tools, external APIs, and enterprise applications. You will work across the full lifecycle, from model and framework selection to evaluation, safety guardrails, deployment, and ongoing LLMOps monitoring. Key Responsibilities

Design and develop scalable

LLM-powered applications

using

Python . Build

RAG pipelines

including document processing, embeddings, vector databases, semantic search, and

reranking . Develop

AI-agent and multi-agent workflows

with tool calling, memory, orchestration, and

human approval steps . Integrate LLMs with internal systems, external APIs, databases, and enterprise applications. Evaluate and select foundation models based on

accuracy ,

latency ,

cost ,

security , and business needs. Improve

prompt quality , retrieval accuracy, response time, and

token usage . Implement

safety guardrails , output validation, access controls, and fallback mechanisms. Build automated evaluation frameworks to assess response quality, hallucination, relevance, and reliability. Containerize applications with

Docker

and deploy on

AWS, Azure, or GCP . Apply

monitoring

and

LLMOps

practices for model performance, cost, latency, errors, and production usage. Collaborate with product, engineering, and business teams to translate requirements into reliable AI solutions. Document technical architecture, design decisions, APIs, and operational processes. Requirements

Bachelor's or Master's degree in Computer Science, Engineering, or a related discipline, or equivalent practical software-development experience. 6-9 years

of professional software-development experience, including strong hands-on experience with

Python . Experience developing backend services and integrating

REST APIs . Hands-on experience building

LLM or Generative AI applications . Practical experience implementing RAG using

embeddings , semantic search, and

vector databases . Experience with

LangChain ,

LangGraph ,

LlamaIndex , or a comparable LLM application framework. Experience integrating foundation models through APIs such as

OpenAI ,

Anthropic Claude ,

Gemini , or

Azure OpenAI . Working knowledge of vector databases including

Pinecone ,

Weaviate ,

Milvus ,

Qdrant ,

Chroma ,

FAISS , or

pgvector . Experience with

Docker

and deployment on at least one of:

AWS ,

Azure , or

GCP . Understanding of prompt engineering, hallucination reduction, output validation, and LLM evaluation. Strong understanding of software engineering practices including

Git , testing, debugging, and clean code. Ability to communicate technical solutions clearly to both technical and non-technical stakeholders. Technologies and Frameworks

Programming and AI:

Python, Retrieval-Augmented Generation (RAG), embeddings, semantic search, reranking, Hugging Face Transformers, PyTorch, LoRA, QLoRA, vLLM, TGI, Ollama Frameworks:

LangChain, LangGraph, LlamaIndex, LangSmith, Langfuse, Arize Phoenix, AutoGen, CrewAI, Semantic Kernel Observability/ML tools:

MLflow, Weights & Biases Model access:

OpenAI, Anthropic Claude, Gemini, Azure OpenAI Vector databases:

Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS, pgvector Cloud and delivery:

Docker, AWS, Azure, GCP, Kubernetes, CI/CD pipelines, REST APIs Benefits

Unlimited PTO. Very generous parental leave, much above industry standards. Entrepreneurial culture where pushing limits and taking risks is everyday business. Open communication with management and company leadership. Small, dynamic teams = massive impact. Medical, Dental and Vision coverage for employees. Access to Disability & Life insurance. Mental health and wellbeing support. Annual bonus program. Employer Stock Purchase Program (ESPP). Yearly team building experiences. Mentorship and sponsorship opportunities. Manager resources and support. Success in the Role

Production-ready AI applications that are accurate, secure, and maintainable. RAG systems that retrieve relevant information and reduce hallucinations. AI-agent workflows that reliably complete business tasks and integrate with existing systems. Measurable improvements in response quality, latency, and inference cost. Clear monitoring of application performance, usage, errors, and model behavior. Salary range:

USD 150,000 - 170,000 per year (US East/West Coast). Work location:

Hybrid remote, San Francisco Bay Area, CA.

#J-18808-Ljbffr