C
Senior Generative AI Engineer
Cogniify
1 hour ago
Full-time
On-site
San Francisco, California, United States
Cogniify is building production-grade Generative AI applications and is hiring a hands-on engineer to design, build, and deploy systems that power LLM-powered products. This role in the San Francisco Bay Area (hybrid) focuses on core capabilities like RAG pipelines, intelligent agent workflows, and reliable integration with internal tools, external APIs, and enterprise applications.
You will work across the full lifecycle, from model and framework selection to evaluation, safety guardrails, deployment, and ongoing LLMOps monitoring.
Key Responsibilities
Design and develop scalable
LLM-powered applications
using
Python . Build
RAG pipelines
including document processing, embeddings, vector databases, semantic search, and
reranking . Develop
AI-agent and multi-agent workflows
with tool calling, memory, orchestration, and
human approval steps . Integrate LLMs with internal systems, external APIs, databases, and enterprise applications. Evaluate and select foundation models based on
accuracy ,
latency ,
cost ,
security , and business needs. Improve
prompt quality , retrieval accuracy, response time, and
token usage . Implement
safety guardrails , output validation, access controls, and fallback mechanisms. Build automated evaluation frameworks to assess response quality, hallucination, relevance, and reliability. Containerize applications with
Docker
and deploy on
AWS, Azure, or GCP . Apply
monitoring
and
LLMOps
practices for model performance, cost, latency, errors, and production usage. Collaborate with product, engineering, and business teams to translate requirements into reliable AI solutions. Document technical architecture, design decisions, APIs, and operational processes. Requirements
Bachelor's or Master's degree in Computer Science, Engineering, or a related discipline, or equivalent practical software-development experience. 6-9 years
of professional software-development experience, including strong hands-on experience with
Python . Experience developing backend services and integrating
REST APIs . Hands-on experience building
LLM or Generative AI applications . Practical experience implementing RAG using
embeddings , semantic search, and
vector databases . Experience with
LangChain ,
LangGraph ,
LlamaIndex , or a comparable LLM application framework. Experience integrating foundation models through APIs such as
OpenAI ,
Anthropic Claude ,
Gemini , or
Azure OpenAI . Working knowledge of vector databases including
Pinecone ,
Weaviate ,
Milvus ,
Qdrant ,
Chroma ,
FAISS , or
pgvector . Experience with
Docker
and deployment on at least one of:
AWS ,
Azure , or
GCP . Understanding of prompt engineering, hallucination reduction, output validation, and LLM evaluation. Strong understanding of software engineering practices including
Git , testing, debugging, and clean code. Ability to communicate technical solutions clearly to both technical and non-technical stakeholders. Technologies and Frameworks
Programming and AI:
Python, Retrieval-Augmented Generation (RAG), embeddings, semantic search, reranking, Hugging Face Transformers, PyTorch, LoRA, QLoRA, vLLM, TGI, Ollama Frameworks:
LangChain, LangGraph, LlamaIndex, LangSmith, Langfuse, Arize Phoenix, AutoGen, CrewAI, Semantic Kernel Observability/ML tools:
MLflow, Weights & Biases Model access:
OpenAI, Anthropic Claude, Gemini, Azure OpenAI Vector databases:
Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS, pgvector Cloud and delivery:
Docker, AWS, Azure, GCP, Kubernetes, CI/CD pipelines, REST APIs Benefits
Unlimited PTO. Very generous parental leave, much above industry standards. Entrepreneurial culture where pushing limits and taking risks is everyday business. Open communication with management and company leadership. Small, dynamic teams = massive impact. Medical, Dental and Vision coverage for employees. Access to Disability & Life insurance. Mental health and wellbeing support. Annual bonus program. Employer Stock Purchase Program (ESPP). Yearly team building experiences. Mentorship and sponsorship opportunities. Manager resources and support. Success in the Role
Production-ready AI applications that are accurate, secure, and maintainable. RAG systems that retrieve relevant information and reduce hallucinations. AI-agent workflows that reliably complete business tasks and integrate with existing systems. Measurable improvements in response quality, latency, and inference cost. Clear monitoring of application performance, usage, errors, and model behavior. Salary range:
USD 150,000 - 170,000 per year (US East/West Coast). Work location:
Hybrid remote, San Francisco Bay Area, CA.
#J-18808-Ljbffr
Design and develop scalable
LLM-powered applications
using
Python . Build
RAG pipelines
including document processing, embeddings, vector databases, semantic search, and
reranking . Develop
AI-agent and multi-agent workflows
with tool calling, memory, orchestration, and
human approval steps . Integrate LLMs with internal systems, external APIs, databases, and enterprise applications. Evaluate and select foundation models based on
accuracy ,
latency ,
cost ,
security , and business needs. Improve
prompt quality , retrieval accuracy, response time, and
token usage . Implement
safety guardrails , output validation, access controls, and fallback mechanisms. Build automated evaluation frameworks to assess response quality, hallucination, relevance, and reliability. Containerize applications with
Docker
and deploy on
AWS, Azure, or GCP . Apply
monitoring
and
LLMOps
practices for model performance, cost, latency, errors, and production usage. Collaborate with product, engineering, and business teams to translate requirements into reliable AI solutions. Document technical architecture, design decisions, APIs, and operational processes. Requirements
Bachelor's or Master's degree in Computer Science, Engineering, or a related discipline, or equivalent practical software-development experience. 6-9 years
of professional software-development experience, including strong hands-on experience with
Python . Experience developing backend services and integrating
REST APIs . Hands-on experience building
LLM or Generative AI applications . Practical experience implementing RAG using
embeddings , semantic search, and
vector databases . Experience with
LangChain ,
LangGraph ,
LlamaIndex , or a comparable LLM application framework. Experience integrating foundation models through APIs such as
OpenAI ,
Anthropic Claude ,
Gemini , or
Azure OpenAI . Working knowledge of vector databases including
Pinecone ,
Weaviate ,
Milvus ,
Qdrant ,
Chroma ,
FAISS , or
pgvector . Experience with
Docker
and deployment on at least one of:
AWS ,
Azure , or
GCP . Understanding of prompt engineering, hallucination reduction, output validation, and LLM evaluation. Strong understanding of software engineering practices including
Git , testing, debugging, and clean code. Ability to communicate technical solutions clearly to both technical and non-technical stakeholders. Technologies and Frameworks
Programming and AI:
Python, Retrieval-Augmented Generation (RAG), embeddings, semantic search, reranking, Hugging Face Transformers, PyTorch, LoRA, QLoRA, vLLM, TGI, Ollama Frameworks:
LangChain, LangGraph, LlamaIndex, LangSmith, Langfuse, Arize Phoenix, AutoGen, CrewAI, Semantic Kernel Observability/ML tools:
MLflow, Weights & Biases Model access:
OpenAI, Anthropic Claude, Gemini, Azure OpenAI Vector databases:
Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS, pgvector Cloud and delivery:
Docker, AWS, Azure, GCP, Kubernetes, CI/CD pipelines, REST APIs Benefits
Unlimited PTO. Very generous parental leave, much above industry standards. Entrepreneurial culture where pushing limits and taking risks is everyday business. Open communication with management and company leadership. Small, dynamic teams = massive impact. Medical, Dental and Vision coverage for employees. Access to Disability & Life insurance. Mental health and wellbeing support. Annual bonus program. Employer Stock Purchase Program (ESPP). Yearly team building experiences. Mentorship and sponsorship opportunities. Manager resources and support. Success in the Role
Production-ready AI applications that are accurate, secure, and maintainable. RAG systems that retrieve relevant information and reduce hallucinations. AI-agent workflows that reliably complete business tasks and integrate with existing systems. Measurable improvements in response quality, latency, and inference cost. Clear monitoring of application performance, usage, errors, and model behavior. Salary range:
USD 150,000 - 170,000 per year (US East/West Coast). Work location:
Hybrid remote, San Francisco Bay Area, CA.
#J-18808-Ljbffr