Q
AI Engineer
Quadcode
3 hours ago
Full-time
On-site
East New York, New York, United States
Quadcode is seeking an AI Engineer to join its AI Platform team. The role focuses on designing, building, and maintaining production generative AI pipelines and LLM-powered features, frequently involving financial and trading data.
Key Responsibilities
Design, build, and maintain LLM-powered features and pipelines for generative AI products, including systems that process financial and trading content
Develop and improve agentic systems and multi-agent workflows, including integrations via MCP (Model Context Protocol) servers and tooling
Implement production-grade backend services in Python on top of LLM provider APIs such as OpenAI, Anthropic, and xAI, and/or on-premise LLM deployments
Build and iterate on prompt engineering and prompt-evaluation pipelines, including structured output generation and validation
Convert prototypes and proofs of concept into reliable, maintainable production services
Monitor and manage trade-offs across LLM cost, latency, and reliability for end-to-end pipelines
Work closely with a small, senior AI Platform team on architecture, code review, and technical direction
Requirements
Production LLM experience : shipped real, live features using OpenAI and/or Anthropic APIs, or with on-premise/open-weight LLMs, not limited to experiments or side projects
Multi-agent workflow experience : hands-on development of multi-agent systems or workflows and work with MCP servers
Strong Python and software engineering foundations
(2 to 4+ years), including REST APIs, SQL, Docker, git, and concurrency, plus demonstrated ability to move from prototype to maintainable production service
Practical LLM engineering skills
including structured outputs, prompt evaluation, and cost-aware design decisions in real environments
English fluency
in both written and verbal communication
Technologies
Python
OpenAI, Anthropic, xAI
MCP (Model Context Protocol)
REST APIs, SQL, Docker, git
Open-weight LLMs, on-premise LLMs
Kubernetes, Kubeflow, Airflow
Hugging Face inference
AWS, GCP
Location and Experience Location:
Georgia (hybrid).
Minimum experience:
2 years.
Benefits
Hybrid work model in a brand-new office in Limassol
Health insurance and mental health services
13th salary and 21 vacation days per year
Sick leave without medical certificate: 3 days per quarter
Catered lunches in the office
Tuition reimbursement (kindergartens/schools)
Onsite Gym
Corporate events & workshops
Bonuses for special events (e.g., child's birth)
Birthday and anniversary gifts
Company-provided laptop
Corporate AI subscriptions (Claude, Gemini, GPT, etc.)
Access to a rewards marketplace with products and language courses, redeemable using the company’s internal currency
Nice to Have
Prior experience working in fintech
Finance and markets knowledge: tickers, earnings, technical indicators, forex/CFDs
Quant background
Experience with model fine-tuning/training, RAG, or LLM optimization techniques
Track record in Kaggle competitions focused on LLMs or multi-agent frameworks
GPU infrastructure experience, including provisioning GPU compute or serving open-weight models (e.g., Hugging Face inference)
Experience with Kubernetes, Kubeflow, or Airflow
Data analysis and statistics experience
Experience with cloud platforms (AWS, GCP, etc.)
Fluent Russian
#J-18808-Ljbffr
Key Responsibilities
Design, build, and maintain LLM-powered features and pipelines for generative AI products, including systems that process financial and trading content
Develop and improve agentic systems and multi-agent workflows, including integrations via MCP (Model Context Protocol) servers and tooling
Implement production-grade backend services in Python on top of LLM provider APIs such as OpenAI, Anthropic, and xAI, and/or on-premise LLM deployments
Build and iterate on prompt engineering and prompt-evaluation pipelines, including structured output generation and validation
Convert prototypes and proofs of concept into reliable, maintainable production services
Monitor and manage trade-offs across LLM cost, latency, and reliability for end-to-end pipelines
Work closely with a small, senior AI Platform team on architecture, code review, and technical direction
Requirements
Production LLM experience : shipped real, live features using OpenAI and/or Anthropic APIs, or with on-premise/open-weight LLMs, not limited to experiments or side projects
Multi-agent workflow experience : hands-on development of multi-agent systems or workflows and work with MCP servers
Strong Python and software engineering foundations
(2 to 4+ years), including REST APIs, SQL, Docker, git, and concurrency, plus demonstrated ability to move from prototype to maintainable production service
Practical LLM engineering skills
including structured outputs, prompt evaluation, and cost-aware design decisions in real environments
English fluency
in both written and verbal communication
Technologies
Python
OpenAI, Anthropic, xAI
MCP (Model Context Protocol)
REST APIs, SQL, Docker, git
Open-weight LLMs, on-premise LLMs
Kubernetes, Kubeflow, Airflow
Hugging Face inference
AWS, GCP
Location and Experience Location:
Georgia (hybrid).
Minimum experience:
2 years.
Benefits
Hybrid work model in a brand-new office in Limassol
Health insurance and mental health services
13th salary and 21 vacation days per year
Sick leave without medical certificate: 3 days per quarter
Catered lunches in the office
Tuition reimbursement (kindergartens/schools)
Onsite Gym
Corporate events & workshops
Bonuses for special events (e.g., child's birth)
Birthday and anniversary gifts
Company-provided laptop
Corporate AI subscriptions (Claude, Gemini, GPT, etc.)
Access to a rewards marketplace with products and language courses, redeemable using the company’s internal currency
Nice to Have
Prior experience working in fintech
Finance and markets knowledge: tickers, earnings, technical indicators, forex/CFDs
Quant background
Experience with model fine-tuning/training, RAG, or LLM optimization techniques
Track record in Kaggle competitions focused on LLMs or multi-agent frameworks
GPU infrastructure experience, including provisioning GPU compute or serving open-weight models (e.g., Hugging Face inference)
Experience with Kubernetes, Kubeflow, or Airflow
Data analysis and statistics experience
Experience with cloud platforms (AWS, GCP, etc.)
Fluent Russian
#J-18808-Ljbffr