IC

Senior LLMOps / AI Platform Engineer

Hiring from
United Arab Emirates
Work type
Hybrid
Posted
Sep 25, 2026
Is this job info correct?

IC Markets Global is one of the most renowned Forex CFD provider, offering trading solutions for active day traders and scalpers as well as traders that are new to the forex market. IC Markets Global offers its clients cutting edge trading platforms, low latency connectivity and superior liquidity.

IC Markets Global is revolutionizing online forex trading. Traders are now able to gain access to pricing previously only available to investment banks and high net worth individuals.

Our management team have significant experience in the Forex, CFD and Equity markets in Asia, Europe and North America. It is this experience that has enabled us to select the best possible technology solutions and hand pick some of the best pricing providers available in the market.


About the Role

We are looking for a Senior LLMOps / AI Platform Engineer to build, deploy, optimize, and operate production-grade LLM and Generative AI infrastructure.

The role combines LLM inference, GPU optimization, Kubernetes, cloud infrastructure, observability, RAG, and AI platform engineering.


Key Responsibilities

  • Deploy and operate self-hosted LLMs using vLLM, SGLang, Ollama.
  • Optimize LLM inference for latency, throughput, concurrency, GPU memory, KV cache, and cost.
  • Manage GPU workloads across multiple NVIDIA GPUs.
  • Deploy and maintain AI services on Kubernetes / AWS EKS using Docker and Helm.
  • Implement LLM reliability mechanisms including health checks, monitoring, automated recovery, and model restart/refresh strategies.
  • Implement observability using Langfuse/LangSmith, OpenTelemetry, Prometheus, and Grafana.
  • Deploy and optimize RAG systems, embedding models, and vector databases such as Qdrant, Milvus.
  • Support AI agents and workflows built with LangChain and LangGraph.
  • Build and maintain CI/CD pipelines for AI services and infrastructure.
  • Troubleshoot production issues across LLMs, GPUs, Kubernetes, networking, and AI applications.


Required Skills

  • Strong Python and FastAPI experience
  • vLLM, Hugging Face and self-hosted LLM deployment
  • Kubernetes, Docker, Helm and AWS
  • NVIDIA GPU inference and performance optimization
  • LangChain / LangGraph/LangSmith
  • RAG, embeddings and vector databases (Qdrant)
  • LLM observability and monitoring
  • PostgreSQL / Redis
  • GitHub Actions / CI/CD
  • Strong production troubleshooting skills

Similar jobs

Apply for this job