Job Description This is a remote position. - Own ML/AI systems end-to-end: data pipelines, model training, serving infrastructure, monitoring, and iteration - Build LLM-powered applications with custom pipelines, prompt management, evaluation, and optimization - Implement multi-agent orchestration systems using LangGraph, CrewAI, or AutoGen for autonomous workflows - Build and optimize RAG pipelines using LlamaIndex with chunking strategies, embedding selection, re-ranking, and evaluation - Deploy and manage LLM inference infrastructure using vLLM or Ollama for on-premise sovereign deployments - Build traditional ML scoring models: churn prediction, propensity scoring, LTV estimation, next-best-action - Design and build feature pipelines using Apache Flink (streaming) and Spark (batch) for real-time and batch ML - Implement MLOps practices: model versioning, registry, drift monitoring, A/B testing, and staged rollouts - Design and implement AI operators for visual low-code canvas (LLM Gateway, RAG Pipeline, Intent Classifier) - Optimize ML inference for latency and throughput at scale (10K+ QPS) - Collaborate with Data Engineering and Platform teams to integrate ML systems with data infrastructure Requirements - 3+ years of hands-on ML/AI engineering with demonstrated end-to-end system ownership - Production experience building LLM-powered applications (not just API consumption) - Hands-on experience with agent orchestration: LangGraph, CrewAI, or AutoGen in production - Production RAG experience with evaluation metrics, hybrid search, and re-ranking strategies - Experience building ML models: churn, propensity, LTV, segmentation, recommendation systems - Hands-on experience with data pipelines: Spark for batch, Flink or Kafka Streams for real-time - Strong Python proficiency: production code structure, async, multiprocessing, profiling, optimization - Experience with vector databases at scale: OpenSearch k-NN, Qdrant, or Milvus - Production MLOps experience: MLflow, experiment tracking, model registry, drift monitoring - Real-time ML inference experience at 1,000+ QPS Good to Have: - Experience at AI-first companies or building AI/ML platforms from scratch - Telco or enterprise data platform background - Experience with LLM fine-tuning: LoRA, QLoRA, PEFT techniques - Experience with embedding models: sentence-transformers, fine-tuning for domain - Kubernetes for ML workload orchestration and GPU scheduling - Knowledge of PII detection (Presidio) and LLM guardrails (NeMo Guardrails)
AI Engineer | Python | AI Agents | Machine Learning | Full Stack | Remote
Eedgetechnology
Machine Learning Engineer
Pomelo
Forward Deployed Machine Learning Engineer
Protege
Machine Learning Engineer (Production)
Empowersstaffing
Machine Learning Engineer (Audio & LLM Stack)
Slash
Lead Machine Learning Engineer (REMOTE)
Recruitytalent