Senior AI Engineer I
- Hiring from
- India
- Work type
- Hybrid
- Posted
Show job descriptionHide job description
You Lead the Way. We've Got Your Back.
With the right backing, people and businesses have the power to progress in incredible ways. When you join Team Amex, you become part of a global and diverse community committed to delivering innovative customer experiences through cutting-edge technology and AI.
At American Express, you will work on next-generation Agentic AI platforms, large-scale data and AI systems, and cloud-native engineering solutions that power intelligent business processes across the enterprise. You will collaborate with world-class engineers, data scientists, architects, and product teams to design, build, deploy, and scale mission-critical AI solutions.
Join Team Amex and help us shape the future of Enterprise AI. We are seeking a highly skilled Senior AI Engineer to lead the design, development, deployment, and scaling of enterprise-grade Agentic AI solutions on Google Cloud Platform (GCP).
This role combines expertise across GenAI, LangGraph workflow orchestration, distributed data engineering, cloud-native architectures, and MLOps. The ideal candidate will build production-ready AI systems capable of processing large-scale enterprise data while ensuring reliability, governance, observability, and performance.
The engineer will play a critical role in establishing scalable AI pipelines, multi-agent architectures, retrieval systems, and intelligent workflow orchestration capabilities that drive business outcomes across the organization.
Agentic AI & LangGraph Development
- Design, develop, and deploy enterprise-scale Agentic AI applications using LangGraph, LangChain, and modern LLM frameworks.
- Build and optimize multi-agent workflows including routing, planning, reasoning, memory, validation, and execution agents.
- Develop production-grade agent orchestration frameworks capable of handling complex business processes.
- Implement workflow state management, checkpointing, human-in-the-loop controls, and failure recovery mechanisms.
- Design intelligent agent collaboration patterns including supervisor-agent, planner-executor, and hierarchical agent architectures.
- Create reusable AI skills, tools, prompts, and workflow libraries that accelerate enterprise AI adoption.
AI Platform Engineering & Deployment
- Build scalable AI services and APIs deployed on GKE (Google Kubernetes Engine).
- Design containerized AI workloads using Kubernetes, Helm, Docker, and GitOps deployment practices.
- Implement CI/CD pipelines for Agentic AI and GenAI applications.
- Optimize inference latency, throughput, resource utilization, and operational costs across AI workloads.
- Support production deployment, monitoring, observability, and incident management for AI systems.
- Design resilient and highly available AI infrastructure supporting enterprise SLAs.
Data Engineering on GCP
- Build distributed data pipelines using Apache Beam, Spark, Dataproc, Dataflow, and BigQuery.
- Design batch and streaming ingestion frameworks using Pub/Sub and event-driven architectures.
- Develop scalable ETL/ELT pipelines supporting AI and analytics workloads.
- Implement enterprise data processing frameworks for structured, semi-structured, and unstructured data.
- Build data quality, lineage, metadata, and governance capabilities across AI pipelines.
- Optimize BigQuery and storage architectures for performance and cost efficiency.
Retrieval & Knowledge Systems
- Design and implement Retrieval Augmented Generation (RAG) platforms.
- Build vector search and semantic retrieval solutions using enterprise knowledge repositories.
- Develop document ingestion, indexing, chunking, embedding, and retrieval pipelines.
- Implement hybrid search architectures combining vector, keyword, and graph-based retrieval capabilities.
- Build knowledge graphs and context-management systems supporting intelligent agents.
Scalability & Reliability Engineering
- Design AI systems capable of processing millions of records and large-scale enterprise datasets.
- Improve workflow reliability through distributed execution, workload partitioning, and fault-tolerant designs.
- Implement asynchronous processing using Pub/Sub, event-driven architectures, and worker-based execution models.
- Build performance monitoring, tracing, and observability frameworks using OpenTelemetry and cloud-native monitoring tools.
- Conduct load testing, performance tuning, and capacity planning activities.
AI Governance & Responsible AI
- Implement model governance, auditability, explainability, and compliance controls.
- Develop automated validation and confidence-scoring frameworks for GenAI outputs.
- Establish evaluation pipelines for model quality, hallucination detection, and business-rule validation.
- Support secure and compliant use of enterprise data in AI systems.
- Partner with governance and risk teams to align AI solutions with enterprise standards.
- Lead architecture reviews and provide technical direction across AI engineering initiatives.
- Mentor junior engineers and establish engineering best practices.
- Drive innovation in Agentic AI, GenAI, LLMOps, and cloud-native engineering.
- Collaborate with product, business, and enterprise architecture teams to deliver strategic AI capabilities.
- Contribute to enterprise AI platforms, reusable frameworks, and long-term technology roadmaps.
- Bachelor's or master’s degree in computer science,Engineering, or related field.
- 8+ years of software engineering experience with AI/ML or GenAI engineering.
- Strong experience building production systems in Python.
- Hands-on experience with LangGraph, LangChain, CrewAI, AutoGen, or equivalent agent orchestration frameworks.
- Experience deploying applications on Google Cloud Platform (GCP).
- Strong background in data engineering, distributed systems, and large-scale data processing.
- Experience with BigQuery, Dataflow, Dataproc, Pub/Sub, Cloud Storage, and GKE.
- Experience building RAG, Vector Search, and Knowledge Graph solutions.
- Experience with Kubernetes, Docker, CI/CD, and infrastructure automation.
- Strong understanding of software design patterns, distributed architectures, and microservices.
- Experience with PostgreSQL, Redis, and NoSQL technologies.
- Knowledge of observability, monitoring, and production support processes.