Cogniify logo

Senior Generative AI Engineer - USA

Cogniify
Posted 1 hour ago
Visa sponsorship
United States
$150K–$170KEngineering & Development
Is this job info correct?

The Role

We are looking for a hands-on Generative AI and LLM Engineer to design, build and deploy production-grade AI applications. The role will focus on developing LLM-powered products, Retrieval-Augmented Generation (RAG) pipelines and intelligent agent workflows using Python, modern AI frameworks and cloud infrastructure.

You will own solutions from requirement understanding and architecture through development, deployment and production monitoring. The ideal candidate combines strong Python backend development experience with practical knowledge of LLMs, vector databases, API integrations, Docker and at least one major cloud platform.

What You Will Do

  • Design and develop scalable LLM-powered applications using Python.

  • Build RAG pipelines using document processing, embeddings, vector databases, semantic search and reranking.

  • Develop AI-agent and multi-agent workflows with tool calling, memory, orchestration and human approval steps.

  • Integrate LLMs with internal systems, external APIs, databases and enterprise applications.

  • Evaluate and select suitable foundation models based on accuracy, latency, cost, security and business requirements.

  • Improve prompt quality, retrieval accuracy, response time and token usage.

  • Implement safety guardrails, output validation, access controls and fallback mechanisms.

  • Build automated evaluation frameworks to measure response quality, hallucination, relevance and reliability.

  • Containerize applications using Docker and deploy them on AWS, Azure or GCP.

  • Implement monitoring and LLMOps practices for model performance, cost, latency, errors and production usage.

  • Collaborate with product, engineering and business teams to convert requirements into reliable AI solutions.

  • Document technical architecture, design decisions, APIs and operational processes.

What Success Looks Like

  • Production-ready AI applications that are accurate, secure and maintainable.

  • RAG systems that retrieve relevant information and reduce hallucinations.

  • AI-agent workflows that reliably complete business tasks and integrate with existing systems.

  • Measurable improvements in response quality, latency and inference cost.

  • Clear monitoring of application performance, usage, errors and model behaviour.

What We're Looking For

  • Bachelor's or Master's degree in Computer Science, Engineering or a related discipline, or equivalent practical software-development experience.

  • 6-9 years of professional software-development experience, including strong hands-on experience with Python.

  • Experience developing backend services and integrating REST APIs.

  • Hands-on experience building LLM or Generative AI applications.

  • Practical experience implementing RAG using embeddings, semantic search and vector databases.

  • Experience with LangChain, LangGraph, LlamaIndex or a comparable LLM application framework.

  • Experience integrating foundation models through APIs such as OpenAI, Anthropic Claude, Gemini or Azure OpenAI.

  • Working knowledge of vector databases such as Pinecone, Weaviate, Milvus, Qdrant, Chroma, FAISS or pgvector.

  • Experience with Docker and deployment on at least one cloud platform: AWS, Azure or GCP.

  • Understanding of prompt engineering, hallucination reduction, output validation and LLM evaluation.

  • Strong understanding of software engineering practices, Git, testing, debugging and clean code.

  • Ability to communicate technical solutions clearly to both technical and non-technical stakeholders.

Technical Frameworks and Toolkit

  • Experience building agentic or multi-agent workflows using LangGraph, AutoGen, CrewAI, Semantic Kernel or similar frameworks.

  • Experience with Hugging Face Transformers, PyTorch or fine-tuning techniques such as LoRA or QLoRA.

  • Knowledge of model serving and inference frameworks such as vLLM, TGI or Ollama.

  • Experience with Kubernetes, CI/CD pipelines and infrastructure automation.

  • Familiarity with observability or LLMOps tools such as LangSmith, Langfuse, Arize Phoenix, MLflow or Weights & Biases.

  • Knowledge of reranking, hybrid search, chunking strategies and retrieval evaluation.

  • Experience implementing AI guardrails, PII protection, prompt-injection prevention and responsible AI practices.

Salary Range: US East/West Coast: $150000 - $170000

Disclaimer: The base salary range is a guideline and may vary based on factors such as candidate experience, specialized skills, and geographical location. Actual compensation may include additional benefits and bonuses.

Perks And Benefits of Working With Us

  • Unlimited PTO.

  • Please ask us about our very generous parental leave, much above industry standards!

  • Entrepreneurial culture where pushing limits and taking risks is everyday business.

  • Open communication with management and company leadership.

  • Small, dynamic teams = massive impact.

  • Medical, Dental and Vision coverage for employees.

  • Access to Disability & Life insurance.

  • Mental health and wellbeing support.

  • Annual bonus program.

  • Employer Stock Purchase Program (ESPP).

  • Yearly team building experiences.

  • Mentorship and sponsorship opportunities.

  • Manager resources and support.

Cogniify is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, or any other protected characteristic

Similar jobs