Fyerx logo

Vector Database Specialist (Pinecone / Milvus / Weaviate)

Hiring from
Probably Worldwide
Work type
Remote
Posted
Sep 23, 2026
Is this job info correct?

Job Description

This is a remote position.

Vector Database Specialist (Pinecone / Milvus / Weaviate)

Job Details
  • Employment Type: Contract
  • Work Mode: Remote
  • Location: Offshore
  • Total Experience Required: 4 to 8 years
  • Relevant Experience Required: 2+ years of dedicated data engineering experience specializing in vector database administration, architectural index design, and high-dimensional semantic search scaling
  • Mandatory Certification: Developer or Administrator certification from a major Vector DB provider (e.g., Pinecone Certified Developer, Milvus Professional) or a major cloud provider Data Engineering Specialty

Job Summary
We are seeking an experienced Vector Database Specialist to design, configure, and optimize the storage infrastructure powering our production-grade GenAI and semantic search applications. The ideal candidate will architect highly scalable vector indexes, write high-throughput embedding ingestion pipelines, configure real-time hybrid search query spaces, and maintain low-latency vector infrastructure handling millions of high-dimensional embeddings.

Key Responsibilities
  • Design, deploy, and govern production vector databases (e.g., Pinecone, Milvus, Weaviate, Qdrant, or pgvector) to manage complex long-term memory structures for LLM applications.
  • Build high-performance embedding ingestion pipelines, managing data chunking strategies, overlap controls, metadata schema extractions, and real-time upsert queues.
  • Optimize high-dimensional vector search spaces, fine-tuning approximate nearest neighbor (ANN) graph parameters, HNSW cluster metrics, IVF index lists, and scalar quantization bounds.
  • Configure advanced hybrid search architectures, engineering unified retrieval execution flows combining semantic vector lookups with traditional full-text keyword querying (BM25).
  • Implement strict metadata filtering schemas, constructing optimized filter patterns to speed up context retrieval times and enforce dynamic domain isolation safety parameters.
  • Monitor cluster metrics and resource optimization loops, tracking vector pod memory allocations, index reconstruction latencies, query-per-second (QPS) thresholds, and compute costs.
  • Collaborate with AI and Data Engineering squads to evaluate text embedding models (e.g., OpenAI, Cohere, Hugging Face) and map vector sizing requirements cleanly to downstream application runtimes.



Requirements

  • 4 to 8 years of core enterprise data engineering, database administration, or backend software development experience, with 2+ dedicated years actively scaling high-dimensional vector database frameworks.
  • Strong technical mastery of Python, advanced SQL, vector similarity distance metrics (Cosine, Euclidean, Dot Product), and data transformation engines (e.g., Spark, dbt).
  • Deep structural understanding of index types (HNSW, IVF, Flat), metadata index caching, memory footprint constraints, and cloud tenant auto-scaling mechanics.
  • Mandatory certification: Official Vector DB specialized credential or a Professional Cloud Data Engineer certificate (AWS/GCP/Azure).

Preferred Qualifications
  • Prior experience implementing real-time change data capture (CDC) architectures to automatically sync operational databases with vector catalogs.
  • Familiarity with orchestration tools like LangChain, LangGraph, or LlamaIndex to structure retrieval steps for Retrieval-Augmented Generation (RAG) pipelines.




Similar jobs

Apply for this job