Shopify is a company of and for entrepreneurs. Our mission is to make commerce better for everyone and for that, we build the tools millions of merchants use to start, run, and grow their businesses. We are digital by default (remote‑first), with teams distributed across North America and beyond. Search is the front door to Shopify stores. Our team owns the pipelines that continuously build and refresh the product‑search index so buyers find the right items fast. We’re a small, pragmatic group that optimizes for impact: batch‑first for scale and cost, with streaming edges where real‑time really moves the needle. The work is high‑leverage and visible — improvements you make to freshness and correctness show up directly in buyer experience and merchant revenue. Own and evolve the data pipelines that power Shopify’s search index. You’ll set clear SLOs for freshness and correctness, design efficient batch flows, add near‑real‑time updates where they materially improve outcomes, and raise the bar on reliability, observability, and recovery. You’ll partner closely with ranking and serving teams to ship changes buyers and merchants feel. Senior candidates will also be considered based on scope. What you’ll do Design batch‑first flows that process billions of documents efficiently; add streaming edges where they create real product impact. Make the data trustworthy: prevent, detect, and contain bad or late data; plan and execute safe backfills and rollbacks. Build the guardrails: instrumentation, alerting, and playbooks that keep pipelines healthy during growth and change. Improve the system: identify bottlenecks, de‑risk failure modes, and deliver measurable SLI/SLA improvements. Collaborate with ranking/serving teams to land end‑to‑end wins; mentor teammates and codify best practices. Qualifications Proven ownership of production data pipelines at significant scale Strong implementation skills in Java and/or Python plus SQL; you’ve written production code beyond orchestration/UI configs. Clear judgment on batch vs streaming and where to add near‑real‑time updates (e.g., CDC/event streams) to materially improve outcomes. Kafka experience; familiarity with Flink/Flink SQL and streaming fundamentals Data correctness mindset On‑call experience, robust observability, and a track record of improving availability/freshness/correctness. Experience setting technical direction across multiple pipelines or domains Bonus points Search indexing exposure (document/feature pipelines; index build/merge mechanics). Experience in our neighborhood of tools (e.g., Spark/MapReduce, BigQuery/GCP, Iceberg/Hudi, Kafka/Flink) Deeper Flink internals (state, checkpointing/savepoints) even if batch is your primary mode.
Principal Research Scientist, Database Systems
MongoDB
Cloud Developer (Remote/LatAm)
Vergo
Staff AI Engineer
Fullpond
Staff Software Engineer
Fullpond
Oracle Forms & PL/SQL Developer
Allied Global Technology Services
Senior Infrastructure Engineer - Streaming Services/Materialization Platform
Shopify