We are looking for a skilled Data Engineer to own and scale our data processing pipelines and search infrastructure. This role is ideal for someone with strong AWS and Elasticsearch expertise, a passion for clean data, and the ability to think deeply about data architecture, scalability, and distributed systems — not just implement pipelines, but design robust foundations for growth and search performance in real-world applications. Responsibilities Own and maintain our data pipelines using AWS Glue (PySpark) and S3 Process and manage large-scale datasets stored in Elasticsearch and MongoDB Build and optimize workflows for data transformation, cleaning, and normalization Improve the performance and relevance of our search system through algorithmic tuning and semantic enhancements (e.g. LLMs, DeepL, embeddings) Monitor data quality, handle schema evolution, and ensure operational stability Collaborate with backend and product teams to ensure fast, reliable, and accurate data-driven features Contribute to data architecture and system design decisions, with a focus on scalability, maintainability, and distributed processing Must-Have Skills Solid experience with AWS data engineering tools (Glue, S3, Lambda) Strong knowledge of Elasticsearch (query design, aggregations, performance tuning) Proficiency in Python, especially with PySpark and Pandas Experience with data quality, schema evolution, and pipeline monitoring Experience with Kafka or other streaming systems Strong understanding of data architecture, data modeling, and scalable system design Ability to reason about distributed data systems and build solutions that scale reliably Nice-to-Have Experience with CDC pipelines (Debezium or similar) Experience designing distributed data platforms Familiarity with LLMs, embeddings, or search ranking improvements Previous experience with B2B datasets, contact enrichment, or lead intelligence Why Join Shape the tech direction and build from the ground up Collaborate closely with the founder on product execution Work with real user feedback and fast iteration cycles Interview Process HR Interview Cultural Fit Interview Technical Interview Benefits 100% remote Competitive salary in USD PTO (vacation, sick leave, holidays) International experience Udemy trainings covered $200 for home-office equipment
Azure Data Engineer/Developer
Helveticminds
Software Engineer - IT Data Protection
Helveticminds
Data Engineer (Contract, Remote)
Empowersstaffing
AI Data Engineer (ML Data Pipelines)
Empowersstaffing
Senior Data Engineer
Transflo
Senior AI Engineer/Data Scientist
Kanini