KR
Salary
$20–$30/hr
USD per hour
Hiring from
Worldwide
Work type
Remote
Posted
Sep 24, 2026
Is this job info correct?

Job Title: Data Engineer


About the Role

We are seeking an experienced Data Engineer to design, build, and operate scalable data pipelines that power analytics, artificial intelligence, and machine learning initiatives across the organization. In this role, you will partner closely with data scientists, analysts, and software engineers to enable data-driven decision making through reliable, high-performance data infrastructure.


Key Responsibilities

Data Pipeline Development & Operations

  • Design, build, and maintain production-grade data pipelines using Apache Airflow for orchestration.

  • Develop and optimize large-scale data processing workflows using Apache Spark.

  • Configure and manage data storage and query engines including Apache Hive and Trino.

  • Implement data quality checks, monitoring systems, and alerting mechanisms to ensure pipeline reliability.

  • Build and maintain dashboards and visualizations using Apache Superset.


Infrastructure & Platform Management

  • Deploy, manage, and operate data infrastructure on Google Cloud Platform (GCP).

  • Containerize applications and manage deployments using Kubernetes and Docker.

  • Implement infrastructure-as-code practices for reproducible and scalable environments.

  • Optimize platform performance, resource utilization, and cloud cost efficiency.


AI/ML Integration

  • Build data pipelines to support machine learning model training and inference.

  • Implement and manage vector databases for embedding storage and similarity search.

  • Collaborate with data scientists on model deployment, monitoring, and retraining workflows.

  • Support prompt engineering initiatives and LLM integration pipelines.

  • Prepare training datasets and features for model fine-tuning and experimentation.


Collaboration & Communication

  • Partner with stakeholders to gather requirements and translate them into technical solutions.

  • Document data architectures, pipelines, and operational best practices.

  • Mentor junior engineers and contribute to team learning and development.

  • Communicate technical concepts effectively to both technical and non-technical audiences.


Required Qualifications

  • 5–10 years of professional experience in data engineering or related roles.

  • Expert-level proficiency in Python with experience building production systems.

  • Hands-on experience configuring and operating Apache Airflow.

  • Strong expertise with Apache Spark for distributed data processing.

  • Practical experience using Apache Hive and Trino for data warehousing and querying.

  • Experience working with Apache Superset or similar BI and visualization tools.

  • Working knowledge of Kubernetes for container orchestration.

  • Demonstrated experience with Google Cloud Platform (BigQuery, Cloud Storage, Dataproc, GKE, etc.).

  • Excellent written and verbal communication skills.

  • Strong problem-solving abilities and attention to detail.


Preferred Qualifications

  • Experience building data pipelines that support AI and machine learning applications.

  • Familiarity with vector databases and embedding workflows.

  • Knowledge of model deployment, fine-tuning, and MLOps practices.

  • Experience with streaming platforms such as Kafka or Pub/Sub.

  • Familiarity with dbt or similar transformation frameworks.

  • Experience implementing CI/CD practices for data pipelines.


What We’re Looking For

Beyond technical expertise, we value engineers who:

  • Take ownership and drive projects to completion.

  • Proactively identify and resolve issues before they impact users.

  • Stay current with evolving data engineering and AI/ML technologies.

  • Collaborate effectively across teams and communicate clearly.

  • Balance pragmatic delivery with engineering excellence.

  • Thrive in fast-paced environments and adapt to shifting priorities.


Technical Environment

  • Languages: Python, SQL, TypeScript

  • Data Processing: Spark, Airflow, Hive, Trino

  • Visualization: Superset

  • Cloud: Google Cloud Platform (GCP)

  • Orchestration: Kubernetes, Docker

  • Version Control: Git

  • AI/ML Tools: Vector databases, model training frameworks, LLM APIs

Location: TBD
Employment Type: Full-Time

Similar jobs

Apply for this job