SRM Tech logo

Senior Data Engineer

Hiring from
Probably Worldwide
Work type
Remote
Posted
Is this job info correct?
Show job description

Job Description

This is a remote position.

Summary:

Role: Senior Data Engineer
Experience: 8+ Years
Mandatory/Core: Python, PySpark, Snowflake, dbt, Apache Iceberg, AWS, SQL
Preferred: AWS Glue, S3, EMR, Lambda, Airflow, Snowpipe/Snowpark, CI/CD, Terraform, Data Modeling
Role Type: Senior hands-on Data Engineer
Focus: Cloud Data Engineering, Lakehouse, Data Transformation, Performance Optimization and Production Engineering


Detailed information:

Senior Data Engineer:

Experience: 8+ years of overall Data Engineering experience, with strong hands-on experience building enterprise-scale cloud data platforms and pipelines.

Primary Skills:

  • Python
  • PySpark / Apache Spark
  • Snowflake
  • dbt (Data Build Tool)
  • Apache Iceberg
  • AWS Data Services
  • Advanced SQL
  • Data Engineering / ETL / ELT
  • Data Lake / Lakehouse architecture

Secondary / Preferred Skills:

  • AWS services such as:
    • S3
    • AWS Glue
    • EMR
    • Lambda
    • Step Functions
    • CloudWatch
    • IAM
  • Apache Airflow or other workflow orchestration tools
  • Snowflake performance optimization and cost optimization
  • Snowpipe / Snowpark
  • Spark performance tuning
  • Data modeling and dimensional modeling
  • Parquet and other columnar data formats
  • Data quality frameworks and automated validation
  • CI/CD for data pipelines
  • Git / GitHub / GitLab
  • Infrastructure as Code such as Terraform or AWS CDK
  • Docker / containerization
  • Data governance, lineage, security, and access control
  • Agile/Scrum delivery experience

Job Description:

We are looking for a Senior Data Engineer with strong hands-on expertise in Python, PySpark, Snowflake, dbt, Apache Iceberg, and AWS to design, develop, and maintain scalable enterprise data solutions.

The candidate should have strong experience working with high-volume data processing, cloud-based data platforms, modern lakehouse architectures, ETL/ELT pipelines, data modeling, performance optimization, and production-grade engineering practices.

The ideal candidate should be capable of independently owning complex data-engineering components, contributing to technical design and architecture decisions, troubleshooting production issues, and providing technical guidance to other engineers.


Key Responsibilities:


1. Data Pipeline Engineering

  • Design, develop, test, and maintain scalable ETL/ELT data pipelines.
  • Develop production-quality data-processing solutions using Python and PySpark.
  • Build reusable frameworks and components for ingestion, transformation, validation, and publishing of data.
  • Process large structured, semi-structured, and distributed datasets.
  • Implement incremental and batch-processing patterns where appropriate.

2. Snowflake Development

  • Design and develop scalable data solutions using Snowflake.
  • Develop complex SQL transformations, data models, views, and reusable data structures.
  • Optimize Snowflake workloads for performance, scalability, and cost.
  • Implement appropriate data-loading and transformation patterns between AWS data platforms and Snowflake.
  • Troubleshoot performance and data-quality issues across Snowflake workloads.

3. dbt Development

  • Build and maintain transformation pipelines using dbt.
  • Develop modular, reusable, maintainable dbt models.
  • Implement dbt tests and documentation.
  • Follow appropriate development practices for source, staging, intermediate, and business-layer transformations.
  • Support automated deployment and CI/CD practices for dbt projects.

4. Apache Iceberg / Lakehouse

  • Design and implement data-lake and lakehouse solutions using Apache Iceberg.
  • Build scalable table structures for large analytical datasets.
  • Work with partitioning, schema evolution, incremental processing, and table-maintenance strategies.
  • Integrate Iceberg-based datasets with Spark and AWS-based data-processing services.
  • Ensure efficient storage and query patterns for high-volume datasets.

5. AWS Data Engineering

  • Design and implement cloud-native data solutions on AWS.
  • Build data-processing workloads leveraging services such as S3, Glue, EMR and Lambda where appropriate.
  • Implement secure access patterns using AWS IAM.
  • Monitor data workloads and troubleshoot operational issues.
  • Participate in designing scalable, reliable, secure, and cost-efficient cloud data architectures.

6. Performance & Scalability

  • Diagnose and optimize Spark/PySpark jobs, SQL queries, Snowflake workloads, and data pipelines.
  • Identify bottlenecks involving compute, storage, partitioning, data skew, transformations, and queries.
  • Design solutions capable of supporting increasing data volumes without unnecessary infrastructure cost.

7. Data Quality & Governance

  • Implement automated data-quality checks across ingestion and transformation layers.
  • Establish proper logging, monitoring, exception handling, and reconciliation mechanisms.
  • Follow organizational standards for data security, governance, lineage, and access controls.
  • Ensure production pipelines are reliable, auditable, and maintainable.

8. Engineering Best Practices

  • Write clean, modular, reusable, testable, and maintainable code.
  • Perform code reviews and enforce engineering standards.
  • Implement unit, integration, and data-validation testing.
  • Use Git-based version control and CI/CD practices.
  • Create and maintain appropriate technical documentation.

9. Senior-Level Responsibilities

  • Independently drive technically complex data-engineering requirements from design through production deployment.
  • Participate in solution design and architecture discussions.
  • Evaluate alternative implementation approaches and recommend appropriate solutions.
  • Troubleshoot complex production and performance issues.
  • Mentor junior and mid-level data engineers.
  • Collaborate with Architects, Product Owners, Business Analysts, Data Scientists, QA, DevOps, and application teams.
  • Translate business/data requirements into scalable technical solutions.
  • Identify technical risks and proactively recommend improvements.

Core Skills Expected

A strong candidate should demonstrate deep hands-on capability, not merely theoretical exposure, in the following areas:

Area

Expected Capability

Python

Advanced, production-quality data engineering development

PySpark

Large-scale distributed processing, optimization and troubleshooting

Snowflake

Development, modeling, optimization and performance tuning

dbt

Models, tests, macros, documentation and deployment practices

Apache Iceberg

Lakehouse/table design, partitioning, schema evolution and optimization

AWS

Hands-on cloud data platform development

SQL

Advanced SQL, query optimization and analytical processing

Data Engineering

ETL/ELT, batch/incremental pipelines, data quality and orchestration

Data Architecture

Data Lake, Data Warehouse and Lakehouse concepts

Engineering Practices

Git, testing, code reviews, CI/CD and production support

Preferred Qualifications

  • Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related discipline.
  • Strong experience delivering enterprise-scale cloud data platforms.
  • Experience migrating legacy data workloads to modern AWS/Snowflake architectures.
  • Experience working with very large datasets and distributed processing.
  • Knowledge of data security and governance practices.
  • Experience working in Agile delivery environments.
  • AWS and/or Snowflake certification is an added advantage.


Similar jobs

Apply for this job