Senior Data Engineer
- Hiring from
- Probably Worldwide
- Work type
- Remote
- Posted
Show job descriptionHide job description
Job Description
This is a remote position.
Summary:
Detailed information:
Senior Data Engineer:
Experience: 8+ years of overall Data Engineering experience, with strong hands-on experience building enterprise-scale cloud data platforms and pipelines.
Primary Skills:
- Python
- PySpark / Apache Spark
- Snowflake
- dbt (Data Build Tool)
- Apache Iceberg
- AWS Data Services
- Advanced SQL
- Data Engineering / ETL / ELT
- Data Lake / Lakehouse architecture
Secondary / Preferred Skills:
- AWS services such as:
- S3
- AWS Glue
- EMR
- Lambda
- Step Functions
- CloudWatch
- IAM
- Apache Airflow or other workflow orchestration tools
- Snowflake performance optimization and cost optimization
- Snowpipe / Snowpark
- Spark performance tuning
- Data modeling and dimensional modeling
- Parquet and other columnar data formats
- Data quality frameworks and automated validation
- CI/CD for data pipelines
- Git / GitHub / GitLab
- Infrastructure as Code such as Terraform or AWS CDK
- Docker / containerization
- Data governance, lineage, security, and access control
- Agile/Scrum delivery experience
Job Description:
We are looking for a Senior Data Engineer with strong hands-on expertise in Python, PySpark, Snowflake, dbt, Apache Iceberg, and AWS to design, develop, and maintain scalable enterprise data solutions.
The candidate should have strong experience working with high-volume data processing, cloud-based data platforms, modern lakehouse architectures, ETL/ELT pipelines, data modeling, performance optimization, and production-grade engineering practices.
The ideal candidate should be capable of independently owning complex data-engineering components, contributing to technical design and architecture decisions, troubleshooting production issues, and providing technical guidance to other engineers.
Key Responsibilities:
1. Data Pipeline Engineering
- Design, develop, test, and maintain scalable ETL/ELT data pipelines.
- Develop production-quality data-processing solutions using Python and PySpark.
- Build reusable frameworks and components for ingestion, transformation, validation, and publishing of data.
- Process large structured, semi-structured, and distributed datasets.
- Implement incremental and batch-processing patterns where appropriate.
2. Snowflake Development
- Design and develop scalable data solutions using Snowflake.
- Develop complex SQL transformations, data models, views, and reusable data structures.
- Optimize Snowflake workloads for performance, scalability, and cost.
- Implement appropriate data-loading and transformation patterns between AWS data platforms and Snowflake.
- Troubleshoot performance and data-quality issues across Snowflake workloads.
3. dbt Development
- Build and maintain transformation pipelines using dbt.
- Develop modular, reusable, maintainable dbt models.
- Implement dbt tests and documentation.
- Follow appropriate development practices for source, staging, intermediate, and business-layer transformations.
- Support automated deployment and CI/CD practices for dbt projects.
4. Apache Iceberg / Lakehouse
- Design and implement data-lake and lakehouse solutions using Apache Iceberg.
- Build scalable table structures for large analytical datasets.
- Work with partitioning, schema evolution, incremental processing, and table-maintenance strategies.
- Integrate Iceberg-based datasets with Spark and AWS-based data-processing services.
- Ensure efficient storage and query patterns for high-volume datasets.
5. AWS Data Engineering
- Design and implement cloud-native data solutions on AWS.
- Build data-processing workloads leveraging services such as S3, Glue, EMR and Lambda where appropriate.
- Implement secure access patterns using AWS IAM.
- Monitor data workloads and troubleshoot operational issues.
- Participate in designing scalable, reliable, secure, and cost-efficient cloud data architectures.
6. Performance & Scalability
- Diagnose and optimize Spark/PySpark jobs, SQL queries, Snowflake workloads, and data pipelines.
- Identify bottlenecks involving compute, storage, partitioning, data skew, transformations, and queries.
- Design solutions capable of supporting increasing data volumes without unnecessary infrastructure cost.
7. Data Quality & Governance
- Implement automated data-quality checks across ingestion and transformation layers.
- Establish proper logging, monitoring, exception handling, and reconciliation mechanisms.
- Follow organizational standards for data security, governance, lineage, and access controls.
- Ensure production pipelines are reliable, auditable, and maintainable.
8. Engineering Best Practices
- Write clean, modular, reusable, testable, and maintainable code.
- Perform code reviews and enforce engineering standards.
- Implement unit, integration, and data-validation testing.
- Use Git-based version control and CI/CD practices.
- Create and maintain appropriate technical documentation.
9. Senior-Level Responsibilities
- Independently drive technically complex data-engineering requirements from design through production deployment.
- Participate in solution design and architecture discussions.
- Evaluate alternative implementation approaches and recommend appropriate solutions.
- Troubleshoot complex production and performance issues.
- Mentor junior and mid-level data engineers.
- Collaborate with Architects, Product Owners, Business Analysts, Data Scientists, QA, DevOps, and application teams.
- Translate business/data requirements into scalable technical solutions.
- Identify technical risks and proactively recommend improvements.
Core Skills Expected
A strong candidate should demonstrate deep hands-on capability, not merely theoretical exposure, in the following areas:
|
Area
|
Expected Capability
|
|
Python
|
Advanced, production-quality data engineering development
|
|
PySpark
|
Large-scale distributed processing, optimization and troubleshooting
|
|
Snowflake
|
Development, modeling, optimization and performance tuning
|
|
dbt
|
Models, tests, macros, documentation and deployment practices
|
|
Apache Iceberg
|
Lakehouse/table design, partitioning, schema evolution and optimization
|
|
AWS
|
Hands-on cloud data platform development
|
|
SQL
|
Advanced SQL, query optimization and analytical processing
|
|
Data Engineering
|
ETL/ELT, batch/incremental pipelines, data quality and orchestration
|
|
Data Architecture
|
Data Lake, Data Warehouse and Lakehouse concepts
|
|
Engineering Practices
|
Git, testing, code reviews, CI/CD and production support
|
Preferred Qualifications
- Bachelor's or Master's degree in Computer Science, Information Technology, Engineering, or a related discipline.
- Strong experience delivering enterprise-scale cloud data platforms.
- Experience migrating legacy data workloads to modern AWS/Snowflake architectures.
- Experience working with very large datasets and distributed processing.
- Knowledge of data security and governance practices.
- Experience working in Agile delivery environments.
- AWS and/or Snowflake certification is an added advantage.