LH
Senior Data Engineer — Azure Databricks & AWS
Logic Hire Solutions LTDAbout The Role
Loghic Hire is seeking a Senior Data Engineer with deep, hands-on expertise in Azure Databricks and AWS to design, build, and operate modern cloud data platforms for our enterprise clients. You will own the end-to-end data lifecycle — from ingestion and transformation to serving and governance — and act as a technical anchor for data engineering across the organization.
This is a build-and-lead role. You will write production-grade PySpark / SQL / Python code, architect medallion (bronze/silver/gold) pipelines, tune Databricks clusters and Delta Lake workloads, and integrate AWS-native services where they add value. You will also mentor engineers, define standards, and partner with analytics, data science, and platform teams.
You will work in a hybrid model based in Lisbon, collaborating on-site with the team and client stakeholders while retaining flexibility for remote work.
Key Responsibilities1. Data Platform Architecture & Design
Loghic Hire is seeking a Senior Data Engineer with deep, hands-on expertise in Azure Databricks and AWS to design, build, and operate modern cloud data platforms for our enterprise clients. You will own the end-to-end data lifecycle — from ingestion and transformation to serving and governance — and act as a technical anchor for data engineering across the organization.
This is a build-and-lead role. You will write production-grade PySpark / SQL / Python code, architect medallion (bronze/silver/gold) pipelines, tune Databricks clusters and Delta Lake workloads, and integrate AWS-native services where they add value. You will also mentor engineers, define standards, and partner with analytics, data science, and platform teams.
You will work in a hybrid model based in Lisbon, collaborating on-site with the team and client stakeholders while retaining flexibility for remote work.
Key Responsibilities1. Data Platform Architecture & Design
- Design and deliver cloud-native data platforms on Azure Databricks and AWS, covering ingestion, storage, processing, serving, and governance layers.
- Architect medallion architectures (Bronze / Silver / Gold) using Delta Lake with ACID transactions, time travel, and schema evolution.
- Define lakehouse vs. warehouse boundaries and integrate with Azure Synapse, Snowflake, Redshift, or Athena as appropriate.
- Design batch, micro-batch, and streaming pipelines using Spark Structured Streaming, Kafka, Event Hubs, or Kinesis.
- Establish data modeling standards — dimensional, Data Vault, and wide-table patterns — fit for analytics and ML consumption.
- Produce architecture decision records (ADRs), design documents, and reference architectures.
- Evaluate and recommend build-vs-buy decisions across the Azure and AWS data ecosystems.
- Data Ingestion & Integration
- Build scalable ingestion frameworks from relational databases, SaaS APIs, event streams, files, and third-party sources.
- Implement CDC (Change Data Capture) patterns using Debezium, Azure Data Factory, AWS DMS, or Databricks Lakeflow Connect.
- Orchestrate pipelines with Apache Airflow, Azure Data Factory, Databricks Workflows, or AWS Step Functions.
- Integrate on-premises and hybrid sources into cloud data platforms securely and reliably.
- Own idempotency, replayability, and backfill strategies for all ingestion paths.
- Implement data contracts between producers and consumers to prevent upstream breakage.
- Transformation & Processing
- Write production-grade PySpark, Spark SQL, and Python for large-scale transformations.
- Optimize Databricks jobs and clusters — partitioning, Z-ordering, liquid clustering, photon, autoscaling, and job clusters vs. all-purpose clusters.
- Tune Spark performance: shuffle, skew handling, broadcast joins, caching, and AQE.
- Build reusable transformation frameworks and libraries that enforce standards across teams.
- Implement Delta Live Tables (DLT) / Lakeflow Declarative Pipelines for declarative, testable pipelines.
- Apply data quality checks using Great Expectations, Soda, or DLT expectations, with quarantine and alerting paths.
- Manage schema evolution, versioning, and backward compatibility rigorously.
- AWS & Azure Cloud Engineering
- Provision and manage data infrastructure with Terraform (and/or Bicep / CloudFormation) as code.
- Work hands-on with AWS services: S3, Glue, Athena, Redshift, Kinesis, Firehose, Lambda, Step Functions, EMR, IAM, KMS.
- Work hands-on with Azure services: Data Lake Storage Gen2, Azure Data Factory, Synapse, Event Hubs, Key Vault, Entra ID, Databricks workspace administration.
- Design cost-efficient storage and compute strategies — lifecycle policies, tiering, spot/preemptible instances, and cluster policies.
- Implement networking and security for data platforms: private endpoints, VNet/VPC injection, NSGs/security groups, and firewalls.
- Own IAM / RBAC / Unity Catalog access models across both clouds.
- Data Governance, Security & Compliance
- Implement Unity Catalog for fine-grained access control, lineage, and auditing on Databricks.
- Integrate with AWS Lake Formation and Azure Purview / Microsoft Purview for cataloging and lineage.
- Enforce data classification, masking, tokenization, and encryption at rest and in transit.
- Ensure compliance with GDPR and other applicable regulations, including right-to-erasure and data residency.
- Establish audit logging, monitoring, and alerting for sensitive data access.
- Partner with security and compliance teams on control evidence and audits.
- Reliability, Observability & Operations
- Define SLAs / SLOs for data freshness, completeness, and availability.
- Instrument pipelines with metrics, logs, and traces (OpenTelemetry, CloudWatch, Azure Monitor, Datadog, Grafana).
- Build data observability using tools like Monte Carlo, Elementary, or custom frameworks.
- Lead root-cause analysis for pipeline failures and drive permanent fixes.
- Implement CI/CD for data pipelines — Git-based workflows, automated testing, and environment promotion.
- Own disaster recovery and business continuity for critical data assets.
- Participate in an on-call / escalation rotation for production data platforms.
- Technical Leadership & Collaboration
- Mentor and coach data engineers — code reviews, design reviews, and pairing.
- Set and enforce coding, testing, and documentation standards.
- Partner with data scientists, analysts, and product teams to translate requirements into robust pipelines.
- Represent Loghic Hire in client architecture and delivery reviews.
- Contribute to internal accelerators, templates, and best-practice playbooks.
- Help interview, onboard, and grow the data engineering practice.
- 10+ years in data engineering, with at least 5+ years building production data platforms on Azure Databricks and AWS.
- Deep, hands-on expertise in Apache Spark / PySpark and Delta Lake, including performance tuning at scale.
- Strong SQL and Python skills; experience writing production-grade, tested, maintainable code.
- Proven experience designing and operating medallion / lakehouse architectures end to end.
- Hands-on with AWS data services (S3, Glue, Athena, Redshift, Kinesis, IAM) and Azure data services (ADLS Gen2, ADF, Synapse, Event Hubs, Key Vault).
- Strong orchestration experience (Airflow, ADF, Databricks Workflows, or Step Functions).
- Experience with CI/CD for data pipelines and Infrastructure as Code (Terraform preferred).
- Solid grasp of data governance, security, and compliance in regulated environments (GDPR, PCI, SOC 2).
- Demonstrated technical leadership — mentoring engineers, setting standards, and driving architecture.
- Excellent communication skills in English; able to work with technical and non-technical stakeholders.
- Databricks Certified Data Engineer Professional or Associate
- AWS Certified Data Analytics — Specialty or AWS Solutions Architect
- Microsoft Certified: Azure Data Engineer Associate
- Experience with Snowflake or Microsoft Fabric
- Exposure to MLOps / feature stores (MLflow, Feast, Databricks Feature Store)
- Knowledge of data mesh principles and domain-oriented data ownership
- Experience with Terraform modules for data platform provisioning
- Familiarity with Kafka / Confluent in production
- Prior consulting or client-facing delivery experience
- Hybrid role based in Lisbon, Portugal — expected on-site presence for team collaboration and client workshops, with flexibility for remote work.
- Must be eligible to work in Portugal / the EU.
- Occasional travel to client sites within Europe may be required.
- Participation in an on-call / escalation rotation for production data platforms.
- Background screening may be required depending on client engagement.
- Competitive compensation aligned with seniority and market
- Hybrid working model based in Lisbon
- Opportunity to work on enterprise-scale cloud data platforms across Azure and AWS
- Access to certifications and continuous learning
- A collaborative, engineering-led culture with real technical ownership
- Exposure to financial services and other regulated industries