LH

Senior Data Engineer — Azure Databricks & AWS

Logic Hire Solutions LTD
Posted 5 hours ago
PortugalHybridData & Analytics
Is this job info correct?
About The Role

Loghic Hire is seeking a Senior Data Engineer with deep, hands-on expertise in Azure Databricks and AWS to design, build, and operate modern cloud data platforms for our enterprise clients. You will own the end-to-end data lifecycle — from ingestion and transformation to serving and governance — and act as a technical anchor for data engineering across the organization.

This is a build-and-lead role. You will write production-grade PySpark / SQL / Python code, architect medallion (bronze/silver/gold) pipelines, tune Databricks clusters and Delta Lake workloads, and integrate AWS-native services where they add value. You will also mentor engineers, define standards, and partner with analytics, data science, and platform teams.

You will work in a hybrid model based in Lisbon, collaborating on-site with the team and client stakeholders while retaining flexibility for remote work.

Key Responsibilities1. Data Platform Architecture & Design

  • Design and deliver cloud-native data platforms on Azure Databricks and AWS, covering ingestion, storage, processing, serving, and governance layers.
  • Architect medallion architectures (Bronze / Silver / Gold) using Delta Lake with ACID transactions, time travel, and schema evolution.
  • Define lakehouse vs. warehouse boundaries and integrate with Azure Synapse, Snowflake, Redshift, or Athena as appropriate.
  • Design batch, micro-batch, and streaming pipelines using Spark Structured Streaming, Kafka, Event Hubs, or Kinesis.
  • Establish data modeling standards — dimensional, Data Vault, and wide-table patterns — fit for analytics and ML consumption.
  • Produce architecture decision records (ADRs), design documents, and reference architectures.
  • Evaluate and recommend build-vs-buy decisions across the Azure and AWS data ecosystems.
  • Data Ingestion & Integration
  • Build scalable ingestion frameworks from relational databases, SaaS APIs, event streams, files, and third-party sources.
  • Implement CDC (Change Data Capture) patterns using Debezium, Azure Data Factory, AWS DMS, or Databricks Lakeflow Connect.
  • Orchestrate pipelines with Apache Airflow, Azure Data Factory, Databricks Workflows, or AWS Step Functions.
  • Integrate on-premises and hybrid sources into cloud data platforms securely and reliably.
  • Own idempotency, replayability, and backfill strategies for all ingestion paths.
  • Implement data contracts between producers and consumers to prevent upstream breakage.
  • Transformation & Processing
  • Write production-grade PySpark, Spark SQL, and Python for large-scale transformations.
  • Optimize Databricks jobs and clusters — partitioning, Z-ordering, liquid clustering, photon, autoscaling, and job clusters vs. all-purpose clusters.
  • Tune Spark performance: shuffle, skew handling, broadcast joins, caching, and AQE.
  • Build reusable transformation frameworks and libraries that enforce standards across teams.
  • Implement Delta Live Tables (DLT) / Lakeflow Declarative Pipelines for declarative, testable pipelines.
  • Apply data quality checks using Great Expectations, Soda, or DLT expectations, with quarantine and alerting paths.
  • Manage schema evolution, versioning, and backward compatibility rigorously.
  • AWS & Azure Cloud Engineering
  • Provision and manage data infrastructure with Terraform (and/or Bicep / CloudFormation) as code.
  • Work hands-on with AWS services: S3, Glue, Athena, Redshift, Kinesis, Firehose, Lambda, Step Functions, EMR, IAM, KMS.
  • Work hands-on with Azure services: Data Lake Storage Gen2, Azure Data Factory, Synapse, Event Hubs, Key Vault, Entra ID, Databricks workspace administration.
  • Design cost-efficient storage and compute strategies — lifecycle policies, tiering, spot/preemptible instances, and cluster policies.
  • Implement networking and security for data platforms: private endpoints, VNet/VPC injection, NSGs/security groups, and firewalls.
  • Own IAM / RBAC / Unity Catalog access models across both clouds.
  • Data Governance, Security & Compliance
  • Implement Unity Catalog for fine-grained access control, lineage, and auditing on Databricks.
  • Integrate with AWS Lake Formation and Azure Purview / Microsoft Purview for cataloging and lineage.
  • Enforce data classification, masking, tokenization, and encryption at rest and in transit.
  • Ensure compliance with GDPR and other applicable regulations, including right-to-erasure and data residency.
  • Establish audit logging, monitoring, and alerting for sensitive data access.
  • Partner with security and compliance teams on control evidence and audits.
  • Reliability, Observability & Operations
  • Define SLAs / SLOs for data freshness, completeness, and availability.
  • Instrument pipelines with metrics, logs, and traces (OpenTelemetry, CloudWatch, Azure Monitor, Datadog, Grafana).
  • Build data observability using tools like Monte Carlo, Elementary, or custom frameworks.
  • Lead root-cause analysis for pipeline failures and drive permanent fixes.
  • Implement CI/CD for data pipelines — Git-based workflows, automated testing, and environment promotion.
  • Own disaster recovery and business continuity for critical data assets.
  • Participate in an on-call / escalation rotation for production data platforms.
  • Technical Leadership & Collaboration
  • Mentor and coach data engineers — code reviews, design reviews, and pairing.
  • Set and enforce coding, testing, and documentation standards.
  • Partner with data scientists, analysts, and product teams to translate requirements into robust pipelines.
  • Represent Loghic Hire in client architecture and delivery reviews.
  • Contribute to internal accelerators, templates, and best-practice playbooks.
  • Help interview, onboard, and grow the data engineering practice.

Tech StackCategoryTechnologiesCloud PlatformsMicrosoft Azure, AWSData PlatformsAzure Databricks, Delta Lake, Delta Live Tables (Lakeflow), Unity CatalogCompute & ProcessingApache Spark, PySpark, Spark SQL, Spark Structured Streaming, Databricks WorkflowsProgrammingPython, SQL, Scala (nice to have), BashOrchestrationApache Airflow, Azure Data Factory, Databricks Workflows, AWS Step FunctionsStreaming & MessagingApache Kafka, Azure Event Hubs, AWS Kinesis, AWS FirehoseStorageADLS Gen2, Amazon S3, Parquet, Delta, Avro, ORCAWS ServicesS3, Glue, Athena, Redshift, EMR, Lambda, DMS, IAM, KMS, Lake FormationAzure ServicesData Lake Storage Gen2, Synapse, Event Hubs, Key Vault, Entra ID, PurviewWarehouses / QuerySnowflake, Redshift, Synapse, Athena, Databricks SQLData QualityGreat Expectations, Soda, DLT Expectations, ElementaryGovernance & CatalogUnity Catalog, AWS Lake Formation, Microsoft PurviewIaC & DevOpsTerraform, Bicep, CloudFormation, GitLab CI / GitHub Actions, Azure DevOpsObservabilityCloudWatch, Azure Monitor, Datadog, Grafana, OpenTelemetrySecurityIAM, RBAC, KMS, Key Vault, encryption, masking, tokenizationCDC & IntegrationDebezium, AWS DMS, Azure Data Factory, Lakeflow ConnectRequirements

  • 10+ years in data engineering, with at least 5+ years building production data platforms on Azure Databricks and AWS.
  • Deep, hands-on expertise in Apache Spark / PySpark and Delta Lake, including performance tuning at scale.
  • Strong SQL and Python skills; experience writing production-grade, tested, maintainable code.
  • Proven experience designing and operating medallion / lakehouse architectures end to end.
  • Hands-on with AWS data services (S3, Glue, Athena, Redshift, Kinesis, IAM) and Azure data services (ADLS Gen2, ADF, Synapse, Event Hubs, Key Vault).
  • Strong orchestration experience (Airflow, ADF, Databricks Workflows, or Step Functions).
  • Experience with CI/CD for data pipelines and Infrastructure as Code (Terraform preferred).
  • Solid grasp of data governance, security, and compliance in regulated environments (GDPR, PCI, SOC 2).
  • Demonstrated technical leadership — mentoring engineers, setting standards, and driving architecture.
  • Excellent communication skills in English; able to work with technical and non-technical stakeholders.

Nice to Have

  • Databricks Certified Data Engineer Professional or Associate
  • AWS Certified Data Analytics — Specialty or AWS Solutions Architect
  • Microsoft Certified: Azure Data Engineer Associate
  • Experience with Snowflake or Microsoft Fabric
  • Exposure to MLOps / feature stores (MLflow, Feast, Databricks Feature Store)
  • Knowledge of data mesh principles and domain-oriented data ownership
  • Experience with Terraform modules for data platform provisioning
  • Familiarity with Kafka / Confluent in production
  • Prior consulting or client-facing delivery experience

Other Requirements

  • Hybrid role based in Lisbon, Portugal — expected on-site presence for team collaboration and client workshops, with flexibility for remote work.
  • Must be eligible to work in Portugal / the EU.
  • Occasional travel to client sites within Europe may be required.
  • Participation in an on-call / escalation rotation for production data platforms.
  • Background screening may be required depending on client engagement.

What We Offer

  • Competitive compensation aligned with seniority and market
  • Hybrid working model based in Lisbon
  • Opportunity to work on enterprise-scale cloud data platforms across Azure and AWS
  • Access to certifications and continuous learning
  • A collaborative, engineering-led culture with real technical ownership
  • Exposure to financial services and other regulated industries

Skills: databricks,cloud,spark,platforms,data,lake,pipelines,azure,sql,aws

Similar jobs