Savvycom - Software Product Development logo

Data Engineer (remote)

Hiring from
Vietnam
Work type
Remote
Posted
Is this job info correct?
Show job description

Company Description Savvycom, founded in 2009, is an AI-driven solution partner and one of Vietnam’s Top 10 Digital Technology Companies, delivering end-to-end technology services and scalable digital solutions for businesses worldwide. The company helps enterprises modernize systems, automate processes, and improve performance through AI, data, and advanced digital platforms. With over 50 AI-ready solutions and a strong global delivery record, Savvycom has supported more than 100 enterprise clients across APAC, ANZ, Japan, South Korea, and the United States. The organization operates 7 global offices, serves clients in 30+ countries, and is recognized as a Top 10 Fintech Company and a Top 100 sustainable enterprise in Vietnam. Savvycom partners with leading technology providers like IBM, Google, AWS, and Softbank and has been featured on major platforms including Forbes, TechinAsia, and TEDx.

The Data Engineer owns the data foundation that every client AI and analytics solution runs on, from ingestion through governed consumption. You will design and operate pipelines, the lakehouse, the semantic layer and data-quality controls that produce authoritative, lineage-traceable and AI-ready data, and work hand-in-hand with AI engineers on the data their agents, retrieval systems and models depend on. This is a hands-on role with real ownership of data architecture at engagement scale, and a chance to help set the data-engineering standard and reusable accelerators for a growing practice.


1. Key Responsibilities

  • Design, build and operate batch and streaming ingestion and integration pipelines from client source systems (ERP, customer, clinical and operational systems, files and APIs) into a governed analytics and AI platform.
  • Build and maintain the lakehouse or warehouse and its transformation layer, modelling data for both BI and AI consumers with tested, version-controlled transformations and clearly documented data models.
  • Design and maintain a governed semantic layer so business logic, metrics and definitions stay consistent across dashboards, analytics and AI systems.
  • Prepare and serve AI-ready data for the engineering team, including curated datasets, embeddings source data, feature inputs and retrieval corpora for RAG and model builds; agree data contracts and interfaces with AI engineers.
  • Own data quality, master data and lineage: validation rules, entity resolution and golden records where needed, and end-to-end lineage from source to report and control so outputs are trustworthy and audit-ready.
  • Embed data governance and privacy from the outset: access control, data classification, de-identification where required, and compliance with Singapore PDPA and relevant APAC cross-border transfer and data-residency rules for each engagement.
  • Take ownership of data architecture at engagement scale and help set the data-engineering standard for the practice.
  • Build reusable pipeline templates and accelerators to move the team from bespoke builds to a repeatable delivery model, and mentor junior data resources as the team grows.
  • Explain data design and technical trade-offs clearly to technical colleagues, client stakeholders and executives.

2. Requirements


Must-have:


  • Approximately 5–10 years of strong, current, hands-on data-engineering experience building and operating production data pipelines and platforms, delivered end-to-end to a governed, production standard.
  • Expert Python and SQL.
  • Production experience with a transformation and modelling framework, particularly dbt, with disciplined, tested, version-controlled transformations.
  • Hands-on delivery experience with Snowflake and/or Databricks, plus at least one major cloud platform (Azure, AWS or GCP) and its data services.
  • Production experience with an orchestration tool such as Dagster, Apache Airflow or dbt Cloud, and experience with batch and streaming ingestion (understanding of Kafka for event transport vs Flink for stream processing) and change-data-capture.
  • Experience with data modelling for analytics (dimensional and/or data vault) and with data-quality and validation frameworks (GX Core / Great Expectations, dbt tests or comparable).
  • Client-facing or executive-stakeholder delivery experience, able to explain design trade-offs to technical and non-technical audiences.
  • Working awareness of Singapore PDPA and how cross-border transfer requirements affect data architecture.
  • Comfortable working in a small, fast-scaling team with a high degree of ownership; professional English communication.

Nice-to-have:

  • Open table and lakehouse formats (Apache Iceberg, Delta Lake) and semantic-layer tools (Cube, dbt Semantic Layer, LookML).
  • Master-data and entity-resolution techniques; data catalogue, lineage and data-contract tooling (Atlan, Collibra, OpenMetadata).
  • Understanding of how data feeds AI systems: embeddings, vector stores (pgvector, Pinecone, Weaviate, Qdrant), chunking and retrieval-corpus preparation for RAG; awareness of ontologies and knowledge graphs (RDF, SPARQL, Neo4j).
  • Infrastructure-as-code (e.g. Terraform) and cost-aware platform design.
  • Awareness of regional data-protection regimes (Japan APPI, Korea PIPA, India DPDP Act, China PIPL/DSL, Hong Kong PDPO, Australia Privacy Act, and Thailand, Malaysia, Indonesia, Philippines) and how they differ on cross-border transfer.
  • Healthcare and life-sciences data experience (clinical, claims, real-world evidence): FHIR, HL7 v2, SNOMED CT, ICD-10/11, LOINC, RxNorm, MedDRA, OMOP CDM, DICOM.
  • Office of the CFO data experience (financial close and reporting, consolidation, controls, risk data) and experience building governed, audit-ready data foundations or contributing reusable accelerators to a growing practice.


Similar jobs

Apply on LinkedIn