Databricks Data Engineer
VtechsolutionsJob Description
This is a remote position.
Location: Remote – United States
Our client, a growing technology services organization supporting federal government programs, is seeking an experienced Databricks Data Engineer to support a federal technology initiative.
This is a remote opportunity available to U.S. citizens residing in the United States. Candidates must currently hold an ACTIVE Secret security clearance or higher. Candidates without an active Secret-level clearance cannot be considered.
The ideal candidate will bring strong hands-on data engineering experience within Databricks, including development of production-grade batch and streaming pipelines, PySpark and SQL transformations, data modeling, ingestion frameworks, and modern data architecture practices.
Responsibilities
- Design, build, and maintain batch and streaming data pipelines using PySpark, SQL, Databricks Workflows, and Delta Live Tables.
- Implement Medallion Architecture across bronze, silver, and gold data layers to support data quality, transformation, and consumption.
- Develop scalable ingestion frameworks for structured, semi-structured, and unstructured data.
- Integrate data from files, databases, APIs, and streaming sources such as Kafka, Kinesis, and Databricks Auto Loader.
- Design dimensional and domain-specific data models supporting analytics and downstream applications.
- Build and consume APIs for integration with downstream systems.
- Optimize Spark workloads for performance and cost through partitioning, caching, cluster sizing, and related techniques.
- Develop and implement data quality checks, validation processes, and pipeline monitoring.
- Maintain documentation covering data flows, lineage, architecture, and integration points.
- Collaborate with engineering, platform, architecture, and federal program stakeholders throughout the development lifecycle.
Requirements
- Active Secret security clearance or higher is required.
- 5+ years of professional data engineering experience.
- 2+ years of hands-on Databricks experience.
- Strong production-level experience with PySpark and SQL.
- Experience developing both batch and streaming data pipelines.
- Strong understanding of data modeling and modern data architecture principles.
- Proficiency with Python.
- Experience working within Git-based development and CI/CD environments.
- Familiarity with Databricks Unity Catalog, including catalogs, schemas, and permissions from a data engineering perspective.
- Experience integrating data from multiple source types, including databases, APIs, files, and streaming platforms.
Preferred Qualifications
- Experience with Databricks Delta Live Tables.
- Experience with Databricks Workflows.
- Experience with Kafka, Kinesis, or similar streaming technologies.
- Experience working within federal, regulated, or security-sensitive environments.
- Bachelor's degree in Computer Science, Engineering, Information Technology, or a related discipline.