Senior Data Engineer (Python / OOP & Modern Data Architecture)
PenbrothersAbout the Role
Are you a Software Engineer at heart who loves solving complex data infrastructure problems? Do you treat data pipelines as production software complete with clean object-oriented architecture, modular design patterns, robust unit testing, and abstract base classes?
We are looking for a Data Engineer to help us build and scale our custom analytics and visualization platform supporting educational leaders, school districts, and non-profits. You won’t just be writing standard SQL transformations or top-to-bottom Pandas scripts; you will architect configurable, extendable, and multi-tenant ETL frameworks that power data-informed decision-making across dozens of organizations
Work Set-up: 100% Work from home
Work Schedule: Night Shift - 9:00 PM to 6:00 AM PH time (aligned with US Business Hours/US Eastern Time)
What You’ll Do (Responsibilities)
- Architect Modular Data Pipelines: Design code-first, object-oriented Python frameworks using SOLID principles to ingest, clean, and standardize high-volume data sets across multiple diverse client schemas.
- Leverage Modern In-Memory Engines: Build high-performance data transformation workflows using Polars and Pandas.
- Manage BigQuery Infrastructure: Optimize columnar warehouse models, partitioning schemes, and query performance in Google BigQuery.
- Build Configurable Frameworks: Develop YAML/JSON configuration-driven pipelines that allow new client onboardings without rewriting core ETL code.
- Enforce Software Quality: Write unit tests using pytest, mock external data connections, establish structured logging/alerting (loguru/custom exceptions), and maintain immaculate version control and documentation.
What We’re Looking For (Technical Skills and Experience)
- Minimum of 4+ years in Data Engineering with Python and SQL as your primary toolset.
- Strong Object-Oriented Background: Proven experience applying OOP concepts (custom classes, inheritance, abstraction, design patterns like Factory/Strategy) to data pipelines.
- Hands-On Polars & Pandas Expertise: Deep experience manipulating and transforming large data frames efficiently.
- At least 2+ years with BigQuery (or equivalent columnar databases like Snowflake/Redshift).
- Software Hygiene: Production experience with pytest, Git workflows, CI/CD, type hinting, and structured exception handling.
Soft Skills
- Ability to work in a fully remote environment (Slack, Zoom)
- Works well with internal stakeholders and can translate needs into technical solutions
- Excited to work in a collaborative team environment with a flat and flexible organizational structure
- Works effectively with diverse stakeholders including school and district leaders
Nice to Have
- EdTech and/or student data experience
- Familiarity with Power BI (our analysts use this extensively)
- Experience working in a Linux environment and with Docker images/containers