MH
Salary
$135K–$160K
USD per year
Hiring from
United States
Work type
Remote
Posted
Sep 24, 2026
Is this job info correct?

Overview

Build the Future

At McGraw Hill, we are dedicated to delivering digital learning experiences that transform education for learners and educators. Our focus is on creating seamless, impactful products that truly benefit our users while supporting growth and collaboration across teams. We foster a culture that values innovation, teamwork, and a balance between career growth and personal well-being.

How can you make an impact?

The Senior Data Engineer will play a key role in advancing McGraw Hill Education's (MHE) product eventing, data platforms, and analytics capabilities. This role focuses on designing scalable, reliable, and high-quality data solutions that transform learning interactions and product events into actionable learning insights.

Embedded within product platform teams, the engineer will collaborate closely with product managers, software engineers, and data teams to define and create meaningful events, build end-to-end event pipelines, and ensure data quality from learning interaction to learning insight.

The ideal candidate brings strong hands-on expertise in AWS, Databricks, and modern event-driven data architectures, with experience designing real-time or near-real-time event pipelines, Delta Lake, and medallion (Bronze/Silver/Gold) data architecture patterns. Proficiency in SQL and Python or Scala for large-scale data transformation is essential. Experience with Databricks Workflows, infrastructure-as-code, and CI/CD practices is a plus.

This role requires a strong focus on product data, event standardization, pipeline reliability, and end-to-end data quality. The engineer will help establish robust data foundations that enable trusted insights across MHE's learning products and platforms, supporting the organization's ability to turn learner interactions into meaningful learning outcomes.

This is a remote position open to applicants authorized to work for any employer within the United States.

What You'll Do

  • Work closely with product platform teams to define and create meaningful learning events, establish event standards, and build scalable, reliable event pipelines that connect learning interactions to actionable learning insights.
  • Demonstrate hands-on experience designing and delivering data solutions on Databricks, including Delta Lake, medallion (Bronze/Silver/Gold) architecture, Unity Catalog, and scalable lakehouse design patterns on AWS.
  • Design, develop, and optimize parallel and distributed ETL/ELT pipelines using Apache Spark (PySpark/Scala) and Databricks. Implement efficient data transformations, joins, aggregations, window functions, and other Spark-native capabilities to support high-volume product and learning event data.
  • Build and maintain real-time and near-real-time event-driven data pipelines that support the ingestion, processing, and transformation of product and learning interactions. Ensure reliable data flow from event creation through curated data layers and learning insights.
  • Establish and maintain end-to-end data quality across event schemas, ingestion pipelines, transformations, and curated data products. Implement validation, monitoring, and alerting to ensure data accuracy, consistency, completeness, and timely delivery.
  • Develop and maintain Databricks Workflows, including dependency management, retry policies, scheduling, monitoring, and alerting. Build maintainable, cloud-native pipeline orchestration solutions.
  • Apply partitioning, caching, broadcast joins, and other performance optimization techniques to improve pipeline efficiency, scalability, and throughput. Use Delta Lake capabilities to support reliable and maintainable data processing.
  • Design and implement scalable data models using Delta tables, including fact tables, dimension tables, star and snowflake schemas, and aggregations where appropriate. Ensure curated data assets support product analytics and learning insights.
  • Apply modern data architecture and governance principles using Unity Catalog, Delta Sharing, and AWS-native lakehouse patterns to support secure, discoverable, and trusted data products.
  • Use Git-based version control, including Databricks Repos/Git folders, and project management tools such as Jira within Agile/Kanban delivery frameworks.
  • Translate product and data requirements into technical designs and production-grade solutions, contributing across the full software development lifecycle—from requirements gathering and architecture through implementation, testing, deployment, and ongoing support.
  • Develop high-quality solution design documentation, including event architecture, data flow diagrams, pipeline specifications, data mappings, and Unity Catalog data asset definitions. Communicate effectively with product, engineering, analytics, and business stakeholders.

What You Bring

  • Deep expertise in modern event-driven data and lakehouse architecture, including product event design, event schemas, real-time and near-real-time event processing, Delta Lake, medallion design patterns, Unity Catalog governance, and cloud-native data solutions on Databricks. Experience connecting product-generated events to reliable data pipelines and actionable learning insights.
  • 5+ years of experience in Data Engineering, with a focus on the following tools and technologies:
  • Databricks — Delta Lake, Delta Live Tables (DLT), Databricks Workflows, Unity Catalog, Databricks SQL, and MLflow, with hands-on experience building scalable event processing and data transformation pipelines.
  • Eventing & Streaming Architecture — Experience designing and implementing event-driven architectures, event schemas, event ingestion, streaming pipelines, and real-time or near-real-time processing using technologies such as Apache Kafka, cloud-native streaming services, or equivalent eventing platforms.
  • AWS & Databricks — S3, Redshift, Glue, Lambda, EMR, Athena (with Iceberg), Step Functions, and IAM, integrated with Databricks as the primary compute and transformation layer where applicable.
  • Scripting and programming languages — Python (PySpark), Scala (Spark), or SQL as primary languages for pipeline development, event processing, and data transformation within Databricks.
  • 3+ years of experience working with cloud platforms, primarily AWS, architecting and operating Databricks environments, including workspace configuration, cluster policies, instance profiles, security, and cost optimization strategies.
  • 1+ years of experience with workflow automation and pipeline orchestration using Databricks Workflows, Apache Airflow (with the Databricks provider), or equivalent cloud-native orchestration tools, including dependency management, retries, monitoring, and alerting.
  • Experience collaborating with product platform teams to define event requirements, establish event schemas, implement event instrumentation, and build reliable pipelines that connect learning interactions to curated data products and learning insights.
  • Strong understanding of end-to-end event data quality, including schema validation, event completeness, data consistency, pipeline observability, and monitoring across event generation, ingestion, processing, and consumption.
  • Experience optimizing high-volume event processing and distributed data pipelines using Spark, Delta Lake, and cloud-native technologies to support scalability, reliability, and efficient resource utilization.

Preferred Experience & Skills:

  • Experience with EdTech domain.

Why work for us?

The work you do at McGraw Hill will be work that matters. We are collectively building experiences that will help shape the future of education. Play your part and experience a sense of fulfilment that will inspire you to even greater heights.

The pay range for this position is between $135,000 - $160,000 annually, however, base pay offered may vary depending on job-related knowledge, skills, experience, and location. An annual bonus plan may be provided as part of the compensation package, in addition to a full range of medical and/or other benefits, depending on the position offered. Click here to learn more about our benefit offerings.

McGraw Hill recruiters always use a “@mheducation.com” email address and/or from our Applicant Tracking System, iCIMS. Any variation of this email domain should be considered suspicious. Additionally, McGraw Hill recruiters and authorized representatives will never request sensitive information in email.

Similar jobs

Apply for this job