Data Platform Engineer
- Hiring from
- United States
- Work type
- Hybrid
- Posted
- Sep 29, 2026
This is Hybrid role (3 days in office /2 days remote)
About your Team:
The Enterprise Architecture organization is looking for a Senior Data Platform Engineer to join the team and help build the next generation of data infrastructure and AI-enabled workflows. In this role, you will design, build, and operate the scalable data lake platform, cloud infrastructure, and self-service tooling that internal teams across IBKR rely on to ingest, govern, and analyze data at scale. You will partner closely with internal development teams and IT leadership to architect platform solutions for diverse use cases across the organization, ranging from cloud-native data lakes and infrastructure as code to AI knowledge bases and advanced analytics. Your work will focus on delivering robust, well-documented data platform capabilities and establishing best practices for data engineering and platform operations across the enterprise.
What will be your responsibilities within IBKR:
- Design, build, and operate the data lake platform and infrastructure, including S3-based storage and the AWS Glue data catalog, delivered as a governed, self-service platform that internal teams can build on
- Design and implement data governance and fine-grained access control across the platform using AWS Lake Formation, ensuring secure, least-privilege, and compliant access to data domains
- Automate infrastructure provisioning as code (e.g., Terraform), giving internal teams a repeatable, self-service path to stand up new platform resources
- Contribute to the GenAI knowledge base platform and Python-based services that internal teams use to build retrieval-augmented (RAG) and knowledge-driven AI applications
- Build or contribute to real-time streaming data pipelines (e.g., using Kafka) for event-driven analytics and data processing, as platform needs evolve
- Build and maintain scalable ETL/ELT pipelines and data crawlers to ingest data from various sources, transforming structured, semi-structured, and unstructured data (text, images, audio, video) for analytics and AI/ML workloads
- Monitor, troubleshoot, and continuously optimize the platform for performance, reliability, data quality, and cost efficiency, treating observability as a core platform capability
- Collaborate with internal development teams, data scientists, and stakeholders to understand requirements, architect platform solutions, and enable teams through clear standards and technical documentation
- Write clean, maintainable, well-tested code following software engineering and data engineering best practices
What required skill’s you need:
Core Data Platform & Cloud Engineering
- 5+ years of hands-on data engineering or data platform experience, including building infrastructure, platforms, or tooling used by other engineering teams
- Strong experience with AWS cloud services, particularly S3, AWS Glue, Athena, and EMR
- Experience with data governance and fine-grained access control, including AWS Lake Formation
- Experience with infrastructure as code (e.g., Terraform) to build and manage cloud infrastructure and governance in a versioned, repeatable way
- Solid understanding of data lake architectures, cataloging, and partitioning strategies
- Proficiency in Python, with experience building and maintaining scalable ETL/ELT pipelines for data processing and automation
- Experience with PySpark on EMR for large-scale data processing
- Strong SQL skills for data analysis and transformation
- Experience with CI/CD practices, version control (Git), and containerization (Docker)
AI/GenAI Platform
- Foundational understanding of GenAI concepts such as RAG architectures, vector databases, or semantic search, with willingness to build deeper hands-on expertise
- Familiarity with AWS AI/ML services (e.g., Bedrock, SageMaker) or equivalent tools
Good to have
- Experience with modern table formats (e.g., Iceberg) and schema evolution at scale
- Experience with Kafka or other streaming platforms for real-time data pipelines
- Experience with Kubernetes/EKS and workflow orchestration tools (e.g., Airflow, AWS Step Functions, Managed Airflow)
- Experience with observability engineering (e.g., CloudWatch, Elastic/ELK stack)
- Experience with data quality frameworks and tooling (e.g., Glue Data Quality)
- Background in building knowledge management systems or enterprise search solutions
- Experience in financial services or other regulated data environments
To be successful in this position, you will have the following:
- Self-motivated and able to handle tasks with minimal supervision.
- Superb analytical and problem-solving skills.
- Excellent collaboration and communication (Verbal and written) skills.
- Outstanding organizational and time management skills.
Company Benefits & Perks
- Competitive salary, annual performance-based bonus, and stock grant awards
- 401(k) retirement plan with competitive company match
- Excellent health and wellness benefits, including medical, dental, and vision benefits. 100% employer-paid medical premiums, with generous employer contributions to dental & vision plans as well.
- Wellness screening and assessments, health coaches, and counseling services through an Employee Assistance Program (EAP)
- Generous paid parental leave (up to 16 weeks paid parental leave for eligible employees)
- Company-paid basic life insurance, accidental death & dismemberment (AD&D), and short- and long-term disability coverage
- Flexible Spending Accounts (Healthcare, Dependent Care, and Commuter FSAs)
- Quarterly fitness stipend to offset costs associated with traditional gym and fitness memberships or fees
- Education reimbursement and professional development opportunities
- Legal services, telehealth access, and voluntary insurance options
- Backup child and adult care support through Care.com
- Daily lunch allowance and fully stocked kitchen with healthy breakfast and snack options