At BairesDev®, we've been leading the way in technology projects for over 15 years. We deliver cutting-edge solutions to giants like Google and the most innovative startups in Silicon Valley.
Our diverse 4,000+ team, composed of the world's Top 1% of tech talent, works remotely on roles that drive significant impact worldwide.
When you apply for this position, you're taking the first step in a process that goes beyond the ordinary. We aim to align your passions and skills with our vacancies, setting you on a path to exceptional career development and success.
Machine Learning + Data Pipeline Technical Leader at BairesDev
In this role, you'll own and redesign robust, scalable data pipeline architecture for foundation-model training data, looking at broken or inefficient pipelines and independently diagnosing what needs to change. You'll anchor design decisions for the team's data pipelines, with other engineers implementing around your direction. This is your opportunity to bring deep technical leadership to a well-defined, high-impact scope, where your ability directly determines the quality of the data that trains foundation models.
What You'll Do
Own and redesign scalable data pipeline architecture for foundation-model training data.
Diagnose and re-architect failing or inefficient pipelines end-to-end.
Design for scale and efficiency across large, heterogeneous, multimodal data.
Anchor technical design decisions and guide implementers through execution.
Establish data observability, traceability, and reliability practices as a discipline.
What We Are Looking For
8+ years of experience in machine learning engineering or data pipeline engineering.
Strong proficiency in Python for ML/AI, with depth in Ray sufficient to redesign pipelines.
Proven, independent ability to diagnose and re-architect failing or inefficient pipelines end-to-end.
Experience designing for scale and efficiency on large heterogeneous multimodal data, including partitioning and worker distribution.
Background in data observability, traceability, and reliability engineering as a discipline.
Demonstrated technical leadership, anchoring design decisions and guiding implementers.
Experience with distributed compute using Kubernetes to optimize containers for pipelines.
Advanced proficiency in English.
Nice To Have
Experience with autonomous vehicle or sensor data such as camera, lidar, or radar.
Fluency in foundation-model or VLM training loops.
Experience with Spark.
Background designing, not just implementing, eval pipelines.
Experience building data-prep frameworks that other engineers build on.
How we do make your work (and your life) easier:
100% remote work (from anywhere).
Excellent compensation in USD or your local currency if preferred
Hardware and software setup for you to work from home.
Flexible hours: create your own schedule.
Paid parental leaves, vacations, and national holidays.
Innovative and multicultural work environment: collaborate and learn from the global Top 1% of talent.
Supportive environment with mentorship, promotions, skill development, and diverse growth opportunities.
Join a global team where your unique talents can truly thrive and make a significant impact!