About Us Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one. Responsibilities Data Processing at Scale: Build and maintain high-throughput image and video data pipelines—cleaning, filtering, augmenting, and transforming multimodal datasets Data Strategy from Product Goals: Translate product objectives and training strategies into concrete data processing plans, defining required dataset characteristics, sourcing strategies, and preparation protocols for model consumption. Iterative Quality Evaluation: Evaluate processed datasets against rigorous quality benchmarks, diagnose data-side failure modes, and iterate on processing strategies until datasets meet the high bar our foundation models demand. Large-Scale Labeling Coordination: Coordinate and drive data labeling efforts with annotation teams, establishing labeling guidelines, reviewing outputs, and ensuring consistency across large-scale annotation campaigns. Requirements You have strong Python programming skills and write clean, modular, production-grade data processing code. You have hands-on experience with ML frameworks such as PyTorch—understanding model data requirements, tensor formats, and training data flow well enough to prepare data that researchers can consume directly. You hold exceptionally high standards for data quality and can assess data reliability, accuracy, and coverage systematically. You are familiar with Computer Vision domain concepts and understand how image and video data characteristics impact downstream generative model performance. You communicate clearly, document your work thoroughly, and collaborate effectively with researchers, engineers, and labeling partners across time zones. Production Quality, Agent Velocity: Your daily workflow runs through AI coding harnesses (e.g., AI agents/assistants), without sacrificing software engineering rigor. You review agent code diffs with the same scrutiny as a team member's PR, recognize AI code generation failure modes, and ship rapidly without introducing technical debt or "slop." Nice to Have Experience with Computer Vision tasks related to image or video generation model training (e.g., diffusion models, autoregressive transformers, GANs). Fluency with annotation formats such as COCO, PASCAL VOC, or custom labeling schemas. Hands-on experience with data orchestration frameworks (e.g., Airflow, Dagster, Prefect, Luigi). Experience with distributed data processing systems (e.g., Ray, Spark, Dask).
Coordinator of Financial Planning & Services
City of Winnipeg
Streets Project Coordinator
City of Winnipeg
French Bilingual Client Service Specialist (24 months)
Legal Aid Ontario
Inside Sales Representative (Bilingual)
LV1 Health Tech
Project Manager
Campus4Tech
Technical Sales Engineer
Engineered Intelligence Inc.