IG

Senior Computer Vision & Multimodal AI Engineer

Hiring from
United Arab Emirates
Work type
Remote
Posted
Oct 2, 2026
Is this job info correct?

Role: Senior Computer Vision & Multimodal AI Engineer


Experience Level: 4+ Years


Employment Type: Full-Time (Paid Position)


Location: Dubai


We’re Hiring: Senior Computer Vision & VLM Engineer (4+ Years Experience)


Are you passionate about bridging the gap between classic computer vision and frontier Vision-Language Models (VLMs)? We are looking for an experienced Senior Computer Vision Engineer to join our team and build next-generation, real-world visual perception and multimodal AI systems.


If you have spent the last 4+ years working directly with camera pipelines, real-time vision processing, deep learning models, and multimodal architectures, we want to hear from you.


What You’ll Do


  • Design & Deploy Vision Pipelines: Architect and optimize end-to-end computer vision pipelines processing multi-camera inputs in real time.
  • Integrate & Fine-Tune VLMs: Build, fine-tune, and deploy state-of-the-art Vision-Language Models (e.g., LLaVA, Florence, Qwen-VL) for complex visual reasoning, grounding, and multimodal understanding.
  • Edge & Cloud Optimization: Quantize, compile, and optimize models (TensorRT, ONNX, OpenVINO) for low-latency deployment on edge hardware and cloud infrastructure.
  • Core CV Engineering: Develop robust solutions for object detection, tracking, semantic segmentation, and camera calibration under challenging real-world conditions.
  • Cross-Functional Collaboration: Partner with data engineers, hardware specialists, and product teams to translate state-of-the-art vision research into production-grade features.


Key Requirements


DomainRequired QualificationsExperience4+ years of hands-on professional experience building and deploying production computer vision systems.Deep Learning & VLMsDemonstrated experience with PyTorch/TensorFlow, fine-tuning VLMs/multimodal models, and prompt engineering/grounding.Camera & Video StreamsDeep understanding of camera hardware, frame capture, OpenCV, RTSP streams, and real-time video stream optimization.Edge & PerformanceProficiency in C++ and Python; experience optimizing models with TensorRT, ONNX, CUDA, or similar frameworks.FundamentalsSolid background in linear algebra, geometry, camera calibration, tracking algorithms, and loss function design.


Preferred / Bonus Qualifications


  • Experience with spatial computing, 3D vision, or SLAM systems.
  • Track record of deploying models to embedded platforms (e.g., NVIDIA Jetson series).
  • Published research or open-source contributions in computer vision, multimodal AI, or vision-language integration.


What We Offer


  • Competitive Salary & Equity: Highly competitive compensation package commensurate with experience.
  • Cutting-Edge Stack: Access to state-of-the-art compute infrastructure and early-stage hardware/models.
  • Flexible Work Environment: Remote-first or hybrid options with flexible working hours.
  • Comprehensive Benefits: Premium health, dental, and vision insurance, 401(k) matching, and continuous learning/conference stipends.


How to Apply

Send your resume, GitHub profile, and a brief note detailing your most impactful computer vision or VLM project to hr@infolabsglobal.ai or apply directly via the link below.


#JobOpening #Hiring #ComputerVision #VLM #MultimodalAI #MachineLearning #DeepLearning #AIJobs #RemoteJobs #CPlusPlus #PyTorch

Similar jobs

Apply on LinkedIn