In this role, you will contribute to building advanced solutions that maximize the capabilities of heterogeneous CPU and GPU clusters while collaborating with a distributed team of more than 100 engineers and scientists driving impactful HPC innovation worldwide.
Our growing R&D team is focused on developing next-generation partial differential equation (PDE) solvers for computational fluid dynamics, seismic analysis, and other high-performance computing applications.
Responsibilities
Design, implement, and optimize high-performance computational kernels for PDE solvers targeting modern CPU and GPU architectures
Contribute to the development and optimization of performance-critical components, helping improve scalability, throughput, and computational efficiency under the guidance of senior engineers.
Collaborate closely with domain scientists, engineers, and researchers to transform mathematical models into efficient, vector-friendly algorithms
Profile and benchmark applications using hardware performance counters, flame graphs, and advanced performance analysis techniques to identify optimization opportunities
Prototype innovative numerical and performance-focused approaches in Python and evolve them into production-ready C++ solutions
Port and tune computational kernels for GPU environments, ensuring effective utilization of accelerator hardware and heterogeneous platforms
Influence technical decisions related to task-based runtimes, mixed-precision arithmetic strategies, and multi-GPU execution models
Contribute to software quality through testing, performance validation, code reviews, and knowledge sharing within the engineering team
Requirements
At least 6 months of commercial experience designing and optimizing finite-difference, finite-volume, or finite-element methods
Strong proficiency in C/C++17 and experience developing SIMD-friendly, cache-aware, high-performance code
Hands-on experience with MPI or another distributed-memory programming model for large-scale numerical computing
Practical knowledge of Linux performance-analysis tools such as perf, FlameGraph, Intel VTune, or similar solutions
Experience working with GPU architectures and programming environments such as CUDA or HIP, including porting CPU kernels to accelerators
Good Python scripting skills for automation, tooling, testing, and data analysis tasks
Strong analytical and problem-solving skills with a passion for algorithm design and optimization
Proficiency in communicating effectively in technical English during daily collaboration, code reviews, and design discussions
Collaborative mindset and willingness to learn, experiment, and contribute to complex R&D initiatives
SoftServe is an equal opportunity employer. Qualified applicants will receive consideration regardless of race, color, ancestry, ethnicity, national origin, religion, sex, sexual orientation, gender identity or expression, age, citizenship, disability, health condition, marital or family status, veteran status, or any other characteristic protected by applicable law.