HPC Performance and Validation Engineer
- Salary
- $180K–$260KUSD
- Hiring from
- United States
- Work type
- Hybrid
- Posted
- Sep 29, 2026
Is this job info correct?
Job Title: HPC Performance and Validation Engineer
Industry: High Performance Computing / AI Infrastructure
Location (city, state): Dallas, TX
Assignment Type: Direct hire
Pay: $180,000–$260,000 base salary, plus a potential $50,000–$100,000 bonus
Work Schedule: Hybrid; three days in the Dallas office and two days remote. The manager determines the in-office days.
Benefits: This position is eligible for 100% paid medical, dental, vision, and 401(k). Additional benefits include 25 days of PTO, an HSA contribution, lunch on office days, and a gym membership.
About The Company:
Our client is expanding the computing infrastructure used for large-scale AI, research, and simulation workloads.
Job Description:
We are seeking an engineer to measure, validate, and improve the performance of distributed HPC systems. You will create repeatable ways to confirm GPU clusters are ready for production workloads, investigate performance problems, and help technical teams make informed infrastructure decisions.
Key Responsibilities:
Relocation assistance may be tailored to the candidate, and TN visa candidates may be considered. The anticipated interview process includes an HR screen, a hiring manager meeting, a discussion of technical experience and concepts, and an onsite visit.
Industry: High Performance Computing / AI Infrastructure
Location (city, state): Dallas, TX
Assignment Type: Direct hire
Pay: $180,000–$260,000 base salary, plus a potential $50,000–$100,000 bonus
Work Schedule: Hybrid; three days in the Dallas office and two days remote. The manager determines the in-office days.
Benefits: This position is eligible for 100% paid medical, dental, vision, and 401(k). Additional benefits include 25 days of PTO, an HSA contribution, lunch on office days, and a gym membership.
About The Company:
Our client is expanding the computing infrastructure used for large-scale AI, research, and simulation workloads.
Job Description:
We are seeking an engineer to measure, validate, and improve the performance of distributed HPC systems. You will create repeatable ways to confirm GPU clusters are ready for production workloads, investigate performance problems, and help technical teams make informed infrastructure decisions.
Key Responsibilities:
- Build automated tests that verify GPU node readiness, health, and utilization across a large HPC environment.
- Benchmark compute, storage, and network performance using established and workload-specific tests.
- Profile AI and research workloads, identify bottlenecks, and work with engineering teams to improve results.
- Develop validation tools and continuous testing workflows using Python, Go, Kubernetes, and CI/CD pipelines.
- Establish monitoring and reporting that make cluster health and performance trends visible.
- Document benchmark findings and use the results to guide architecture and capacity decisions.
- Lead technical investigations and collaborate with infrastructure and research teams as the platform grows.
- Experience benchmarking, profiling, and tuning large-scale GPU clusters and HPC workloads.
- Strong understanding of accelerator performance and tools such as NVIDIA Nsight, DCGM, ClusterKit, or MLPerf.
- Experience testing network and storage performance, including InfiniBand or RoCE environments.
- Proficiency building automated validation tools with Python or Go in Linux environments; Kubernetes experience is also needed.
- Familiarity with observability tools such as Prometheus, Grafana, OpenTelemetry, or the ELK stack.
- Ability to lead complex technical work, explain findings clearly, and influence decisions across teams.
- A degree is preferred, but relevant experience is more important.
Relocation assistance may be tailored to the candidate, and TN visa candidates may be considered. The anticipated interview process includes an HR screen, a hiring manager meeting, a discussion of technical experience and concepts, and an onsite visit.