Platform Engineer – Developer Experience (DevEx)
- Hiring from
- Probably Worldwide
- Work type
- Remote
- Posted
- Sep 29, 2026
Platform Engineer – Developer Experience (DevEx)
Primary responsibilities
Reporting to the Platform Engineering Manager, you will join Ledger's Platform Engineering team as a Platform Engineer with a Developer Experience (DevEx) focus. You will embed alongside product engineering teams to remove friction from the build, package, and deploy path, owning the tooling, golden paths, and self-service platforms that let engineers ship faster and more safely, without needing to become infrastructure experts themselves.
This role sits at the intersection of Platform Engineering and the product teams it supports. You will partner closely with product engineering leads (and with the DevOps engineers embedded in those teams, who own pipeline configuration) to make sure the platform's runners, registries, delivery tooling, and observability stack are fast, reliable, secure, and easy to consume. Turning platform capabilities into products that engineering teams actually want to use.
We are looking for someone who thinks about infrastructure the way a product manager thinks about a product: who is the user, what is their friction, and how do we measure whether we made their day better. You will bring strong systems and automation experience together with a genuine interest in developer experience, documentation, and cross-team collaboration.
In this role, you will:
- CI/CD platform & runner fleet
- Own and evolve the self-hosted GitHub Actions runner fleet, including Linux runners managed via Actions Runner Controller (ARC), self-hosted macOS runners, and the Terraform modules that provision and scale self-hosted infrastructure.
- Right-size runner classes (CPU/memory/storage) against real workload profiles, balancing cost efficiency against engineering velocity, and supporting engineering teams in understanding and owning the resource footprint of their own jobs.
- Support the migration away from GitHub-hosted runners, including evaluating trade-offs between containerized and VM-based execution for workloads with special hardware requirements (e.g. KVM access for Android emulators).
- Partner with embedded DevOps engineers on pipeline and dependency-graph optimizations (build triggers, job fan-out, monorepo change detection) to reduce queue times and unnecessary CI spend — you own the platform and runner layer these pipelines execute on.
- Investigate and resolve macOS virtualization performance issues (e.g. Tart) and help evaluate alternative approaches such as remote simulator hosts.
Artifact & package management
- Administer the JFrog Platform (Artifactory, Xray, Curation) as the organization's primary artifact repository, including repository structure, access control, vulnerability scanning policies, and curation/allow-listing of upstream dependencies.
- Provide light-touch, minimal support for GitHub Container Registry (GHCR) and GitHub Packages where teams use them alongside JFrog.
- Operate and maintain ChartMuseum as the organization's Helm chart repository, hosted on S3, including storage lifecycle, access, and chart promotion workflows.
- Define and enforce container image ownership boundaries — platform-owned hardened/minimal base images (e.g. Docker Hardened Images, Chainguard) versus engineering-owned project extensions — and drive Dockerfile build and layer optimization best practices across teams.
Progressive delivery & infrastructure automation
- Operate two ArgoCD / Argo Rollouts environments, one serving internal tooling and one serving product workloads supporting GitOps-based, progressive delivery (canary/blue-green) patterns for engineering teams.
- Author, maintain, and support Helm charts and Kustomize overlays as the standard configuration management approach for Kubernetes workloads, providing golden-path templates that reduce boilerplate for product teams.
- Design, review, and evolve Terraform and Terragrunt modules for infrastructure-as-code, including maintaining a public/internal module registry that product and platform teams can consume in a self-service manner.
Observability & reliability
- Improve and help unify the observability stack; Grafana, Prometheus, and Alertmanager. Fixing metric accuracy issues, building meaningful dashboards, and defining SLOs/alerting that reflect real platform and pipeline health.
- Work across Datadog and VictoriaLogs/VictoriaMetrics to reduce tool sprawl, and help define a coherent strategy for where metrics, logs, and events should live.
- Provide platform-level telemetry (runner utilization, build queue times, artifact scan results, deployment health) that gives engineering teams and leadership a single, trustworthy view of pipeline and platform performance.
Internal developer platform, policy & modern workloads
- Evaluate and help operate Crossplane as a control-plane layer for self-service, Kubernetes-native infrastructure provisioning, complementing the existing Terraform/Terragrunt IaC estate.
- Contribute to a developer portal / internal catalog strategy (e.g. Backstage, Cortex, or Port) so engineers have a single place to discover services, ownership, docs, and golden-path templates.
- Implement and maintain policy-as-code guardrails using Kyverno and/or Open Policy Agent (OPA) across the Kubernetes and CI/CD estate (e.g. image provenance, resource limits, security baselines).
- Support Knative-based serverless/event-driven workloads on Kubernetes where teams need scale-to-zero or request-driven compute.
- Roll out OpenTelemetry instrumentation and pipelines to standardize traces, metrics, and logs across services, feeding the Grafana/Prometheus/Datadog/VictoriaLogs stack with consistent, correlated telemetry.
Developer experience & enablement
- Act as the platform's embedded point of contact for one or more product engineering teams: understand their day-to-day friction, prioritize platform work accordingly, and close the loop on delivery.
- Build self-service tooling, templates, and documentation (golden paths) so engineers can provision runners, publish artifacts, and deploy without needing deep platform expertise.
- Clearly document ownership boundaries between platform and product teams (e.g. base images vs. application extensions) and help engineering teams operate confidently within them.
- Contribute to the platform roadmap, bringing a developer-experience lens to prioritization — and measure success in terms of engineer-facing outcomes such as build time, queue time, and self-service adoption, not just infrastructure uptime.
Requirements
Experience & technical skills
- 5+ years in a Platform Engineering, DevOps, SRE, or infrastructure engineering role, ideally supporting product engineering teams directly.
- Hands-on experience with GitHub Actions, including self-hosted runner architectures (Actions Runner Controller / Kubernetes-based Linux runners, self-hosted macOS runners) and provisioning them via Terraform.
- Practical knowledge of the JFrog Platform (Artifactory, Xray, and ideally Curation) for artifact management and dependency security.
- Familiarity with container registries beyond a single vendor (GHCR/GitHub Packages) and Helm chart repositories such as ChartMuseum on object storage (S3).
- Experience with GitOps continuous delivery using ArgoCD, and ideally Argo Rollouts for progressive delivery (canary/blue-green) across multiple clusters/environments.
- Strong Kubernetes configuration management skills with Helm and Kustomize.
- Solid Infrastructure-as-Code experience with Terraform and Terragrunt, including designing and consuming shared/reusable modules.
- Experience with Crossplane (or a comparable Kubernetes-native control-plane approach) for building self-service infrastructure APIs, alongside deep Terraform proficiency for the underlying IaC.
- Experience building or operating an internal developer portal / software catalog such as Backstage, Cortex, or Port, including modeling service ownership and scorecards.
- Hands-on experience with policy-as-code enforcement in Kubernetes using Kyverno and/or Open Policy Agent (OPA / Gatekeeper) — e.g. image provenance, resource limits, security baselines.
- Familiarity with Knative or similar Kubernetes-native serverless/eventing frameworks for scale-to-zero or request-driven workloads.
- Experience building, hardening, and optimizing container images (Dockerfile best practices, minimal/hardened base images such as Docker Hardened Images or Chainguard).
- Working knowledge of an observability stack — Grafana, Prometheus, and Alertmanager — plus exposure to Datadog and/or VictoriaLogs/VictoriaMetrics, and practical experience instrumenting services and pipelines with OpenTelemetry (traces, metrics, logs).
- Comfortable in Linux/Unix environments, with solid scripting ability (e.g. Bash, Python, or Go) and strong Git fundamentals.
- AWS cloud experience (compute, storage/S3, networking fundamentals) is a plus.
Soft skills
- Genuine product mindset toward internal platforms: you think in terms of the engineer as a customer, and you seek out and act on their feedback.
- Strong cross-functional collaborator, comfortable embedding with a product engineering team and building trust with people who don't share your infrastructure background.
- Clear written and verbal communicator, able to explain infrastructure concepts (e.g. Kubernetes, runner sizing, build performance) to engineers without a platform background, and able to tailor the level of detail to the audience.
- Pragmatic and cost-conscious: able to balance reliability, security, developer speed, and infrastructure cost, and to make and explain trade-offs rather than defaulting to more resources.
- Ownership mindset with strong personal accountability; comfortable working autonomously and driving initiatives from discovery through to rollout.
- Curious and continuously improving: proactively investigates root causes (e.g. inconsistent metrics, flaky runners) rather than just resolving symptoms.
- Comfortable with ambiguity and evolving priorities, and able to sequence and negotiate scope in a fast-moving, PI-based planning environment.
- Empathetic and collaborative when defining ownership boundaries between platform and product teams, favoring clarity and shared documentation over friction.
Nice to have
- Experience supporting mobile/device-testing CI (e.g. Android emulator/KVM requirements, iOS simulators) in a self-hosted runner context.
- Experience with developing platform API and CLI for developers to consume blueprints and patterns using either Kratix or Crossplane tooling.
- Experience with large Kubernetes cluster fleet automation from provisioning to upgrade to decommission
- Experience operating in a large monorepo, including build-graph/dependency-aware CI triggering.
- Prior experience defining or operating SLOs/SLAs and error budgets for internal platform services.