Remote | Senior Platform Engineer — $60–$130/hour
24-MAGWe are sharing a specialised consulting opportunity for experienced Senior Platform Engineers with strong expertise in cloud infrastructure, distributed systems, platform engineering, DevOps, SRE, production infrastructure, and resilience engineering to contribute to an advanced AI training and cloud-infrastructure evaluation project.
Selected professionals will design realistic cloud-infrastructure tasks and reinforcement-learning environments that test an AI system's ability to design, deploy, secure, scale, troubleshoot, and recover production-grade systems. The work requires substantial hands-on ownership of production infrastructure, strong systems judgement, and the ability to create reproducible environments, deterministic tests, reference solutions, and intentionally defective variants. No prior experience in AI is required.
Key Responsibilities
Cloud Infrastructure Task Design
- Create realistic infrastructure tasks involving production-grade cloud systems
- Develop scenarios spanning distributed systems, networking, security, scalability, and reliability
- Define technically meaningful success criteria
- Incorporate practical production constraints into benchmark scenarios
- Design tasks that require substantive infrastructure reasoning
Platform Engineering & Systems Architecture
- Apply platform-engineering expertise to complex infrastructure problems
- Evaluate architecture across compute, storage, networking, and service layers
- Analyse trade-offs involving reliability, scalability, security, and maintainability
- Design realistic system topologies
- Support technically rigorous infrastructure evaluation workflows
Distributed Systems
- Develop scenarios involving scalable APIs and distributed services
- Evaluate queues, autoscaling, durable storage, and service coordination
- Model partial-failure and degraded-service conditions
- Assess distributed-system behaviour under realistic workloads
- Identify weaknesses affecting resilience or scalability
Production Infrastructure Ownership
- Apply hands-on experience managing production platforms
- Design environments reflecting real operational constraints
- Evaluate deployment, scaling, security, observability, and recovery workflows
- Incorporate practical incident and failure patterns into tasks
- Use production experience to strengthen benchmark realism
Reproducible Environment Development
- Build reproducible and containerised infrastructure environments
- Create valid reference solutions
- Develop intentionally defective variants for evaluation
- Define consistent environment setup and execution requirements
- Ensure technical tasks can be reproduced reliably
Validation & Automated Testing
- Develop deterministic integration tests
- Create load and scalability tests
- Build security and access-control validation
- Implement failure-injection, deployment, and recovery tests
- Verify infrastructure configuration, topology, and runtime behaviour
Networking & IAM
- Design scenarios involving private networking and service connectivity
- Apply IAM and least-privilege principles
- Evaluate service-to-service security
- Identify insecure or overly permissive access patterns
- Validate network and identity configurations against task requirements
Observability & SLOs
- Incorporate logging, metrics, tracing, and operational telemetry
- Define measurable SLOs and reliability criteria
- Evaluate system behaviour during failure conditions
- Design tasks requiring observability-driven troubleshooting
- Assess whether infrastructure provides sufficient operational visibility
Deployment, Rollback & Disaster Recovery
- Develop scenarios involving rolling deployments
- Evaluate rollback strategies
- Design disaster-recovery and restoration tasks
- Assess infrastructure response to failed releases or degraded services
- Validate recovery mechanisms and resilience assumptions
Infrastructure Automation
- Write infrastructure automation or testing tools
- Develop utilities for environment setup and validation
- Debug containerised systems
- Support repeatable provisioning and teardown workflows
- Apply relevant programming languages to infrastructure automation
Reinforcement Learning Environment Development
- Create environments that test AI decision-making across infrastructure workflows
- Define measurable task requirements
- Build golden reference implementations
- Construct defective or adversarial variants
- Support reliable evaluation of multi-step technical reasoning
Peer Review & Technical Quality
- Review tasks developed by other experts
- Identify ambiguity, unrealistic assumptions, or validation gaps
- Assess technical correctness and benchmark difficulty
- Improve reproducibility and grading reliability
- Maintain strong engineering-quality standards across project deliverables
Ideal Profile
- Senior-level experience in cloud infrastructure, platform engineering, DevOps, systems engineering, or SRE
- Demonstrated personal ownership of a production platform or production infrastructure
- Strong knowledge of distributed systems
- Experience with scalable APIs
- Strong understanding of queues, autoscaling, durable storage, and partial-failure scenarios
- Practical IAM and least-privilege experience
- Strong knowledge of private networking and service-to-service security
- Experience with observability and measurable SLOs
- Experience with rolling deployments and rollback strategies
- Familiarity with disaster recovery and resilience engineering
- Ability to write infrastructure automation or testing tools
- Strong debugging skills in containerised environments
- Experience with Terraform or OpenTofu is advantageous
- Experience with AWS, Azure, GCP, Kubernetes, or multi-cloud infrastructure is valuable
- Background building internal developer platforms, edge infrastructure, or shared platform services is beneficial
- Experience with chaos engineering, fault injection, local cloud emulators, or resilience testing is advantageous
- Experience creating technical evaluations, automated grading systems, or AI environments is helpful but not required
- No prior AI-training or model-evaluation experience is required
Engagement Details
- Independent contractor engagement
- Fully remote
- Displayed compensation range: $60–$130/hour
- Actual compensation structure is output-based, with payment made per task that meets project specifications
- Task completion time may vary depending on individual experience and workflow
- Minimum weekly submission requirements apply; the source does not specify the exact number of required tasks
- Work will involve cloud architecture, distributed systems, networking, IAM, observability, production infrastructure, resilience testing, deployment workflows, and reinforcement-learning environments
- Strong hands-on production infrastructure experience is central to this engagement
- Roles are typically filled within approximately 48 hours
- Selected experts are expected to begin initial tasks within approximately 24–48 hours after onboarding
- Project scope, technical environments, benchmark requirements, and evaluation standards may evolve depending on project needs
- Work must be completed without using confidential or proprietary information belonging to any employer, client, cloud provider, software organisation, infrastructure environment, or other third party
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy