24-MAG logo

Remote | Senior Technical Architect — $60–$130/hour

24-MAG
Posted 13 hours ago
United StatesRemote$60–$130/hrEngineering & Development
Is this job info correct?

We are sharing a specialised consulting opportunity for experienced Senior Technical Architects with strong expertise in cloud infrastructure, distributed systems, platform engineering, DevOps, SRE, production architecture, and resilience engineering to contribute to an advanced AI training and cloud-infrastructure evaluation project.

Selected professionals will design realistic cloud-infrastructure tasks and reinforcement-learning environments that test an AI system's ability to design, deploy, secure, scale, troubleshoot, and recover production-grade systems. The work requires substantial hands-on ownership of production infrastructure, strong systems judgement, and the ability to create reproducible environments, deterministic tests, and technically rigorous reference solutions. No prior experience in AI is required.

Key Responsibilities

Cloud Architecture & Systems Design

  • Design realistic infrastructure scenarios involving production-grade cloud systems
  • Create tasks spanning distributed systems, networking, security, scalability, and reliability
  • Evaluate architectural trade-offs across performance, availability, and operational complexity
  • Define clear system requirements and expected behaviours
  • Apply practical production experience to benchmark design

Distributed Systems

  • Develop scenarios involving scalable APIs and distributed services
  • Evaluate queues, durable storage, autoscaling, and service coordination
  • Model partial-failure and degraded-service conditions
  • Design tasks that test distributed-systems reasoning
  • Assess system behaviour under realistic failure scenarios

Production Infrastructure Ownership

  • Apply hands-on experience managing production platforms
  • Design environments reflecting real operational constraints
  • Evaluate deployment, scaling, security, observability, and recovery workflows
  • Incorporate practical infrastructure trade-offs into task design
  • Identify common production failure patterns and operational weaknesses

Infrastructure Environment Development

  • Build reproducible and containerised technical environments
  • Create valid reference implementations
  • Develop intentionally defective variants for evaluation purposes
  • Define deterministic environment setup and execution requirements
  • Ensure tasks can be consistently reproduced and validated

Infrastructure Validation & Testing

  • Develop deterministic integration tests
  • Create load and scalability tests
  • Build security and access-control validation
  • Implement failure-injection and resilience tests
  • Validate deployment, rollback, and recovery behaviour

Networking & Security

  • Design scenarios involving private networking and service connectivity
  • Apply IAM and least-privilege principles
  • Evaluate service-to-service security
  • Identify insecure or overly permissive infrastructure configurations
  • Test model understanding of secure cloud-architecture practices

Observability & Reliability Engineering

  • Incorporate logging, metrics, tracing, and operational telemetry
  • Define measurable SLOs and reliability criteria
  • Evaluate system behaviour during partial failures
  • Design realistic incident and troubleshooting scenarios
  • Assess observability quality and operational diagnosability

Deployment & Recovery

  • Develop scenarios involving rolling deployments
  • Evaluate rollback strategies and failure recovery
  • Design disaster-recovery tasks
  • Test infrastructure resilience under deployment failures
  • Assess whether recovery approaches preserve service availability and data integrity

Infrastructure Automation

  • Write infrastructure automation or testing tools
  • Develop utilities supporting environment setup and validation
  • Debug containerised infrastructure
  • Support repeatable provisioning and teardown workflows
  • Apply appropriate programming languages to infrastructure tasks

Reinforcement Learning Environment Development

  • Create technical environments that test AI decision-making
  • Define measurable success criteria
  • Build golden reference solutions
  • Construct defective or adversarial variants
  • Support evaluation of AI performance across multi-step infrastructure workflows

Peer Review & Technical Quality

  • Review tasks developed by other technical experts
  • Identify ambiguity, unrealistic assumptions, or validation gaps
  • Assess benchmark difficulty and technical correctness
  • Improve task reproducibility and grading reliability
  • Maintain strong engineering standards across evaluation content

Ideal Profile

  • Senior-level experience in technical architecture, cloud infrastructure, platform engineering, DevOps, systems engineering, or SRE
  • Demonstrated personal ownership of production infrastructure or a production platform
  • Strong knowledge of distributed systems
  • Experience with scalable APIs and service architectures
  • Strong understanding of queues, autoscaling, durable storage, and partial-failure scenarios
  • Practical IAM and least-privilege experience
  • Strong knowledge of private networking and service-to-service security
  • Experience with observability and measurable SLOs
  • Experience with rolling deployments and rollback strategies
  • Familiarity with disaster recovery and resilience engineering
  • Ability to write infrastructure automation or testing tools
  • Strong debugging skills in containerised environments
  • Experience with Terraform or OpenTofu is advantageous
  • Experience with AWS, Azure, GCP, Kubernetes, or multi-cloud environments is valuable
  • Background building internal developer platforms, edge infrastructure, or shared platform services is beneficial
  • Experience with chaos engineering, fault injection, local cloud emulators, or resilience testing is advantageous
  • Experience creating technical evaluations, automated grading systems, or AI environments is helpful but not required
  • No prior AI-training or model-evaluation experience is required

Engagement Details

  • Independent contractor engagement
  • Fully remote
  • Displayed compensation range: $60–$130/hour
  • Actual compensation structure is output-based, with payment made per task that meets project specifications
  • Task completion time may vary depending on individual experience and workflow
  • Minimum weekly submission requirements apply; the source does not specify the exact number of required tasks
  • Work will involve cloud architecture, distributed systems, networking, IAM, observability, production infrastructure, resilience testing, deployment workflows, and AI evaluation environments
  • Strong hands-on production infrastructure experience is central to this engagement
  • Roles are typically filled within approximately 48 hours
  • Selected experts are expected to begin initial tasks within approximately 24–48 hours after onboarding
  • Project scope, technical environments, benchmark requirements, and evaluation standards may evolve depending on project needs
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, cloud provider, software organisation, infrastructure environment, or other third party

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy

Similar jobs