Talent Shore logo

Platform Engineering Lead

Talent Shore
Posted 1 hour ago
South AfricaRemoteEngineering & Development
Is this job info correct?

Job Description

Platform Engineering Lead
Remote - South Africa
6 month contract to start - It will then roll into a Full time role thereafter


Our client is standing up as an independent company, and technology is at the centre of how we get there. We are working through three phases: Unify — becoming one business on one operational platform, quickly and safely; Extend — redesigning our customer and operational journeys to be digital-first, removing the manual steps and handoffs that slow us down; and Transform — reaching a state where systems run repeatable work by default and our people spend their time on what only people can do: teaching, solutioning, relationships, and judgement. This is not a transformation project with a finish line. It is a permanent shift in how people and technology work together at our company.

Humans do the work. Technology carries the weight.



Role Snapshot


 Title: Platform Automation Engineering Lead
 Location: Remote , South Africa
 Reporting to: CIO
 Employment Type: Full-time

 Direct reports:

o M365/Intune Engineer
o Azure Engineer
o Network Engineer



We are hiring a hands-on Platform Automation Engineering Lead to build a code-first, AI-enabled platform capability across Microsoft 365 and Azure, while establishing a reliable operating model for core network and connectivity services. This is a player-coach leadership role: you are accountable for platform outcomes and expected to execute directly on the highest-risk, highest-value technical work. This role is more than M365/Azure administration. It owns how we build and run platform automation as an engineering discipline: issues-to-code, policy-as-code, drift detection, agent-assisted delivery, MCP tool surfaces, and safe release patterns.



What You Will Own


 Platform automation engineering strategy and delivery across M365, Intune, Azure, and core network services.

 Operational reliability and security posture for identity, endpoint, cloud, and connectivity layers.

 Configuration-as-code standards, CI/CD controls, and change governance for managed platforms.

 Agent-enabled platform delivery model (issue quality, PR quality, review discipline, and release safety).

 MCP/tooling layer that exposes governed platform capabilities to agents and automation flows.

 Service health, incident response, problem management, and continuous hardening cadence.

 Team outcomes for three platform engineers (M365/Intune, Azure, Network).

 Monthly executive reporting to CIO on platform health, risk, performance, and priorities.



Key Responsibilities


 Lead the day-to-day platform engineering function spanning M365, Intune, Azure, and network operations.

 Define and execute a 12-month platform roadmap balancing reliability, security, modernization, and cost control.

 Own M365 and Intune engineering standards across identity, endpoint compliance, access policy, and governance.

 Own Azure engineering standards across landing zones, monitoring, security controls, backup/recovery, and platform services.

 Operate a code-first platform model where significant changes are designed, reviewed, tested, and released through versioned pipelines.

 Use AI-enabled engineering tooling to accelerate delivery while maintaining strict review and quality gates.

 Build or adopt MCP-compatible platform tool surfaces so agents can perform governed platform actions safely.

 Partner with network engineering to ensure resilient connectivity, secure access, and operational readiness.

 Drive incident and problem management discipline, including root-cause analysis and permanent fixes.

 Ensure all significant platform changes follow version control, peer review, and controlled release practices.

 Establish clear ownership boundaries and quarterly delivery goals for each direct report.

 Mentor and develop team capability while maintaining hands-on technical contribution.



First 90 Days Deliverables


 Platform health baseline across M365, Azure, Intune, and network: risks, gaps, ownership, and mitigation plan.

 Stabilization plan and execution for top-priority reliability and security issues.

 Operating cadence established for incidents, major changes, and service reviews.

 Team charter in place for M365/Intune, Azure, and Network engineers with clear outcomes and accountabilities.

 Code-first operating baseline in place: issue templates, PR quality standards, release and rollback controls.

 Initial AI-enabled delivery loop established with measurable quality gates for agent-assisted changes.

 Executive dashboard live with key platform KPIs, risks, and 30/60/90-day priorities.



Success Measures (First 12 Months)


 Platform reliability and security posture measurably improved across M365, Azure, Intune, and network layers.

 Change success rate improved with fewer high-impact incidents linked to platform changes.

 Critical risks reduced through a managed hardening and remediation backlog.

 Strong service ownership behavior established across the three-engineer team.

 Faster incident resolution and improved post-incident fix completion.

 Clear roadmap delivery against agreed quarterly outcomes.

 High-quality code-first delivery: materially improved first-pass PR/CI success and reduced rework.

 Practical AI-enabled engineering adoption that increases throughput without reducing control quality.



Equal Opportunity

We are committed to fair and inclusive hiring. Applicants are encouraged to apply even if they do not meet every preferred qualification.



Application Instructions

Candidates should provide:

 CV

 A short note covering:

o A code-first platform capability they built (not just administered)
o A cross-domain platform incident they stabilized and what changed permanently
o How they used AI-enabled engineering to increase delivery speed while preserving control quality






Requirements



 8-12+ years in platform engineering, cloud engineering, or infrastructure engineering roles.

 Demonstrated leadership of mixed platform teams and cross-domain operations.

 Strong hands-on expertise in Microsoft 365 and Azure platform engineering.

 Practical experience with Intune, identity and access controls, and endpoint governance.

 Strong scripting and automation engineering capability (PowerShell and/or Python) with production delivery history.

 Experience with API-driven platform operations, YAML/pipeline patterns, and configuration-as-code.

 Proven history of shipping real AI-assisted engineering work in production environments.

 Strong operational background in incident/problem/change management for production environments.

 Ability to lead outcomes across teams without direct line authority.

 Strong communication skills for technical, business, and executive audiences.



Preferred Qualifications


 Experience in SaaS, enterprise software, or multi-country operational environments.

 Familiarity with infrastructure-as-code/configuration-as-code pipelines and policy-driven governance.

 Experience with MCP, agent tooling, or secure platform tool exposure patterns.

 Experience with Azure AI Foundry, Logic Apps, APIM, or related automation integration services.

 Exposure to modern network security patterns and cloud connectivity architecture.

 Relevant certifications across Azure, M365, security, or network engineering.


Quick Screen (Must-Have Evidence)

 Led production platform automation using version control and CI/CD.

 Deep hands-on experience across M365, Intune, and Azure.

 Incident-to-permanent-fix ownership across multiple platform domains.

 Player-coach leadership: leads engineers while still shipping technical work.

 Real AI-assisted engineering use in production with review and release controls.



Ways of Working

 Outcome-based performance with clear quarterly commitments.

 Player-coach leadership: direct technical execution combined with team development.

 High accountability for uptime, security, and service quality.

 Structured operating cadence over reactive firefighting.

 Inclusive, respectful, and high-performance team culture.


Fast Reject Signals

 Strong platform administration profile but weak engineering evidence.

 No proof of release quality gates, rollback discipline, or peer review.

 People manager profile with limited hands-on technical ownership.





Benefits

Remote working with one of the global leaders in e-learning
Opportunity to fully own the platform automation function & team as the business embarks on an exciting journey after being recently acquired
Work autonomously, supporting a CIO that you can learn from while adding great value


Similar jobs