Platform Operations Engineer
AcclaimDescription
This isn't classic ticket support. We're looking for someone who becomes the platform owner in the moment of an incident, not just passing the problem along, but understanding what's actually going on, driving communication through to full resolution, and being able to clearly explain what happened and why.
If you've written bug reports that developers could act on immediately without follow-up questions, that skill translates directly here - just applied to production incidents instead of test cases.
Requirements
- Experience working with monitoring and logging systems
- Ability to read and analyze logs (Grafana, Kibana, Loki)
- Basic issue localization (network, DNS, service connectivity)
- Basic understanding of Kubernetes: kubectl logs, kubectl describe, kubectl get
- Understanding of application configuration: Helm values, ConfigMaps, environment variables
- Infrastructure-level troubleshooting (service availability, node status, resources)
- Deep knowledge of our specific platform isn't required going in — we'll train you. Kubernetes/Helm experience is a plus but
- Background in QA/tech support
Nice to Have
- Experience with Helm
- Familiarity with CI/CD pipelines
- Understanding of microservices architecture
Soft Skills
- Ability to ask precise clarifying questions
- Structured problem description when escalating
- Independence and ownership
- A genuine desire to understand the issue, not just pass it along
- Growth-oriented mindset — this is an entry point, not a final destination; our platform evolves fast, and the role has a natural path toward DevOps/development for those who want it
- Willingness to work night shifts covering European hours; candidates based near the EST timezone are preferred
Responsibilities
- Receiving and processing requests via Telegram, Slack, and email
- Initial diagnostics: clarifying the nature of the issue and gathering details from the user
- Issue localization: identifying which component or service an incident relates to
- End-to-end incident ownership: not just escalation, but owning the incident from detection through resolution, keeping stakeholders updated along the way
- Basic infrastructure troubleshooting: logs, service status, configurations
- Executing deterministic runbook actions (e.g., telephony service failure identified → restart service → test call → confirm recovery)
- Escalating to L2/L3 (DevOps, developers) with prepared context
- Creating tickets and bug reports in the tracker
- Collaborating with the team to build and maintain runbooks and knowledge base entries for recurring issues
What we offer
- The team has built award-winning AI products for tech corporations - devices, voice assistants, products that are actually in the world
- Cutting-edge tech stack: Speech Technologies, NLP, Generative AI (LLMs, diffusion models), voice-first agentic architecture with privacy-first and on-premises deployment
- High engineering bar and real ownership - the team cares about what actually works in production, not what looks good in a demo, and you'll see the impact of your work directly
- Fast career progression - a senior-heavy team and a high volume of real problems means you grow faster than you would anywhere else
- Startup pace with enterprise stability - real clients, real revenue, no bureaucracy
- Fully remote across Europe
- 21 vacation days + public holidays + 5 sick days
- Private English lessons via Preply