Description review

Senior Site Reliability Engineer

Playson · Europe · back to the listing

HR standards

74/100

solid

Title ↔ description

96/100

strong

Reads as

Site Reliability Engineer

100% confident

What this role officially is

ICT system administrator — ESCO, the EU occupation classification

ICT system administrators are responsible for the upkeep, configuration, and reliable operation of computer and network systems, servers, workstations and peripheral devices. They may acquire, install, or upgrade computer components and software; automate routine tasks; write computer programs; troubleshoot; train and supervise staff; and provide technical support. They ensure optimum system integrity, security, backup and performance.

Also known as: enterprise administrator, IT system administrator, ICT systems administrator, ICT sysadmins, IT systems administrator, ICT sysadmin

How others title the same work

Large employers

  • Senior Site Reliability / Gitops Engineer Canonical
  • Senior Site Reliability / Gitops Engineer Canonical Ltd.
  • Senior Site Reliability Engineer Canonical Ltd.
  • Site Reliability / Gitops Engineer Canonical Ltd.
  • Site Reliability Engineer Canonical Ltd.

Startups

  • Senior Software Engineer, Product Infrastructure Atlas
  • Site Reliability Engineer Beam

What the listing never says

  • 31 bullet points. Long requirement lists deter qualified candidates, who read them as hard gates. Scope clarity
  • No pay range published. Candidates cannot tell whether applying is worth their time. Pay transparency

The listing, marked up

Nothing in the wording of this listing tripped a check. The scores above still judge how complete and coherent it is.

About the Role

We’re looking for a Senior Site Reliability Engineer to join our Infrastructure Squad - a lean & senior team where ownership is high and expectations are even higher. This is a deeply hands-on role at the core of a high-traffic system, where you’ll be directly responsible for maintaining reliability, performance, and stability in a fast-paced environment.

You’ll be working on real-time production challenges, handling incidents, managing alerts, and being part of a critical on-call rotation. This role requires resilience, strong decision-making under pressure, and a proactive mindset to continuously improve systems operating at scale.

If you thrive in high-load environments, enjoy solving complex production issues, and want to have a direct impact on systems used by millions - this is the place for you.

Key Responsibilities


Own system reliability by actively monitoring platform health, managing alerts, and responding to incidents in real time


Participate in 24/7 on-call rotations, taking full ownership of production stability in a high-traffic (5–7k RPS) environment


Investigate incidents, perform root cause analysis, and implement long-term fixes to prevent recurrence


Build and continuously improve monitoring, alerting, and observability across the Kubernetes (EKS) ecosystem


Deploy, manage, and optimise infrastructure using Terraform, Helm, and GitOps tools (Flux/ArgoCD)


Drive automation and proactively improve system resilience, reducing manual intervention and recurring issues


Maintain and evolve CI/CD pipelines and infrastructure-as-code practices


Collaborate closely with engineering teams to support deployments and minimise user impact in a live environment


Introduce and integrate new tools and technologies to enhance scalability, reliability, and performance


Handle environment-specific requests and ensure smooth day-to-day platform operations under constant load

Requirements


Strong hands-on experience with Kubernetes (deployment, scaling, troubleshooting) in high-load environments


Experience with GitOps tools such as FluxCD or ArgoCD


Proven experience in incident response, root cause analysis, and postmortems in production systems


Solid experience with AWS, Terraform, Docker, and CI/CD pipelines


Experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, and logging stacks like ELK or CloudWatch


Strong understanding of networking concepts and protocols


Proficiency in at least one scripting language (e.g. Python, Go, Node.js)


Experience working with version control systems (Git)


Familiarity with incident management tools like PagerDuty, Opsgenie, or similar


Ability to operate effectively in a fast-paced, high-pressure environment with strong ownership and accountability


Proactive, resilient mindset with a focus on continuous improvement and system stability

What We Offer


Competitive Salary


Quarterly Bonuses


Unlimited Paid Time Off


Unlimited Paid Sick Leave


Remote & Flexible Working


Private Medical Insurance


Financial Support for Life Events


Professional Development Budget


International Exposure


Regular Company Events

*Benefits may vary depending on location and contractual agreement

Recruitment Process

1. HR Interview (30-45 min)

2. Technical interview (90 min)

4. Final Interview with C-level (60 min)

By submitting your application, you acknowledge that your personal data will be processed in accordance with our Privacy Policy.

How this was produced

Highlights are found by rule, not by a model: each one is a phrase matched at a known position, and every note is a template we wrote. The two scores come from a typed-decision model (Jev) that reads the listing against the official role definition and real listings for the same role, and returns probabilities rather than prose — it never writes any of the words on this page, and never chooses what to highlight.

Deterministic penalty applied to the HR score: 8 points (from 82 before penalties). Reviewed 21 Sep 2026.