Archived listing. This role was posted over 30 days ago. We keep it for reference, but the employer may have already filled it. See today's verified listings.

Lead Site Reliability Engineer - Imunify Reliability Platform (remote-only)

CloudLinux Anywhere in the World

Apply now Save · sign in Alert me to jobs like this

This one is closed

Live roles like this one

See every "Site Reliability Engineer" role →

Posted 22 Aug 2026
Last seen 22 Aug 2026
Location Anywhere in the World
Lifecycle mature
Grade D

This listing scored 52/100, which is a D. It lost the most ground on pay transparency.

-10 Ghost-job penalty — Deducted for signals that this posting may not be a real, currently-open role — staleness, repeated relisting, or talent-pool language.

Every figure above is arithmetic over the posting itself — its salary field, its text, its age, its tags and how many sources carry it. How the grades work →

kubernetes python rust

The problem you'd own

Imunify360 is a multi-layer Linux server security suite — WAF, IDS/IPS, malware scanning and cleanup, proactive defence, patch management, reputation — running as an agent on hundreds of thousands of customer servers, backed by a cloud estate of scanning, correlation and signature-delivery services on our own bare metal.

Roughly 70 components currently ship without a defined service level indicator. Some are internal services we can scrape. Many are agent-side subsystems running on machines we do not own, reporting through a heartbeat we designed for something else. There is monitoring, and there are dashboards, and there is no coherent answer to the question "is this component doing its job right now, and how would we know if it stopped?".

We know the cost of that gap precisely, because we recently paid it: a security control was silently disabled across a large fraction of the fleet for 61 days. Every dashboard was green. The telemetry reported a ruleset version but not whether the control that consumed it was switched on, so a configuration change was indistinguishable from a broken updater. Three independent safety mechanisms existed and all three were gated behind the same condition that caused the failure.

Your job is to make that class of failure detectable in hours instead of months, across the whole product line, and to build the system that keeps it detectable as the product changes.

This is a greenfield charter inside a brownfield estate. You are not inheriting an SRE team, an SLO framework or a paging culture. You are defining them, with the engineering leads, and then making them stick.

What you'll do

1. Define what "working" means for ~70 components

2. Build the collection system

3. Build alerting and alert management

4. Build escalation

Requirements

What you'll bring

Required (Must-haves):

Valuable (Nice-to-haves):

Not this role

First year, in outcomes

30 days: Component inventory with named owners. SLI taxonomy and tiering agreed. 3 pilot components fully instrumented end to end as the reference implementation.

90 days: Collection pipeline in production. Tier-1 components (the ones whose failure is a customer security exposure) carry SLO, alert, runbook, owner. Escalation routing live for tier-1.

180 days: All ~70 components have a defined SLI and an owner. Alert taxonomy enforced; page volume and actionable-rate measured and published. Squad on-call operating.

365 days: Mean time to detect a silent control-degradation is under 24 hours, measured, against a 61-day baseline. Error-budget policy influences release decisions. The function is documented well enough that hire #2 and #3 are additive, not archaeological.

How we work

Remote-first and async across nine time zones. Weekly PO sync and architecture sync; monthly demo and OKR review; quarterly architecture summit. Decisions land as ADRs. Every output carries an owner and a due date. Postmortems are blameless and published, and we correct ourselves on the record when we get something wrong.

Benefits

What's in it for you?

By applying for this position, you consent to the processing of your personal data as described in our Privacy Policy (https://cloudlinux.com/candidate-privacy-notice), which provides detailed information on how we maintain and handle your data.

Apply for this role Opens apply.workable.com — verified as the employer's own application page

Quick question · anonymous · one tap

Would you apply to this job?

Answer to see what other job seekers said.

Your turn · no account needed

Help the next applicant

You may know something about this listing that we cannot see from here. One tap. No account needed. Signed-in reports earn points once the evidence agrees with you.

I know what it pays

What you were offered, quoted in an interview, or paid in this role. A range is fine.

Sign in with Google to earn points for reports — 100 confirmed points buy a week of Early Access.

Where this listing came from

  1. 22 Aug 2026 Real Work From Anywhere first sighting

Seen on 1 board over 46 days.