Description review
Senior Cloud Engineer, SRE
Seismic · United States · back to the listing
HR standards
61/100
needs work
Title ↔ description
93/100
strong
Reads as
Site Reliability Engineer
100% confident
What this role officially is
ICT system administrator — ESCO, the EU occupation classification
ICT system administrators are responsible for the upkeep, configuration, and reliable operation of computer and network systems, servers, workstations and peripheral devices. They may acquire, install, or upgrade computer components and software; automate routine tasks; write computer programs; troubleshoot; train and supervise staff; and provide technical support. They ensure optimum system integrity, security, backup and performance.
Also known as: enterprise administrator, IT system administrator, ICT systems administrator, ICT sysadmins, IT systems administrator, ICT sysadmin
How others title the same work
Large employers
- Senior Site Reliability / Gitops Engineer Canonical
- Senior Site Reliability / Gitops Engineer Canonical Ltd.
- Senior Site Reliability Engineer Canonical Ltd.
- Site Reliability / Gitops Engineer Canonical Ltd.
- Site Reliability Engineer Canonical Ltd.
Startups
- Senior Software Engineer, Product Infrastructure Atlas
- Site Reliability Engineer Beam
What the listing never says
- 23 bullet points. Long requirement lists deter qualified candidates, who read them as hard gates. Scope clarity
- No location or timezone policy stated, so a candidate cannot tell where they may work from. Scope clarity
The listing, marked up
Nothing in the wording of this listing tripped a check. The scores above still judge how complete and coherent it is.
Overview
Seismic is seeking a Senior Cloud Engineer to advance reliability across our AWS, Azure, IBM Cloud, and OCI environments, working hands-on to build automation, improve observability, and strengthen incident response as part of a globally distributed SRE org. We prioritize a cloud agnostic approach to architecture, automation, and engineering standards to support velocity and scale. You will contribute to our reliability roadmap, provide technical guidance on reliability best practices, participate in major incident response, and work collaboratively across Product & Engineering, Security, and Customer-facing teams.
Who you are:
• Experience in a production facing SRE role supporting a complex SaaS environment.
• You approach discussions about existing solutions, processes, and proposals with curiosity, respect, and a collaborative mindset.
• Strong technical judgment across distributed systems, multi-cloud environments, Kubernetes, networking, various infrastructure technologies, GitOps, and CI/CD.
• Experience establishing and maturing SRE principles and practices, including SLOs, error budgets, observability, capacity planning, incident response, and toil elimination.
• Proficient in using observability data to resolve high-severity incidents and dig deeper into root cause during postmortems.
• Experience leading through high-pressure incidents and communicating clearly with technical teams, executives, customer-facing stakeholders, and third-party vendors.
What you'll be doing:
Operating Model
• Design, build, and maintain automation and tooling that reduces operational toil.
• Contribute to maturing reliability practices in partnership with Product & Engineering leaders.
• Contribute to an inclusive, high-accountability culture that encourages curiosity, collaboration, and blameless improvement.
Incident Management and Operational Excellence
• Actively participate in the health and continuous improvement of the incident-management lifecycle, including detection, engagement, escalation, mitigation, stakeholder communication, post-incident review, and corrective-action follow-through.
• Participate in a 12-hour follow-the-sun on-call rotation within the Global SRE team.
• Ensure incident practices are customer-centered, data-driven, blameless, and consistent across teams while preserving accurate severity and escalation decisions.
• Use alert, incident, support, and SLO trends to move the organization from reactive response toward proactive risk reduction.
• Build measurable feedback loops that connect incident learning to engineering standards, service maturity, product priorities, and vendor actions.
• Work closely with application engineering teams and incorporate their feedback to improve developer experience and reduce toil.
Reliability Strategy and Service Maturity
• Partner with service owners to ensure production readiness standards are met before each release stage.
• Provide an SRE point of view on capacity planning, resilience testing, game days, disaster-recovery readiness, and modernization of fragile or legacy workloads.
• Partner with Product and Engineering leaders to document critical customer workflows, define health expectations, surface dependencies early, and align reliability investment with business priorities.
• Participate in cross-team reliability engagements, influencing outcomes without relying on direct authority.
• Build strategic relationships with vendors in the observability, incident response, and cloud infrastructure domains.
AI-First Reliability Engineering
• Responsibly adopt AI-assisted and agentic workflows for alert triage, incident mitigation, postmortems, trend analysis, capacity planning, SLO analysis, and self-service knowledge.
• Keep qualified humans in the decision loop for production-impacting actions.
• Improve the context available to reliability workflows by strengthening service metadata, observability data, incident records, runbooks, architecture documentation, and corrective-action quality.
What we have for you:
At Seismic, we’re committed to providing benefits and perks for the whole self. To explore our benefits available in each country, please visit the Global Benefits page.
If you are an individual with a disability and would like to request a reasonable accommodation as part of the application or recruiting process, please click here.
Seismic is an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to gender, age, race, religion, or any other classification which is protected by applicable law.
We are committed to fair and equitable compensation practices.Seismic's annual base salary range for this position will vary based on applicant's location, experience, job level, skills, and abilities as well as internal equity and alignment market data. The range listed below is the minimum to the maximum of our target hiring range. Seismic's salary range for this position is: USD $150,000.00/Yr. - USD $175,000.00/Yr.This position is also eligible to participate in Seismic's incentive plans in addition to base salary. The actual incentive amount will very and will be subject to the terms and conditions set in the applicable incentive plan.
Originally posted on Himalayas
Seismic is seeking a Senior Cloud Engineer to advance reliability across our AWS, Azure, IBM Cloud, and OCI environments, working hands-on to build automation, improve observability, and strengthen incident response as part of a globally distributed SRE org. We prioritize a cloud agnostic approach to architecture, automation, and engineering standards to support velocity and scale. You will contribute to our reliability roadmap, provide technical guidance on reliability best practices, participate in major incident response, and work collaboratively across Product & Engineering, Security, and Customer-facing teams.
Who you are:
• Experience in a production facing SRE role supporting a complex SaaS environment.
• You approach discussions about existing solutions, processes, and proposals with curiosity, respect, and a collaborative mindset.
• Strong technical judgment across distributed systems, multi-cloud environments, Kubernetes, networking, various infrastructure technologies, GitOps, and CI/CD.
• Experience establishing and maturing SRE principles and practices, including SLOs, error budgets, observability, capacity planning, incident response, and toil elimination.
• Proficient in using observability data to resolve high-severity incidents and dig deeper into root cause during postmortems.
• Experience leading through high-pressure incidents and communicating clearly with technical teams, executives, customer-facing stakeholders, and third-party vendors.
What you'll be doing:
Operating Model
• Design, build, and maintain automation and tooling that reduces operational toil.
• Contribute to maturing reliability practices in partnership with Product & Engineering leaders.
• Contribute to an inclusive, high-accountability culture that encourages curiosity, collaboration, and blameless improvement.
Incident Management and Operational Excellence
• Actively participate in the health and continuous improvement of the incident-management lifecycle, including detection, engagement, escalation, mitigation, stakeholder communication, post-incident review, and corrective-action follow-through.
• Participate in a 12-hour follow-the-sun on-call rotation within the Global SRE team.
• Ensure incident practices are customer-centered, data-driven, blameless, and consistent across teams while preserving accurate severity and escalation decisions.
• Use alert, incident, support, and SLO trends to move the organization from reactive response toward proactive risk reduction.
• Build measurable feedback loops that connect incident learning to engineering standards, service maturity, product priorities, and vendor actions.
• Work closely with application engineering teams and incorporate their feedback to improve developer experience and reduce toil.
Reliability Strategy and Service Maturity
• Partner with service owners to ensure production readiness standards are met before each release stage.
• Provide an SRE point of view on capacity planning, resilience testing, game days, disaster-recovery readiness, and modernization of fragile or legacy workloads.
• Partner with Product and Engineering leaders to document critical customer workflows, define health expectations, surface dependencies early, and align reliability investment with business priorities.
• Participate in cross-team reliability engagements, influencing outcomes without relying on direct authority.
• Build strategic relationships with vendors in the observability, incident response, and cloud infrastructure domains.
AI-First Reliability Engineering
• Responsibly adopt AI-assisted and agentic workflows for alert triage, incident mitigation, postmortems, trend analysis, capacity planning, SLO analysis, and self-service knowledge.
• Keep qualified humans in the decision loop for production-impacting actions.
• Improve the context available to reliability workflows by strengthening service metadata, observability data, incident records, runbooks, architecture documentation, and corrective-action quality.
What we have for you:
At Seismic, we’re committed to providing benefits and perks for the whole self. To explore our benefits available in each country, please visit the Global Benefits page.
If you are an individual with a disability and would like to request a reasonable accommodation as part of the application or recruiting process, please click here.
Seismic is an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to gender, age, race, religion, or any other classification which is protected by applicable law.
We are committed to fair and equitable compensation practices.Seismic's annual base salary range for this position will vary based on applicant's location, experience, job level, skills, and abilities as well as internal equity and alignment market data. The range listed below is the minimum to the maximum of our target hiring range. Seismic's salary range for this position is: USD $150,000.00/Yr. - USD $175,000.00/Yr.This position is also eligible to participate in Seismic's incentive plans in addition to base salary. The actual incentive amount will very and will be subject to the terms and conditions set in the applicable incentive plan.
Originally posted on Himalayas