Description review
Senior Site Reliability Engineer
Endear · United States · back to the listing
HR standards
78/100
solid
Title ↔ description
96/100
strong
Reads as
Site Reliability Engineer
100% confident
What this role officially is
ICT system administrator — ESCO, the EU occupation classification
ICT system administrators are responsible for the upkeep, configuration, and reliable operation of computer and network systems, servers, workstations and peripheral devices. They may acquire, install, or upgrade computer components and software; automate routine tasks; write computer programs; troubleshoot; train and supervise staff; and provide technical support. They ensure optimum system integrity, security, backup and performance.
Also known as: enterprise administrator, IT system administrator, ICT systems administrator, ICT sysadmins, IT systems administrator, ICT sysadmin
How others title the same work
Large employers
- Senior Site Reliability / Gitops Engineer Canonical
- Senior Site Reliability / Gitops Engineer Canonical Ltd.
- Senior Site Reliability Engineer Canonical Ltd.
- Site Reliability / Gitops Engineer Canonical Ltd.
- Site Reliability Engineer Canonical Ltd.
Startups
- Senior Software Engineer, Product Infrastructure Atlas
- Site Reliability Engineer Beam
What the listing never says
- 32 bullet points. Long requirement lists deter qualified candidates, who read them as hard gates. Scope clarity
- No pay range published. Candidates cannot tell whether applying is worth their time. Pay transparency
The listing, marked up
Nothing in the wording of this listing tripped a check. The scores above still judge how complete and coherent it is.
Our Story
At Endear, we’re building a modern CRM for retail teams—starting with the frontline. Our software helps sales associates have more personal, effective customer conversations through AI-powered tools that drive measurable revenue.
Despite retail being significantly larger than eCommerce, most software overlooks the in-store experience. Endear helps brands turn real customer relationships into growth through intuitive software and thoughtful design.
As Endear grows, we are investing in the reliability and platform systems that keep our product fast, resilient, and easy for engineering teams to operate.
Position Overview
We’re hiring a Senior Site Reliability Engineer to become Endear’s first dedicated reliability hire. This is a hands-on builder role for someone who wants to solve the underlying systems problems that create on-call burden—not simply respond to pages.
You will own the work of making Endear’s systems more observable, reliable, and scalable. You’ll investigate recurring incidents, drive root-cause fixes, improve alerting and incident playbooks, and partner with engineers on database, queue, and event-processing reliability.
You will also help establish a nearshore triage layer for routine, well-understood issues, so product engineers can spend more time building product and less time on operational interruptions.
What You’ll Accomplish
In your first 6 months, you’ll…
• Build a clear view of Endear’s highest-impact reliability risks, recurring incidents, and on-call pain points.
• Establish a prioritized reliability backlog and drive root-cause fixes for the most important issues.
• Improve alert quality, severity definitions, escalation paths, and runbooks for common incidents.
• Strengthen observability, queue health, database capacity planning, and operational readiness ahead of peak retail periods.
• Create the foundation for a nearshore triage process for low-priority, repeatable issues.
In your first year, you’ll…
• Make on-call materially quieter and less disruptive for product engineers.
• Build scalable systems for observability, alerting, incident response, database reliability, and queue/event-processing health.
• Own and improve the nearshore triage relationship, playbooks, and escalation process.
• Help establish the technical roadmap and future resourcing plan for Endear’s broader platform and reliability function.
You’ll Thrive in This Role If You…
• Have deep hands-on experience with Kubernetes, production databases, and event-driven systems.
• Have operated and improved high-volume production systems with meaningful reliability, performance, and data-scale requirements.
• Enjoy finding root causes, fixing repeat incidents, and building tooling that makes engineers’ lives easier.
• Have experience with observability, alerting, incident response, capacity planning, and operational runbooks.
• Can work effectively as a senior IC: owning complex technical work directly while coordinating across teams.
• Are comfortable in a lean environment where priorities move quickly and you will need to make practical trade-offs.
• Bring experience from a scaling, mid-size company rather than only an early-stage startup or hyperscaler environment.
• Have GCP or ClickHouse experience, which are strong pluses.
About the Team
You’ll partner with:
• JP Grace, CTO: Align on reliability priorities, technical risks, and the roadmap for improving on-call health.
• Engineering team: Partner on root-cause fixes, platform improvements, and architecture decisions that improve reliability.
• Nearshore triage partner: Build and maintain playbooks, escalation paths, and expectations for routine incident handling.
• Product and Support: Help ensure issues are surfaced, prioritized, and resolved with the right level of urgency.
Endear is a lean, remote team where individuals have broad ownership. This role will directly shape how the company handles production reliability as it grows.
Our Hiring Process
• Recruiter screen — 30 minutes
• Behavioral interview with CTO — 60 minutes
• Technical panel with Engineering — 60 minutes
• Final conversation with Co-Founders
• Offer 🎉
Compensation & Benefits
• Base salary: $140,000-180,000
• Fully remote, U.S.-based role
• Comprehensive healthcare, including medical, dental, and vision, plus a 401(k) plan
• Monthly stipend for co-working and home-office setup
• Flexible PTO and unlimited vacation
• Opportunity to build Endear’s first dedicated reliability function from the ground up
Apply Even If You Don’t Check Every Box
If this role excites you but you’re unsure if you meet every requirement, reach out anyway. We care about skills, motivation, and how you think more than perfect resumes.
Originally posted on Himalayas
At Endear, we’re building a modern CRM for retail teams—starting with the frontline. Our software helps sales associates have more personal, effective customer conversations through AI-powered tools that drive measurable revenue.
Despite retail being significantly larger than eCommerce, most software overlooks the in-store experience. Endear helps brands turn real customer relationships into growth through intuitive software and thoughtful design.
As Endear grows, we are investing in the reliability and platform systems that keep our product fast, resilient, and easy for engineering teams to operate.
Position Overview
We’re hiring a Senior Site Reliability Engineer to become Endear’s first dedicated reliability hire. This is a hands-on builder role for someone who wants to solve the underlying systems problems that create on-call burden—not simply respond to pages.
You will own the work of making Endear’s systems more observable, reliable, and scalable. You’ll investigate recurring incidents, drive root-cause fixes, improve alerting and incident playbooks, and partner with engineers on database, queue, and event-processing reliability.
You will also help establish a nearshore triage layer for routine, well-understood issues, so product engineers can spend more time building product and less time on operational interruptions.
What You’ll Accomplish
In your first 6 months, you’ll…
• Build a clear view of Endear’s highest-impact reliability risks, recurring incidents, and on-call pain points.
• Establish a prioritized reliability backlog and drive root-cause fixes for the most important issues.
• Improve alert quality, severity definitions, escalation paths, and runbooks for common incidents.
• Strengthen observability, queue health, database capacity planning, and operational readiness ahead of peak retail periods.
• Create the foundation for a nearshore triage process for low-priority, repeatable issues.
In your first year, you’ll…
• Make on-call materially quieter and less disruptive for product engineers.
• Build scalable systems for observability, alerting, incident response, database reliability, and queue/event-processing health.
• Own and improve the nearshore triage relationship, playbooks, and escalation process.
• Help establish the technical roadmap and future resourcing plan for Endear’s broader platform and reliability function.
You’ll Thrive in This Role If You…
• Have deep hands-on experience with Kubernetes, production databases, and event-driven systems.
• Have operated and improved high-volume production systems with meaningful reliability, performance, and data-scale requirements.
• Enjoy finding root causes, fixing repeat incidents, and building tooling that makes engineers’ lives easier.
• Have experience with observability, alerting, incident response, capacity planning, and operational runbooks.
• Can work effectively as a senior IC: owning complex technical work directly while coordinating across teams.
• Are comfortable in a lean environment where priorities move quickly and you will need to make practical trade-offs.
• Bring experience from a scaling, mid-size company rather than only an early-stage startup or hyperscaler environment.
• Have GCP or ClickHouse experience, which are strong pluses.
About the Team
You’ll partner with:
• JP Grace, CTO: Align on reliability priorities, technical risks, and the roadmap for improving on-call health.
• Engineering team: Partner on root-cause fixes, platform improvements, and architecture decisions that improve reliability.
• Nearshore triage partner: Build and maintain playbooks, escalation paths, and expectations for routine incident handling.
• Product and Support: Help ensure issues are surfaced, prioritized, and resolved with the right level of urgency.
Endear is a lean, remote team where individuals have broad ownership. This role will directly shape how the company handles production reliability as it grows.
Our Hiring Process
• Recruiter screen — 30 minutes
• Behavioral interview with CTO — 60 minutes
• Technical panel with Engineering — 60 minutes
• Final conversation with Co-Founders
• Offer 🎉
Compensation & Benefits
• Base salary: $140,000-180,000
• Fully remote, U.S.-based role
• Comprehensive healthcare, including medical, dental, and vision, plus a 401(k) plan
• Monthly stipend for co-working and home-office setup
• Flexible PTO and unlimited vacation
• Opportunity to build Endear’s first dedicated reliability function from the ground up
Apply Even If You Don’t Check Every Box
If this role excites you but you’re unsure if you meet every requirement, reach out anyway. We care about skills, motivation, and how you think more than perfect resumes.
Originally posted on Himalayas