This one is closed
Live roles like this one
-
C
4h ago
Rapinno Tech United States
-
B
4h ago
Software Engineers (Generalists)
ICF United States $99k - $168k/yr
-
B
6h ago
Cybersecurity Engineering & Operations Director
Onit United States $150k - $190k/yr
-
A
6h ago
Nodeworthy Remote $4k - $6k/mo
See every "Cloud Performance Engineering" role →
Get new “Cloud Performance Engineering” roles by email
One email a day with what is new in "Cloud Performance Engineering". Nothing new, no email.
We confirm the address first, and every mail carries an unsubscribe link. Alerts are ours, not a third party's.
Why this grade This listing scored 30/100, which is an F. It lost the most ground on freshness. See the breakdown
- Description depth 20 / 20 How much the posting actually says about the work, measured in characters of real text.
- Pay transparency 12 / 25 A published salary range, worth more than any other single factor because it is what a candidate cannot find out without applying.
- Remote clarity 8 / 15 Whether "remote" means anywhere, or is quietly restricted to one country.
- Corroboration 5 / 10 Whether more than one source carries this listing.
- Role specificity 0 / 10 Whether the listing is tagged well enough to tell what the role actually is.
- Freshness 0 / 15 How recently it was posted. Older postings are likelier to be filled or abandoned.
-15 Ghost-job penalty — Deducted for signals that this posting may not be a real, currently-open role — staleness, repeated relisting, or talent-pool language.
Every figure above is arithmetic over the posting itself — its salary field, its text, its age, its tags and how many sources carry it. How the grades work →
Apply today and find plenty of reasons to SMILE!
The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure platforms for Smile Digital Health, its clients, and partners.
This role designs and automates performance testing frameworks, integrates them into CI/CD pipelines, and uses observability tools to proactively detect and resolve bottlenecks. Working closely with engineering, product, and security teams, the SRE ensures systems meet strict SLAs for performance and availability while driving continuous optimization across multiple cloud platforms.
Responsibilities:
Collaborate with our Security Operations teams to help define and implement best practices around Cloud Service Provider configuration for Azure and other cloud providers.
Develop, implement and coordinate a multi-tenant approach around service offerings for DB, Container platform, Authentication, Certificates, and Product Registries etc.
Design and maintain performance testing strategies, framework, and environments in the cloud.
Develop and maintain cost/utilization tracking and attribution processes for all Cloud Service Providers.
Create documentation around Cloud Service Provider offerings detailing use cases, best practices, and implementation details.
Develop and maintain technical relationships with our core Cloud Service Providers.
Implement and maintain a secure and scalable infrastructure platform for delivering Cloud Services applications.
Ensure that internal and external SLA’s meet and exceed expectations, and ensure that system centric KPIs are continuously monitored and improved.
Create tools for automating deployment, monitoring and operations of the overall platform.
Participate in an on-call rotation to provide application support, incident management, and troubleshooting.
Provide ongoing maintenance and support of internal tools, improve system health and reliability.
Assist customers with the on-site deployments when needed.
Implement and manage observability tools (logging, metrics, tracing) for performance insights, Otel and Grafana Stack preferred
Requirements:
Demonstrated expertise in cloud service providers and best practices around implementation and configuration, preferably managing Azure on behalf of multiple teams for a company that delivers SaaS products.
Proven experience working with microservices architecture, with a strong focus on Java-based services.
Experience in applying chaos engineering practices to evaluate and enhance system resiliency.
Skilled in troubleshooting performance issues, including analyzing time consumption, allocating resources, and recommending optimizations.
Familiar with performance testing methodologies and tools to assess system behavior under load.
Experience with deployment and usage of observability tools such as Prometheus and the Grafana suite.
Proven experience designing and executing performance test plans (load, stress, soak, and spike testing) to validate that application services sustain 500+ transactions per second (TPS) within defined latency and error-rate thresholds.
Hands-on experience with performance/load testing tools such as JMeter, Gatling, Azure Load Testing
Experience tuning and validating autoscaling (Kubernetes/OpenShift HPA, Azure scale sets) to ensure required TPS is met under variable load without breaching cost or resource constraints
Experience tuning Kafka (partitioning, consumer group sizing, throughput/latency trade-offs) and other messaging/queueing components to sustain target transaction rates.
Experience with Azure-native monitoring and diagnostics (Azure Monitor, Application Insights, Log Analytics) to correlate throughput, latency, and error metrics during test execution.
Proven experience with Security and Compliance (SOC2, HIPAA, ISO27001) best practices and how to implement controls that support high-velocity software delivery teams.
Proficiency in Terraform, Ansible or Chef.
Expertise in troubleshooting, support escalation, on-call process optimization and documenting knowledge.
Passionate about Infrastructure as code, automation, and developing solutions that help developers move quickly and safely.
Familiarity with infrastructure management and operations lifecycle concepts and ecosystem.
Experience operating and maintaining production systems in a Linux and public cloud environment.
You have prior experience working in high-performance or distributed systems, while we strive to hire at a variety of experience levels.
Working knowledge of industry best practices regarding information security
Previous experience building or maintaining a large-scale Cloud service.
Proven ability to prioritize and track multiple projects in parallel.
Some of the benefits we offer:
* Remote Work Environment* Flexible Time Away From Work Policy including PTO, Personal and Sick Days* Competitive Salary and Health/Medical Benefits* RRSP/TFSA/401K Employee Contribution* Life and Disability* Employee Assistance Program* FHIR Study Program and Skillsoft Learning* Super HAPI Fun ClubSmile's core values include respect, inclusion, embracing our differences, and celebrating shared values because our people are the foundation of our success. We are big on creating a sense of belonging and empowering each other to bring our authentic selves to work. We are dedicated to fostering a workplace that values diversity, equity, and inclusion.We welcome and encourage candidates of all backgrounds to apply. Candidates are encouraged to inform us if they wish to discuss or require accommodations during interviews or while working at Smile.Originally posted on Himalayas
Apply for this role Opens himalayas.app — the link as listed; we have not yet verified it is the employer's own page
Quick question · anonymous · one tap
Would you apply to this job?
Answer to see what other job seekers said.
Your turn · no account needed
Help the next applicant
You may know something about this listing that we cannot see from here. One tap. No account needed. Signed-in reports earn points once the evidence agrees with you.
I know what it pays
Sign in with Google to earn points for reports — 100 confirmed points buy a week of Early Access.
Where this listing came from
- 13 Jul 2026 Himalayas first sighting
Seen on 1 board over 0 days.