Cloud Performance Engineering - Site Reliability Engineer
$110k - $125k per yearSmile Digital Health
Working for a company like Smile Digital Health means supporting our mandate for #BetterGlobalHealth . We strive towards this goal every day, and the results can be seen in the impact of our innovative health data platform and data management solutions, which are used in over 20 countries. We were #19 on Deloitte's Technology Fast 50 Ranking for 2024!
Smile Digital Health makes it easy for healthcare stakeholders to collect and exchange data with our leading FHIR-based data liberation platform.
At its heart, the Smile platform enables people and organizations to better manage healthcare data. We help generate and liberate structured healthcare data to ensure effective delivery across care teams and health systems bringing #BetterGlobalHealth to patients everyday!
Apply today and find plenty of reasons to SMILE!
The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure platforms for Smile Digital Health, its clients, and partners.
This role designs and automates performance testing frameworks, integrates them into CI/CD pipelines, and uses observability tools to proactively detect and resolve bottlenecks. Working closely with engineering, product, and security teams, the SRE ensures systems meet strict SLAs for performance and availability while driving continuous optimization across multiple cloud platforms.
Responsibilities:
Collaborate with our Security Operations teams to help define and implement best practices around Cloud Service Provider configuration for Azure and other cloud providers.
Develop, implement and coordinate a multi-tenant approach around service offerings for DB, Container platform, Authentication, Certificates, and Product Registries etc.
Design and maintain performance testing strategies, framework, and environments in the cloud.
Develop and maintain cost/utilization tracking and attribution processes for all Cloud Service Providers.
Create documentation around Cloud Service Provider offerings detailing use cases, best practices, and implementation details.
Develop and maintain technical relationships with our core Cloud Service Providers.
Implement and maintain a secure and scalable infrastructure platform for delivering Cloud Services applications.
Ensure that internal and external SLA’s meet and exceed expectations, and ensure that system centric KPIs are continuously monitored and improved.
Create tools for automating deployment, monitoring and operations of the overall platform.
Participate in an on-call rotation to provide application support, incident management, and troubleshooting.
Provide ongoing maintenance and support of internal tools, improve system health and reliability.
Assist customers with the on-site deployments when needed.
Implement and manage observability tools (logging, metrics, tracing) for performance insights, Otel and Grafana Stack preferred
Requirements:
Demonstrated expertise in cloud service providers and best practices around implementation and configuration, preferably managing Azure on behalf of multiple teams for a company that delivers SaaS products.
Proven experience working with microservices architecture, with a strong focus on Java-based services.
Experience in applying chaos engineering practices to evaluate and enhance system resiliency.
Skilled in troubleshooting performance issues, including analyzing time consumption, allocating resources, and recommending optimizations.
Familiar with performance testing methodologies and tools to assess system behavior under load.
Experience with deployment and usage of observability tools such as Prometheus and the Grafana suite.
Proven experience designing and executing performance test plans (load, stress, soak, and spike testing) to validate that application services sustain 500+ transactions per second (TPS) within defined latency and error-rate thresholds.
Hands-on experience with performance/load testing tools such as JMeter, Gatling, Azure Load Testing
Experience tuning and validating autoscaling (Kubernetes/OpenShift HPA, Azure scale sets) to ensure required TPS is met under variable load without breaching cost or resource constraints
Experience tuning Kafka (partitioning, consumer group sizing, throughput/latency trade-offs) and other messaging/queueing components to sustain target transaction rates.
Experience with Azure-native monitoring and diagnostics (Azure Monitor, Application Insights, Log Analytics) to correlate throughput, latency, and error metrics during test execution.
Proven experience with Security and Compliance (SOC2, HIPAA, ISO27001) best practices and how to implement controls that support high-velocity software delivery teams.
Proficiency in Terraform, Ansible or Chef.
Expertise in troubleshooting, support escalation, on-call process optimization and documenting knowledge.
Passionate about Infrastructure as code, automation, and developing solutions that help developers move quickly and safely.
Familiarity with infrastructure management and operations lifecycle concepts and ecosystem.
Experience operating and maintaining production systems in a Linux and public cloud environment.
You have prior experience working in high-performance or distributed systems, while we strive to hire at a variety of experience levels.
Working knowledge of industry best practices regarding information security
Previous experience building or maintaining a large-scale Cloud service.
Proven ability to prioritize and track multiple projects in parallel.
$110,000 - $125,000 a year
Smile discloses that artificial intelligence (AI) may be used in portions of the recruitment and selection process, such as resume screening or application assessment. All hiring decisions are ultimately made by qualified human decision-makers, and AI tools are used to support — not replace — fair and equitable hiring practices.
This position is a new role, created to support Smile’s continued growth and commitment to operational excellence.
Some of the benefits we offer:
* Remote Work Environment
* Flexible Time Away From Work Policy including PTO, Personal and Sick Days
* Competitive Salary and Health/Medical Benefits
* RRSP/TFSA/401K Employee Contribution
* Life and Disability
* Employee Assistance Program
* FHIR Study Program and Skillsoft Learning
* Super HAPI Fun Club
Smile's core values include respect, inclusion, embracing our differences, and celebrating shared values because our people are the foundation of our success. We are big on creating a sense of belonging and empowering each other to bring our authentic selves to work. We are dedicated to fostering a workplace that values diversity, equity, and inclusion.
We welcome and encourage candidates of all backgrounds to apply. Candidates are encouraged to inform us if they wish to discuss or require accommodations during interviews or while working at Smile.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
$115k - $130k per year
...innovation, accountability, and results. We are a high-growth, high-performance organization that values individuals who are driven,... ...and continuously raising the bar. Kaseya is hiring a Site Reliability Engineer to keep our production systems healthy as we scale. You'll...PerformanceLong term contractFull timeWorldwide$110k - $120k per year
...Discretionary Annual Bonus – Rewarding both individual and company performance Comprehensive Benefits – Health and dental coverage with... ...programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring...PerformanceFull timeTemporary workInternshipWork at officeRemote work- ...Generative AI inference solution in the world, over 10 times faster than GPU-based hyperscale cloud inference services. About The Role Join Cerebras as a Performance & Reliability Engineer within our innovative Co-Design and Next Generation Team. Our groundbreaking CS-3...PerformanceFull time
- ...good fit for you! About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a... ...will help maintain, optimize, and ensure the reliability and performance of the systems that power our cloud infrastructure across...PerformanceRemote jobLong term contractPermanent employmentFull time
$100k - $125k per year
...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer... ...will play a crucial role in enhancing the reliability, performance, and scalability of our systems and services. You will...PerformanceFull timeWork at officeFlexible hours- ...Description The SRE Role · SREs are engineers with the right mix of knowledge and... ...observation to entire systems to improve reliability, performance and operability). · We constantly evaluate... ...such as puppet, chef or ansible. · Cloud technologies and platforms such as AWS...PerformanceFull time
$154k - $200k per year
...America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with... ...the design and evolution of major systems within our multi-cloud, multi-region, active-active content serving platform that serves...PerformanceLong term contractFull time- ...About the Role Fivetran is looking for a high-performance engineer to be a part of a team of Site Reliability Engineers. You will be working closely with engineering teams, product managers, as well as support and sales engineers to build the future of the Fivetran Data...PerformanceFull timeWork at officeRemote work
- ...the world and our goal is preserving uncensored Internet access and online privacy for all. Right now we are looking for a Site Reliability Engineer to help us tame DNS. This position can be hybrid at our Toronto offices or remote in Canada only but you MUST reside in the...Full timeDirect hireWork at officeRemote work
$140k - $182k per year
...America, Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and... ...help inform the evolution of major systems within our multi-cloud, multi-region, active-active content serving platform that serves...PerformanceFull time- ...because learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident... ...structured reliability work such as capacity forecasting, performance tuning, and controlled failure testing. Improve how the...PerformanceFull timeFor contractorsWork at officeWorldwide3 days per week
- ...learning never stops. The Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team of engineers dedicated... .... By guiding this team, you will actively evolve our legacy Cloud Operations & Support function into a modern, world-class SRE...PerformanceLong term contractFull timeFor contractorsWork at officeWorldwide3 days per week
- ...belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a... ...and organization. As a Senior SRE, you’ll help scale our cloud platform, collaborate across teams to promote...PerformanceFull timeWork at officeLocal areaRemote workWorldwideMonday to fridayFlexible hours
- ...and how we use AI in our recruiting process here . The Site Reliability Engineering organization at Pinterest is accountable for ensuring... ...and maximize the speed of innovation Manage capacity and performance to help scale our infrastructure both on public and private...PerformanceFull timeWork at office
- ...tenant SaaS platform running reliably for customers around the world... ...bar across the wider engineering org. What You'll Be Doing... ...such as capacity forecasting, performance tuning, and controlled failure... ..., containerized systems, and cloud native infrastructure, including...PerformanceFull timeImmediate startFlexible hours
$100k per year
...on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency.... ...deployments. This role sits at the intersection of site reliability, infrastructure operations, and customer engineering, ensuring our systems are reliable, observable,...PerformancePermanent employmentFull time- ...increase visibility, and optimize spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s systems are highly available, scalable, and performant. This role blends deep technical expertise with strong collaboration...PerformanceFull timeManual laborLocal areaFlexible hours
- ...Cohere is a team of researchers, engineers, designers, and more, who are passionate... ...Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want... ...NLP applications? We are looking for a Site Reliability Engineer to join the Model...PerformanceFull timeWork at officeRemote workFlexible hours
$243k - $297k per year
...resilient businesses. As Relay continues to scale, the reliability, performance, and resilience of our platform are no longer just... ...role responsible not only for guiding a strong team of Site Reliability Engineers, but for shaping how reliability strategy influences engineering...PerformanceLong term contractFull timeInternshipWork at officeTrial periodFlexible hours$130k - $180k per year
...work in Canada (visa or sponsorship won't be provided) Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure (those with Azure will be prioritized first) About Us: We're one...PerformanceFull timeRemote workVisa sponsorshipWork visaFlexible hours$136k - $187k per year
...users worldwide. Our commitment to reliability is a key foundation of our product... ...availability expectations is a core engineering focus. As a Senior Site Reliability Engineer, you'll join our... ..., improving the availability, performance, and observability of our services....PerformanceLocal areaRemote workWorldwide- ...Senior Site Reliability Engineer - Edge Location : Ottawa/Toronto, On-Site Reports to: Head of Security The Role You own the edge... ...embedded/SBC experience in real operational environments, not just cloud VMs. Real DDIL experience — built systems that operate...Full time
$144k - $200k per year
**The Team** Platform Engineering is the department within SRE that is... ...organization. Among these are our multi-cloud-provider Kubernetes... ...developing and maintaining the reliable and globally connected multi-cloud... ...* We are seeking a talented Site Reliability Engineer (SRE)...Full timeWork at officeRemote workWorldwideFlexible hours$104.24k - $143.3k per year
.... The role will provide an opportunity to work with Microsoft Azure cloud technologies in a team that adopts DevOps methodologies at an innovative and growing company. Our Lead Site Reliability Engineers will provide a stable infrastructure platform throughout our clients...Full timeWork at officeImmediate startRemote workShift workWeekend work2 days per week- ...Group, a division of Fox Corporation. About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are... ...a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through...PerformanceRemplacementFull timeContract workTemporary workFlexible hours
- ...and capabilities in others. About the Role: As a Site Reliability Engineer , you’ll join the global Platform SRE team responsible for... ...scale , automating operations, and continuously improving performance, resilience, and deployment pipelines. What You’ll Do:...PerformanceFull time
$150k - $250k per year
...About The Role We're looking for a Senior Site Reliability Engineer to help us run one of the most exciting GPU clusters around—our Toronto... ...That means troubleshooting issues as they arise, monitoring performance, developing automation to make our lives easier, and working...PerformanceFull time- ...Platform Engineer – DevOps, Site Reliability Engineering (SRE) & Dynatrace Required Skills • Strong experience as a Platform Engineer with... ...of high availability, reliability, scalability, and performance engineering for mission-critical applications. • Hands...PerformancePermanent employment
$141k - $191k per year
...experience in Service Management, working with cloud providers, software development, and... ...Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational... ...the Role: In this opportunity as Site Reliability Engineering Manager , you will be...PerformanceWork at officeLocal areaFlexible hours2 days per week3 days per week- ...Site Reliability Engineer Location: Toronto ON Work Model: Hybrid (2 days per week in-person at the Toronto office preferred) Required... ...) and DevOps practices to ensure high system availability performance and reliability across distributed environments....PerformanceFull timeWork at office2 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Cloud Performance Engineering - Site Reliability Engineer. Be the first to apply!
- site reliability engineer intern Toronto, ON
- site reliability engineer remote Toronto, ON
- site reliability engineer Toronto, ON
- senior site reliability engineer Toronto, ON
- building performance specialist Toronto, ON
- performance engineer Toronto, ON
- acting performance Toronto, ON
- performance testing Toronto, ON
- website developer Toronto, ON
- site safety Toronto, ON
