Staff Site Reliability Engineer
- Remote job
This is a hands-on senior engineering role focused on improving production resilience, strengthening security, driving operational excellence, and enhancing the developer experience across the organization.
In this role, you will design, build, and evolve the foundational systems, tooling, and operational practices that enable engineering teams to ship secure, reliable, and scalable software with confidence. You will help establish reliability standards, define service level objectives (SLOs), improve observability, automate operational processes, and drive incident management and post-incident learning practices that strengthen platform stability over time.
Partnering closely with Engineering, Security, Platform, and Product teams, you will architect scalable distributed systems, optimize Kubernetes and AWS-based infrastructure, and build automated delivery pipelines that support rapid and safe software releases. You will play a key role in reducing operational toil, improving system performance, increasing platform reliability, and ensuring that our infrastructure can support continued business growth.
❗ This is a full-time permanent position
❗ This is an existing vacancy
Location: This is a remote location open to candidates legally authorized to work in Canada.
What you will be doing: Drive reliability engineering initiatives and operational excellence for mission-critical services running on AWS and Kubernetes. Design, implement, and continuously improve deployment, release, and rollback strategies across complex distributed systems. Establish secure-by-default CI/CD pipelines with robust automation, governance, and policy-driven controls. Enhance platform observability through metrics, logs, tracing, and actionable alerting to improve system visibility and operational efficiency. Define, implement, and mature Service Level Indicators (SLIs), Service Level Objectives (SLOs), and reliability standards across the organization. Lead response efforts for high-severity incidents, ensuring timely resolution, effective communication, and meaningful post-incident reviews that drive continuous improvement. Partner closely with engineering teams to strengthen platform standards, improve service resilience, optimize runtime performance, and embed reliability best practices. Mentor and guide engineers on cloud-native technologies, site reliability engineering principles, and operational excellence practices, fostering a culture of continuous learning and accountability. What you will bring: 8+ years of experience in Site Reliability Engineering (SRE), Platform Engineering, DevOps, or related cloud-native engineering roles. Deep expertise in AWS services, including EKS, IAM, VPC, Lambda, CloudFront, S3, and cloud networking/security best practices. Advanced experience operating and scaling production Kubernetes environments. Strong hands-on experience with Istio service mesh, including traffic management, security, observability, and resiliency. Proven expertise with Infrastructure as Code (IaC), preferably using AWS CDK. Experience building and managing CI/CD pipelines using GitHub Actions or similar platforms. Strong troubleshooting, performance optimization, and incident management experience in distributed systems. Excellent communication, collaboration, and technical leadership skills. Observability & Reliability Experience designing and operating monitoring, logging, tracing, and alerting solutions for cloud-native platforms. Strong knowledge of AWS CloudWatch, OpenTelemetry, AWS X-Ray, and Kubernetes observability tooling. Experience defining and operationalizing SLIs, SLOs, alerting strategies, runbooks, and reliability metrics. Proven ability to leverage observability data to improve service reliability, reduce incident impact, and optimize operational performance. Software Engineering & Platform Development Strong proficiency in TypeScript and Node.js for platform engineering, automation, and operational tooling. Experience building and maintaining scalable backend services, APIs, and event-driven systems. Deep understanding of Kubernetes architecture, controllers, Gateway API, ingress management, and service networking. Experience implementing zero-trust architectures, mTLS, and service-to-service security controls. Commitment to high-quality engineering practices, including automated testing, code reviews, and observability-driven development. Strong understanding of resilience engineering, including autoscaling, disruption management, failure testing, and safe deployment strategies. Nice to Have Experience with progressive delivery practices such as canary, blue/green, and feature-flag-based deployments. Experience working in regulated, compliance-driven, or security-sensitive SaaS environments. Familiarity with FinOps principles and cost optimization strategies for cloud platforms. Experience building internal developer platforms and self-service engineering tooling. Cloud-native certifications such as CKA, CKAD, CKS, KCSA, or KCNA. Kubestronaut certification or equivalent advanced Kubernetes expertise is highly regarded.- As a Senior Ground Segment Engineer, you play a pivotal role in building software to operate our fleet of in-orbit satellites, and to prepare... ...engineering , with deep hands-on networking: routing, VPN/site-to-site, segmentation, DNS, firewalling Hands-on Software-Defined...SuggestedLong term contractTemporary workWork at officeFlexible hours
- ...audience, human or otherwise. The Role: Acquia is seeking a Staff AI Engineer to join our AI Core Engineering team. This is first and... ...Temporal, Pydantic and LangFuse are your primary tools; enterprise reliability, observability, and scale are your standards. You'll also...SuggestedRemote jobLocal area
$80k - $95k per year
...Running. Improve What's Next. We're partnering confidentially with a leading Canadian food manufacturer seeking a Mechanical Reliability Engineer to play a key role in improving equipment reliability, maintenance strategy, and plant performance within a highly automated manufacturing...SuggestedLong term contractPermanent employmentFull timeFor contractorsRelocation package- ...flow expansion, we sit at the intersection of product/platform engineers and financial partners, connecting them to ensure that everyone... ...ensure that data flowing through our partner and internal systems is reliable and correct Collaborate across the company, including...Suggested
$150k - $210k per year
...moments to travelers around the world. We’re looking for a Staff Software Engineer to join the Checkfront team. Checkfront is an established... ...maintaining and improving the product with work spanning reliability, performance, retention, maintainability, and new feature development...SuggestedFull time- ...quality considerations. You will work with many partner teams: peer engineering organizations to build robust/reusable/standardized access... ...ll directly build well-tested APIs, UX/UI pieces, and analytics engines for defining and measuring the full-stack Dashboard...Full time
$96k - $113k per year
...People like you! Job Description We are seeking a dedicated Reliability Engineer to focus on long-term monthly and quarterly issues. This role... ...we obtain when you apply for a position through this career site. If you do not consent to the terms of this Privacy Notice, please...Long term contractTemporary workWorldwide$186.37k - $223.64k per year
...NVIDIA, Microsoft, and Salesforce – trust Grafana Labs to ensure reliability of their applications and systems, resolve incidents quickly... ...for what could be a truly career-defining opportunity. Staff Backend Engineer - Mimir Query, Databases This is a remote position. We...Remote jobLocal areaFlexible hours$98.6k - $147.9k per year
...a culture of safety and accountability. Join Us in Leading Safe and Reliable Surface Maintenance Operations! Role Overview Reporting to the Reliability Superintendent, the Reliability Engineer plays a fundamental role within the Fixed Plant Maintenance Department....Long term contractTemporary workInternshipShift work$147.6k - $184.5k per year
...Coursera achieve a seamless experience across its platform. As a Staff Product Designer, you will lead the design strategy and... ...for a major product area. You'll partner closely with Product, Engineering, Research, and Data Science across multiple teams to shape scalable...InternshipWork at officeWorldwide$130k - $150k per year
...future of energy with us. POSITION SUMMARY The role of Site Manager for Project Management directly supports the company’s mission... ...and tracks all the quality control, transportation, and engineering documents for the project. Prepares daily and weekly job progress...Long term contractFor contractorsFor subcontractorWork at office$39 - $43 per hour
...future of energy with us. POSITION SUMMARY The role of the Site Quality Technician directly supports the company’s mission to... ...Nonconformities by updating with new information, communicating with Nordex Engineering, Repair teams and other involved with nonconformity. Other...Hourly payLong term contractWeekend work$90.7k - $113.4k per year
...heard of. We started as a simple shareware site in 1993 and have since grown into a stable... ..., and automation-minded Data Platform Engineer to join our Data Engineering team. This is... ...engineering role for someone who enjoys building reliable systems, writing maintainable automation,...Remote work$145k - $155k per year
...We are looking for an experienced backend engineer who enjoys solving complex problems, sharing knowledge, and building reliable software. As a Senior Backend Software Engineer... ...Learn more about Tucows, our businesses, culture and employee benefits on our site here ....Remote work- ...Systems Inc. is looking for a SysAdmin DevOps Engineer for a full-time position (40 hours per... ...practices to optimize efficiency and reliability. Maintain documentation and providing... ...a result of seamless integration with on-site processes. Svitla Systems' global mission...Full timeRemote workWorldwide
$103k - $140k per year
...a Senior Developer Operations Engineer based in Canada. This is a... ...ownership over infrastructure reliability, operational efficiency, deployment... ...focus on Google Kubernetes Engine (GKE). Develop, maintain,... ...DevOps, Developer Operations, Site Reliability Engineering, or a...Remote jobFull timeWorldwideFlexible hours- ...our people to be their best. As a Data Engineer, you’ll help expand, optimize, and maintain... ...environments. Ensure data quality, reliability, and observability through robust monitoring... ...ensure training data quality, and build reliable data feeds for analytical and predictive...Remote jobLong term contractInternshipFlexible hours
- ...nearly 30 years, our expert teams have helped our clients deliver reliable products that users enjoy interacting with. Our 100% Canada‑... ...of application Summary We are seeking an experienced AI Engineer with 5+ years of experience designing, developing, and deploying...Contract workInternship
$110k - $120k per year
...intensive industries, where performance, reliability, and precision are essential. With a strong... ...the world. Their success is built on engineering excellence, practical innovation, and a... ...Abbotsford, BC. The first three months are on-site to support onboarding, product training,...Permanent employmentFull timeRelocation package$68.81k - $114.71k per year
...Europe, Australasia and Africa. Working for us, you’ll help ensure that Canada’s critical services and assets are readily available, reliable, and capable for our defence and civil customers and contribute to Babcock’s purpose: “Creating a safe and secure world, together.” Our...ApprenticeshipInternshipMonday to fridayShift workDay shift- ...Impact: As a Senior Frontend Engineer on the squad, you'll build... ...directly through the booking engine. Our Booking Engine Team:... ...revenue-critical path, so speed, reliability, and getting the details right... ...Agencies: Our Careers Site is only for individuals seeking...Work at officeLocal areaRemote workWork from homeWorldwideHome officeWeekend work
- ...an Impact: As a Senior Backend Engineer on the Booking Engine squad, you'll implement new architectures... ...the revenue-critical path, so speed, reliability, and getting the details right... ...and Recruiting Agencies: Our Careers Site is only for individuals seeking a job...Long term contractWork at officeLocal areaRemote workWork from homeWorldwideHome officeWeekend work
- ...career. About the team The Data Foundations team drives Data Engineering and Data Apps and Tooling work across Stripe, enabling Stripes... ...forecast the future potential performance of the business and reliably measure ongoing attainment toward targets Build data...Internship
$116k - $159.5k per year
...advertising and marketing channels. The Engineering Operations team builds and operates the... ...on infrastructure as code, operational reliability, and scalable platform practices across... ...with Terraform and Terrateam. Support reliable, secure deployments across cloud and...Local areaRemote workWork from homeHome office$180k - $220k per year
...Principal Software Engineer Location: Remote (Anywhere in Canada) Company Overview... ...for engineering quality, performance, and reliability Identify and prevent overengineering,... ...applications ~ Proven experience operating at a staff or principal level, influencing multiple...Long term contractFull timeRemote work- ...We are seeking a highly skilled Senior Application Engineer with deep T-SQL expertise to join the Liquid Credit Engineering Team, a... ...critical production issues in a timely manner to maintain system reliability and minimize business impact T-SQL Expertise: Strong hands-on experience...Remote job
- Lone Wolf is looking to hire a Software Engineering Manager to lead and manage a small, focused... ...clients to sign documents securely, reliably, and with full legal validity. You will... ...intuitive, easy to maintain, responsive, reliable, scalable, and performant. You will play...Remote job
- ...nearly 30 years, our expert teams have helped our clients deliver reliable products that users enjoy interacting with. Our 100% Canada‑... ...for defining and leading the DevOps strategy, practices, and engineering standards that enable the successful delivery and operation of...Contract workInternship
$140k - $150k per year
...London, ON Area (On-Site) Full-Time | Permanent $140,000 – $150,000 + Bonus + Benefits + Pension Lead the Engineering Team. Shape Canada's Infrastructure. We're partnering confidentially with a respected Canadian infrastructure manufacturer seeking an...Long term contractPermanent employmentFull timeFor contractors$100k - $135k per year
...Intermediate Software Engineer Location: Remote (Anywhere in Canada) Company Overview eDynamic Learning is celebrating 18 years... ...teams, and contribute to improving code quality, system reliability, and engineering practices. You will partner with senior engineers...Long term contractFull timeRemote work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!
- assistant ingénieur Canada
- project engineer assistant project manager Canada
- stage assistant ingénieur Canada
- assistant electrical engineer Canada
- site reliability engineer intern Canada
- site reliability engineer Canada
- senior site reliability engineer Canada
- site reliability engineer remote Canada
- site safety Canada
- site carpenter Canada


