Staff Site Reliability Engineer
- Remote job
This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer based in Canada.
This is a high-impact reliability leadership role within a fully remote engineering organization operating globally.
You will be the first dedicated SRE, helping establish reliability practices across multiple engineering teams and critical production systems.
The role combines hands-on engineering with organization-wide influence, covering observability, incident response, operational readiness, and resilience.
You will work closely with engineering leadership, infrastructure specialists, architects, and product teams to make reliability measurable and actionable.
A major focus will be embedding SRE principles into engineering culture rather than simply owning individual services.
You will also help shape how AI is used for incident investigation, operational tooling, observability, and safe system operations.
The position offers substantial autonomy to define standards, coach engineers, and build practices that scale with the organization.
Define and implement SLIs and SLOs for critical production request paths, ensuring reliability objectives are visible, measurable, reviewed, and connected to engineering decisions.
Introduce and champion error budgets as a practical framework for balancing reliability investments with product and feature delivery.
Establish and maintain the reliability metrics used by engineering leadership to evaluate progress and identify areas requiring investment.
Strengthen the complete incident management lifecycle, including detection, response, communication, escalation, postmortems, and follow-up actions.
Improve alert quality, anomaly detection, escalation processes, and shared operational tooling in collaboration with infrastructure teams.
Lead reliability assessments for high-risk changes and new services, covering production readiness, capacity, failure modes, rollback strategies, and operational risks.
Introduce deliberate failure testing, game days, and chaos exercises to identify weaknesses and validate safe operational limits before incidents occur.
Work directly with engineering teams on complex reliability challenges through focused engagements, leaving behind stronger practices and clear ownership.
Coach Staff and Lead engineers to become reliability advocates within their respective teams and help establish distributed SRE ownership.
Develop lightweight, repeatable operational standards covering production readiness, on-call practices, runbooks, change safety, and service operability.
Partner with architects and technical leads to ensure reliability and failure tolerance are incorporated into system design rather than addressed after deployment.
Remain hands-on during production incidents and investigations, building tooling, dashboards, automation, and reference implementations where appropriate.
Promote effective use of AI for incident investigation, telemetry analysis, postmortem development, runbook creation, observability, and reliability tooling.
Help structure operational data, alerts, dashboards, and runbooks so that both engineers and AI agents can safely interpret and act on production signals.
Contribute production fixes and improvements directly through code and infrastructure changes rather than limiting the role to recommendations and reviews.
10+ years of engineering experience, including at least 3 years in SRE, production engineering, or a reliability-focused Staff Engineer role operating across multiple teams.
Demonstrated experience owning reliability at a platform or organizational level rather than only for an individual service.
Deep practical experience designing and implementing SLIs, SLOs, and error budgets, including successfully driving adoption across product and engineering teams.
Strong incident leadership experience, including managing high-severity, customer-facing incidents and leading effective postmortems that result in measurable improvements.
Advanced understanding of distributed-system failure modes, including database and cache saturation, cascading failures, retry storms, capacity constraints, graceful degradation, and load shedding.
Strong hands-on experience with Kubernetes, AWS, and modern observability platforms such as Datadog or comparable technologies.
Ability to read and write production code in Go, TypeScript, or a similar language, as well as work with infrastructure as code.
Demonstrated ability to influence teams without direct authority and successfully change engineering practices across an organization.
Strong coaching and mentoring skills, with evidence of developing engineers into effective reliability owners.
Exceptional written and verbal communication skills, with the ability to clearly communicate incidents, risks, technical trade-offs, and reliability priorities to both engineers and executives.
Strong preference for asynchronous, documented decision-making and clear technical communication.
Practical experience using AI tools for incident investigation, telemetry analysis, runbook and postmortem development, and engineering tooling.
Understanding of how operational data, alerts, dashboards, and runbooks should be structured to support safe AI-assisted diagnosis and operations.
Pragmatic approach to reliability, with the ability to balance operational risk, engineering investment, delivery speed, and business priorities.
Experience in fraud detection, identity, payments, or other real-time and adversarial environments is an asset.
Experience with multi-region architectures, cell-based architectures, or failure-isolation strategies is a plus.
Experience operating Elasticsearch, Redis, DynamoDB, or Kafka at scale and understanding their failure modes is beneficial.
Familiarity with FinOps and cloud infrastructure cost-versus-reliability trade-offs is an advantage.
Must be authorized to work from the hiring location; visa sponsorship is not provided.
Fully remote working environment.
Opportunity to become the first dedicated Site Reliability Engineer and establish organization-wide reliability practices.
High level of autonomy and direct influence over engineering standards, operational practices, and platform reliability.
Opportunity to work across multiple engineering teams and critical production systems.
Close collaboration with engineering leadership, architects, infrastructure teams, and technical leads.
Opportunity to shape AI-assisted reliability practices and the future of production operations.
Strong focus on professional growth, technical leadership, coaching, and knowledge sharing.
Inclusive, globally distributed engineering environment that values diverse perspectives and backgrounds.
For US-based employees, the stated cash compensation range is $177,000–$240,000 USD , with actual offers varying according to factors such as experience, skills, education, certifications, and market conditions. Compensation may differ for other hiring locations.
Remote work eligibility is subject to applicable regulatory and security requirements in the candidate's location.
Requirements
Benefits
$160k - $180k per year
...donor fundraising. Position Summary: We are seeking a Staff Software Engineer, Payments to provide technical leadership as Kindsight... ...payments partner, how the platform should scale, and how we build reliable systems for moving and reconciling real donor funds. This...SuggestedInternship$80k - $95k per year
...Running. Improve What's Next. We're partnering confidentially with a leading Canadian food manufacturer seeking a Mechanical Reliability Engineer to play a key role in improving equipment reliability, maintenance strategy, and plant performance within a highly automated manufacturing...SuggestedLong term contractPermanent employmentFull timeFor contractorsRelocation package$150k - $210k per year
...moments to travelers around the world. We’re looking for a Staff Software Engineer to join the Checkfront team. Checkfront is an established... ...maintaining and improving the product with work spanning reliability, performance, retention, maintainability, and new feature development...SuggestedFull time$148k - $185k per year
...Coursera achieve a seamless experience across its platform. As a Staff Product Designer, you will lead the design strategy and... ...for a major product area. You'll partner closely with Product, Engineering, Research, and Data Science across multiple teams to shape scalable...SuggestedFull timeInternshipWork at office$45 - $50 per hour
...future of energy with us. POSITION SUMMARY The role of Site Quality Lead directly supports the company’s mission to develop... ...diploma or equivalent (GED) required. ~ Degree in technical or engineering field, or equivalent work experience ~3+ years’ experience preferred...SuggestedHourly payLong term contractFor contractorsFor subcontractorWeekend work$130k - $150k per year
...future of energy with us. POSITION SUMMARY The role of Site Manager for Project Management directly supports the company’s mission... ...and tracks all the quality control, transportation, and engineering documents for the project. Prepares daily and weekly job progress...Long term contractFor contractorsFor subcontractorWork at office- ...RESPONSIBILITIES At SpryPoint we value collaborate work environments, automation, learning, and delivering value to our users. As a Software Engineer III at SpryPoint, you will be building and integrating interactive web applications, services, and apps that real people will...Summer workInternshipLocal areaRemote workHome officeFlexible hours
$115k - $120k per year
...energy with us. POSITION SUMMARY The role of Assistant Site Manager directly supports the company’s mission to develop and deliver... ...phase. WHAT YOU HAVE ~ Associate’s in mechanical engineering or related discipline, or equivalent related work experience accepted...Long term contractFor contractorsFor subcontractorWork at office- ...of our offices or hubs, or a co-working space near you. Data Engineering is unique at Coursera. Our team doesn't simply build reports on... ...Architect scalable data models and build efficient and reliable ETL pipelines to bring the data into our core data lake Design...Full timeRemote workWork from homeWorldwideFlexible hours
$120k - $150k per year
...are seeking a Senior Software Engineer, Payments to join a newly... ...You will work closely with our Staff Engineer and Director of... ...help establish the patterns for reliability, reconciliation, integration,... ...Payments partner. Design reliable transaction-processing flows...$100k - $120k per year
...Position Summary: We are seeking an Intermediate Software Engineer, Payments to join a newly forming Payments Engineering... ..., AdvRM, and Connect. You'll work alongside Senior and Staff engineers to build reliable backend services, integrate with our selected payments...$148.3k - $171k per year
...leads the industry with its low fees, advanced trading engine, strong security and reliability, and robust compliance framework. Learn more at . All roles... ...for designing, developing, testing, and operating reliable backend services that support Binance.US products and...Full timeLocal areaRemote workWork from home- ...Description Job Opportunity: Integration Senior Software Engineer Role Overview Job Title: Integration... ...– September 2027) Work Location: Remote (No on-site requirement) Security Clearance: Reliability Status (required prior to contract start) Language...Contract workRemote work
$110k - $120k per year
...intensive industries, where performance, reliability, and precision are essential. With a strong... ...the world. Their success is built on engineering excellence, practical innovation, and a... ...Abbotsford, BC. The first three months are on-site to support onboarding, product training,...Permanent employmentFull timeRelocation package$110k - $150k per year
...applications and next steps. Our partner is looking for a Structural Engineer - Project Manager based in Canada. This is a senior... ...consistent project delivery. Mentor junior engineers and technical staff while promoting technical quality, efficiency, knowledge sharing,...Remote jobLong term contractFor contractors$170.4k - $234.3k per year
...and next steps. Our partner is looking for a Senior Full Stack Engineer based in Canada. This role offers the opportunity to build and... ...developer ecosystems. ~ Work collaboratively across teams to develop reliable systems and improve how software is designed, tested, deployed,...Remote jobLong term contractWork from homeFlexible hours$5000 per month
...applications and next steps. Our partner is looking for a Lead Software Engineer (.NET) based in Canada. This role offers the opportunity to... ..., resolve complex backend and data challenges, and help ensure reliable delivery. You will collaborate closely with Product and...Remote jobFull timeHome office$114k - $172k per year
...workflow, and built in reporting. It runs in the cloud on a secure, reliable, extensively audited platform and integrates deeply with on... ...systems. We are looking for an experienced Senior UI Software Engineer to work on our Onboarding and Lifecycle Management (LCM) Platform...Full timeRemote workFlexible hours- ...nearly 30 years, our expert teams have helped our clients deliver reliable products that users enjoy interacting with. Our 100% Canada‑... ...Summary We are seeking an experienced Infrastructure Engineer to join our Infrastructure & DevOps team. This role is responsible...Contract workInternshipRemote work
$95k - $125k per year
...can work from anywhere in Canada. The Senior Manager, Data Engineering leads a team of data engineers, data analysts and business system... ...including operability, availability, performance, reliability and security Implement rigorous data validation framework in...Full timeHome officeFlexible hoursusd5 000 per month
...Palo Alto · Full-time · On-site About the Company We're an early-stage, YC-backed hardware company building robotics hardware for demanding field environments. Our systems are already in the hands of users, and their feedback drives what we build next. We went from concept...Full timeInternshipRelocationOverseas$240k - $280k per year
...donor fundraising. Position Summary: The Director, Solutions Engineering is responsible for leading Kindsight’s high-impact, customer-... ...product demonstrations ranging from virtual overviews to multi-day on-site enterprise presentations. Establish demo standards, playbooks...$180k - $220k per year
...Principal Software Engineer Location: Remote (Anywhere in Canada) Company Overview... ...for engineering quality, performance, and reliability Identify and prevent overengineering,... ...applications ~ Proven experience operating at a staff or principal level, influencing multiple...Long term contractFull timeRemote work- ...manage a small agile team that turns operational friction into reliable, auditable tooling. This role sits at the intersection of translation... ...and manage a small agile team focused on localization engineering, automation, linguistic data, analytics, and workflow tooling,...Full timeInternshipWork at officeRemote work
- ...well. This role exists at the intersection of product and engineering, and it leans hard into both. You’ll own the full lifecycle of... ...inconsistent documentation, rate limits, and auth schemes, and made it reliable. Webhooks are the primary integration pattern for Twilio,...Full time
- ...nearly 30 years, our expert teams have helped our clients deliver reliable products that users enjoy interacting with. Our 100% Canada‑... ...of application Summary We are seeking an experienced AI Engineer with 5+ years of experience designing, developing, and deploying...Contract workInternship
$150k - $185k per year
...Summary We’re looking for a Senior Data Engineer to join our Product Engineering team and... ...and platform capabilities that move data reliably from complex source systems into trusted,... ...You’ll Do Design, build, and maintain reliable ELT and ETL pipelines using Snowflake,...Long term contract$68.81k - $114.71k per year
...Europe, Australasia and Africa. Working for us, you’ll help ensure that Canada’s critical services and assets are readily available, reliable, and capable for our defence and civil customers and contribute to Babcock’s purpose: “Creating a safe and secure world, together.” Our...ApprenticeshipInternshipMonday to fridayShift workDay shift- ...secure. Job Description Fires Systems Engineer GDIT has an immediate career... ...performance. Ensure high availability, reliability, and resilience of system and network infrastructure... ...of in-service and future systems, and on-site and remote support of field units...Contract workImmediate startRemote workWorldwide
- ...AUTHORISED TO WORK IN CANADA. WE ARE UNABLE SPONSOR VISAS AT THIS TIME. What do we need Dotmatics is seeking a Senior Full Stack Engineer with an understanding of both Node.js and React to join our team. You will be responsible for developing and implementing software...Full timeRemote workVisa sponsorshipFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!
- project engineer assistant project manager Canada
- assistant ingénieur Canada
- assistant electrical engineer Canada
- stage assistant ingénieur Canada
- senior site reliability engineer Canada
- site reliability engineer Canada
- site reliability engineer intern Canada
- site safety Canada
- site maintenance Canada
- site carpenter Canada


