Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Staff Site Reliability Engineer

Canada
  • Remote job

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Staff Site Reliability Engineer based in Canada.

This is a high-impact reliability leadership role within a fully remote engineering organization operating globally.
You will be the first dedicated SRE, helping establish reliability practices across multiple engineering teams and critical production systems.
The role combines hands-on engineering with organization-wide influence, covering observability, incident response, operational readiness, and resilience.
You will work closely with engineering leadership, infrastructure specialists, architects, and product teams to make reliability measurable and actionable.
A major focus will be embedding SRE principles into engineering culture rather than simply owning individual services.
You will also help shape how AI is used for incident investigation, operational tooling, observability, and safe system operations.
The position offers substantial autonomy to define standards, coach engineers, and build practices that scale with the organization.

Accountabilities
  • Define and implement SLIs and SLOs for critical production request paths, ensuring reliability objectives are visible, measurable, reviewed, and connected to engineering decisions.

  • Introduce and champion error budgets as a practical framework for balancing reliability investments with product and feature delivery.

  • Establish and maintain the reliability metrics used by engineering leadership to evaluate progress and identify areas requiring investment.

  • Strengthen the complete incident management lifecycle, including detection, response, communication, escalation, postmortems, and follow-up actions.

  • Improve alert quality, anomaly detection, escalation processes, and shared operational tooling in collaboration with infrastructure teams.

  • Lead reliability assessments for high-risk changes and new services, covering production readiness, capacity, failure modes, rollback strategies, and operational risks.

  • Introduce deliberate failure testing, game days, and chaos exercises to identify weaknesses and validate safe operational limits before incidents occur.

  • Work directly with engineering teams on complex reliability challenges through focused engagements, leaving behind stronger practices and clear ownership.

  • Coach Staff and Lead engineers to become reliability advocates within their respective teams and help establish distributed SRE ownership.

  • Develop lightweight, repeatable operational standards covering production readiness, on-call practices, runbooks, change safety, and service operability.

  • Partner with architects and technical leads to ensure reliability and failure tolerance are incorporated into system design rather than addressed after deployment.

  • Remain hands-on during production incidents and investigations, building tooling, dashboards, automation, and reference implementations where appropriate.

  • Promote effective use of AI for incident investigation, telemetry analysis, postmortem development, runbook creation, observability, and reliability tooling.

  • Help structure operational data, alerts, dashboards, and runbooks so that both engineers and AI agents can safely interpret and act on production signals.

  • Contribute production fixes and improvements directly through code and infrastructure changes rather than limiting the role to recommendations and reviews.

  • Requirements

    • 10+ years of engineering experience, including at least 3 years in SRE, production engineering, or a reliability-focused Staff Engineer role operating across multiple teams.

    • Demonstrated experience owning reliability at a platform or organizational level rather than only for an individual service.

    • Deep practical experience designing and implementing SLIs, SLOs, and error budgets, including successfully driving adoption across product and engineering teams.

    • Strong incident leadership experience, including managing high-severity, customer-facing incidents and leading effective postmortems that result in measurable improvements.

    • Advanced understanding of distributed-system failure modes, including database and cache saturation, cascading failures, retry storms, capacity constraints, graceful degradation, and load shedding.

    • Strong hands-on experience with Kubernetes, AWS, and modern observability platforms such as Datadog or comparable technologies.

    • Ability to read and write production code in Go, TypeScript, or a similar language, as well as work with infrastructure as code.

    • Demonstrated ability to influence teams without direct authority and successfully change engineering practices across an organization.

    • Strong coaching and mentoring skills, with evidence of developing engineers into effective reliability owners.

    • Exceptional written and verbal communication skills, with the ability to clearly communicate incidents, risks, technical trade-offs, and reliability priorities to both engineers and executives.

    • Strong preference for asynchronous, documented decision-making and clear technical communication.

    • Practical experience using AI tools for incident investigation, telemetry analysis, runbook and postmortem development, and engineering tooling.

    • Understanding of how operational data, alerts, dashboards, and runbooks should be structured to support safe AI-assisted diagnosis and operations.

    • Pragmatic approach to reliability, with the ability to balance operational risk, engineering investment, delivery speed, and business priorities.

    • Experience in fraud detection, identity, payments, or other real-time and adversarial environments is an asset.

    • Experience with multi-region architectures, cell-based architectures, or failure-isolation strategies is a plus.

    • Experience operating Elasticsearch, Redis, DynamoDB, or Kafka at scale and understanding their failure modes is beneficial.

    • Familiarity with FinOps and cloud infrastructure cost-versus-reliability trade-offs is an advantage.

    • Must be authorized to work from the hiring location; visa sponsorship is not provided.

    • Benefits

      • Fully remote working environment.

      • Opportunity to become the first dedicated Site Reliability Engineer and establish organization-wide reliability practices.

      • High level of autonomy and direct influence over engineering standards, operational practices, and platform reliability.

      • Opportunity to work across multiple engineering teams and critical production systems.

      • Close collaboration with engineering leadership, architects, infrastructure teams, and technical leads.

      • Opportunity to shape AI-assisted reliability practices and the future of production operations.

      • Strong focus on professional growth, technical leadership, coaching, and knowledge sharing.

      • Inclusive, globally distributed engineering environment that values diverse perspectives and backgrounds.

      • For US-based employees, the stated cash compensation range is $177,000–$240,000 USD , with actual offers varying according to factors such as experience, skills, education, certifications, and market conditions. Compensation may differ for other hiring locations.

      • Remote work eligibility is subject to applicable regulatory and security requirements in the candidate's location.

Vacancy posted 4 hours ago
Similar jobs that could be interesting for youBased on the Staff Site Reliability Engineer in Canada vacancy
  • $160k - $180k per year

     ...donor fundraising. Position Summary: We are seeking a Staff Software Engineer, Payments to provide technical leadership as Kindsight...  ...payments partner, how the platform should scale, and how we build reliable systems for moving and reconciling real donor funds. This... 
    Suggested
    Internship

    Kindsight

    Canada
    19 days ago
  • $80k - $95k per year

     ...Running. Improve What's Next. We're partnering confidentially with a leading Canadian food manufacturer seeking a Mechanical Reliability Engineer to play a key role in improving equipment reliability, maintenance strategy, and plant performance within a highly automated manufacturing... 
    Suggested
    Long term contract
    Permanent employment
    Full time
    For contractors
    Relocation package

    Remote People

    Canada
    a month ago
  • $150k - $210k per year

     ...moments to travelers around the world. We’re looking for a Staff Software Engineer to join the Checkfront team. Checkfront is an established...  ...maintaining and improving the product with work spanning reliability, performance, retention, maintainability, and new feature development... 
    Suggested
    Full time

    Checkfront

    Canada
    more than 2 months ago
  • $148k - $185k per year

     ...Coursera achieve a seamless experience across its platform. As a Staff Product Designer, you will lead the design strategy and...  ...for a major product area. You'll partner closely with Product, Engineering, Research, and Data Science across multiple teams to shape scalable... 
    Suggested
    Full time
    Internship
    Work at office

    Coursera

    Canada
    7 days ago
  • $45 - $50 per hour

     ...future of energy with us. POSITION SUMMARY The role of Site Quality Lead directly supports the company’s mission to develop...  ...diploma or equivalent (GED) required. ~ Degree in technical or engineering field, or equivalent work experience ~3+ years’ experience preferred... 
    Suggested
    Hourly pay
    Long term contract
    For contractors
    For subcontractor
    Weekend work

    Nordex Group

    Canada
    6 days ago
  • $130k - $150k per year

     ...future of energy with us. POSITION SUMMARY The role of Site Manager for Project Management directly supports the company’s mission...  ...and tracks all the quality control, transportation, and engineering documents for the project. Prepares daily and weekly job progress... 
    Long term contract
    For contractors
    For subcontractor
    Work at office

    Nordex Group

    Canada
    4 days ago
  •  ...RESPONSIBILITIES At SpryPoint we value collaborate work environments, automation, learning, and delivering value to our users. As a Software Engineer III  at SpryPoint, you will be building and integrating interactive web applications, services, and apps that real people will... 
    Summer work
    Internship
    Local area
    Remote work
    Home office
    Flexible hours

    SpryPoint

    Canada
    15 hours ago
  • $115k - $120k per year

     ...energy with us. POSITION SUMMARY The role of Assistant Site Manager directly supports the company’s mission to develop and deliver...  ...phase. WHAT YOU HAVE ~ Associate’s in mechanical engineering or related discipline, or equivalent related work experience accepted... 
    Long term contract
    For contractors
    For subcontractor
    Work at office

    Nordex Group

    Canada
    4 days ago
  •  ...of our offices or hubs, or a co-working space near you. Data Engineering is unique at Coursera. Our team doesn't simply build reports on...  ...Architect scalable data models and build efficient and reliable ETL pipelines to bring the data into our core data lake Design... 
    Full time
    Remote work
    Work from home
    Worldwide
    Flexible hours

    Coursera

    Canada
    2 days ago
  • $120k - $150k per year

     ...are seeking a Senior Software Engineer, Payments to join a newly...  ...You will work closely with our Staff Engineer and Director of...  ...help establish the patterns for reliability, reconciliation, integration,...  ...Payments partner. Design reliable transaction-processing flows... 

    Kindsight

    Canada
    19 days ago
  • $100k - $120k per year

     ...Position Summary: We are seeking an Intermediate Software Engineer, Payments to join a newly forming Payments Engineering...  ..., AdvRM, and Connect. You'll work alongside Senior and Staff engineers to build reliable backend services, integrate with our selected payments... 

    Kindsight

    Canada
    19 days ago
  • $148.3k - $171k per year

     ...leads the industry with its low fees, advanced trading engine, strong security and reliability, and robust compliance framework. Learn more at . All roles...  ...for designing, developing, testing, and operating reliable backend services that support Binance.US products and... 
    Full time
    Local area
    Remote work
    Work from home

    binance.us

    Canada
    12 hours ago
  •  ...Description Job Opportunity: Integration Senior Software Engineer Role Overview Job Title: Integration...  ...– September 2027) Work Location: Remote (No on-site requirement) Security Clearance: Reliability Status (required prior to contract start) Language... 
    Contract work
    Remote work

    Oomple

    Canada
    23 days ago
  • $110k - $120k per year

     ...intensive industries, where performance, reliability, and precision are essential. With a strong...  ...the world. Their success is built on engineering excellence, practical innovation, and a...  ...Abbotsford, BC. The first three months are on-site to support onboarding, product training,... 
    Permanent employment
    Full time
    Relocation package

    Stoakley-Stewart Consultants

    Canada
    more than 2 months ago
  • $110k - $150k per year

     ...applications and next steps. Our partner is looking for a Structural Engineer - Project Manager based in Canada. This is a senior...  ...consistent project delivery. Mentor junior engineers and technical staff while promoting technical quality, efficiency, knowledge sharing,... 
    Remote job
    Long term contract
    For contractors
    Canada
    4 hours ago
  • $170.4k - $234.3k per year

     ...and next steps. Our partner is looking for a Senior Full Stack Engineer based in Canada. This role offers the opportunity to build and...  ...developer ecosystems. ~ Work collaboratively across teams to develop reliable systems and improve how software is designed, tested, deployed,... 
    Remote job
    Long term contract
    Work from home
    Flexible hours
    Canada
    4 hours ago
  • $5000 per month

     ...applications and next steps. Our partner is looking for a Lead Software Engineer (.NET) based in Canada. This role offers the opportunity to...  ..., resolve complex backend and data challenges, and help ensure reliable delivery. You will collaborate closely with Product and... 
    Remote job
    Full time
    Home office
    Canada
    4 hours ago
  • $114k - $172k per year

     ...workflow, and built in reporting. It runs in the cloud on a secure, reliable, extensively audited platform and integrates deeply with on...  ...systems. We are looking for an experienced Senior UI Software Engineer to work on our Onboarding and Lifecycle Management (LCM) Platform... 
    Full time
    Remote work
    Flexible hours

    Okta

    Canada
    15 hours ago
  •  ...nearly 30 years, our expert teams have helped our clients deliver reliable products that users enjoy interacting with. Our 100% Canada‑...  ...Summary We are seeking an experienced Infrastructure Engineer to join our Infrastructure & DevOps team. This role is responsible... 
    Contract work
    Internship
    Remote work

    PLATO

    Canada
    4 days ago
  • $95k - $125k per year

     ...can work from anywhere in Canada. The Senior Manager, Data Engineering leads a team of data engineers, data analysts and business system...  ...including operability, availability, performance, reliability and security Implement rigorous data validation framework in... 
    Full time
    Home office
    Flexible hours

    Heart & Stroke

    Canada
    5 days ago
  • usd5 000 per month

     ...Palo Alto · Full-time · On-site About the Company We're an early-stage, YC-backed hardware company building robotics hardware for demanding field environments. Our systems are already in the hands of users, and their feedback drives what we build next. We went from concept... 
    Full time
    Internship
    Relocation
    Overseas

    Simonyan Consulting

    Canada
    1 day ago
  • $240k - $280k per year

     ...donor fundraising. Position Summary: The Director, Solutions Engineering is responsible for leading Kindsight’s high-impact, customer-...  ...product demonstrations ranging from virtual overviews to multi-day on-site enterprise presentations. Establish demo standards, playbooks... 

    Kindsight

    Canada
    21 days ago
  • $180k - $220k per year

     ...Principal Software Engineer Location: Remote (Anywhere in Canada) Company Overview...  ...for engineering quality, performance, and reliability Identify and prevent overengineering,...  ...applications ~ Proven experience operating at a staff or principal level, influencing multiple... 
    Long term contract
    Full time
    Remote work

    eDynamic Learning

    Canada
    more than 2 months ago
  •  ...manage a small agile team that turns operational friction into reliable, auditable tooling. This role sits at the intersection of translation...  ...and manage a small agile team focused on localization engineering, automation, linguistic data, analytics, and workflow tooling,... 
    Full time
    Internship
    Work at office
    Remote work

    Apertera

    Canada
    13 days ago
  •  ...well.   This role exists at the intersection of product and engineering, and it leans hard into both. You’ll own the full lifecycle of...  ...inconsistent documentation, rate limits, and auth schemes, and made it reliable. Webhooks are   the primary integration pattern for Twilio,... 
    Full time

    Symend

    Canada
    13 days ago
  •  ...nearly 30 years, our expert teams have helped our clients deliver reliable products that users enjoy interacting with. Our 100% Canada‑...  ...of application Summary We are seeking an experienced AI Engineer with 5+ years of experience designing, developing, and deploying... 
    Contract work
    Internship

    PLATO

    Canada
    a month ago
  • $150k - $185k per year

     ...Summary We’re looking for a Senior Data Engineer to join our Product Engineering team and...  ...and platform capabilities that move data reliably from complex source systems into trusted,...  ...You’ll Do Design, build, and maintain reliable ELT and ETL pipelines using Snowflake,... 
    Long term contract

    Kindsight

    Canada
    29 days ago
  • $68.81k - $114.71k per year

     ...Europe, Australasia and Africa. Working for us, you’ll help ensure that Canada’s critical services and assets are readily available, reliable, and capable for our defence and civil customers and contribute to Babcock’s purpose: “Creating a safe and secure world, together.” Our... 
    Apprenticeship
    Internship
    Monday to friday
    Shift work
    Day shift

    Babcock International

    Canada
    a month ago
  •  ...secure. Job Description Fires Systems Engineer GDIT has an immediate career...  ...performance. Ensure high availability, reliability, and resilience of system and network infrastructure...  ...of in-service and future systems, and on-site and remote support of field units... 
    Contract work
    Immediate start
    Remote work
    Worldwide

    General Dynamics Information Technology

    Canada
    more than 2 months ago
  •  ...AUTHORISED TO WORK IN CANADA. WE ARE UNABLE SPONSOR VISAS AT THIS TIME. What do we need Dotmatics is seeking a Senior Full Stack Engineer with an understanding of both Node.js and React to join our team. You will be responsible for developing and implementing software... 
    Full time
    Remote work
    Visa sponsorship
    Flexible hours

    Dotmatics

    Canada
    2 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Staff Site Reliability Engineer. Be the first to apply!