Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Full-time

TOTEM Recruteur de talent

Employment Status: Permanent
Schedule: 40 hours/week – 100% remote work

Job Description

We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability, performance, and resilience of high-availability SaaS platforms.

Working in an AWS and Kubernetes environment, you will help design, automate, monitor, and continuously improve the infrastructure supporting critical cloud-based services. This role combines hands-on engineering with operational leadership, giving you direct ownership of system availability, scalability, and incident response.

You will be involved throughout the service lifecycle, from architecture and launch preparation to production monitoring and continuous improvement.


Responsibilities

  • Partner with engineering teams during system design, capacity planning, launch readiness, and production deployment.
  • Monitor and improve service availability, latency, performance, and overall system health.
  • Identify recurring operational issues and implement sustainable solutions that improve scalability and resilience.
  • Define and evolve observability practices, including dashboards, alerts, SLOs, and SLIs.
  • Build and maintain automated infrastructure using Terraform and CI/CD pipelines.
  • Develop automation and operational tooling to reduce manual intervention and support self-healing systems.
  • Coordinate incident response and act as Incident Commander during critical production events.
  • Facilitate blameless post-incident reviews and ensure that corrective actions are completed.
  • Use AI-assisted engineering tools responsibly to accelerate development and operational workflows.
  • Maintain clear technical documentation and contribute to the continuous improvement of SRE practices.


Required profile

  • Significant experience in Site Reliability Engineering, Cloud Engineering, DevOps, or a similar infrastructure-focused role.
  • ​​​​​​​Experience supporting complex or large-scale SaaS environments with high availability requirements.
  • Strong hands-on knowledge of AWS services and architecture, including multi-account environments, VPC, EC2, and EKS.
  • Proven experience operating and troubleshooting Kubernetes environments at scale.
  • Strong knowledge of Infrastructure as Code, particularly Terraform.
  • Experience building or maintaining CI/CD pipelines using GitLab, Jenkins, or comparable tools.
  • Experience with enterprise observability platforms such as Datadog, Prometheus, Grafana, or equivalent solutions.
  • Strong scripting skills using Python, Bash, or a similar language.
  • Direct experience participating in on-call rotations, coordinating incident response, and conducting post-incident reviews.
  • Familiarity with Java or .NET application environments is considered an asset.
  • Ability to communicate clearly and collaborate with development, infrastructure, security, and operations teams.
  • Must be legally authorized to work in Canada.


What to Expect

  • Remote-first work environment within Canada.
  • ​​​​​​​Occasional visits to a local office or participation in in-person meetings may be required, representing less than 10% of the role.
  • Participation in a scheduled on-call rotation is required.


Does this opportunity sound like a good fit for you? Apply now through our website or by sending your resume to View email address on emploisti.com.

Thank you for your interest in this position; only candidates who meet our client’s requirements will be contacted.

 

The masculine gender is used as a neutral form.

 

#totemtech

Vacancy posted 4 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Toronto, ON vacancy
  • $110k - $120k per year

     ...professional development support, discounts through Perkopolis, and recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring the availability, performance, and resilience of the... 
    Suggested
    Full time
    Temporary work
    Internship
    Work at office
    Remote work

    Momentum Financial Services Group

    Toronto, ON
    22 hours ago
  • $100k - $125k per year

     ...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,...  ...culture Tech at Tipalti  Our tech teams are the engine behind our business. Tipalti’s tech ecosystem is extremely... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Tipalti

    Toronto, ON
    22 hours ago
  •  ...Docebians around the world and help us reinvent the way people learn, because learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident response while also shaping the underlying infrastructure that... 
    Suggested
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    22 hours ago
  •  ...learning never stops. The Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team of engineers dedicated...  ...Product, Engineering, and Support leadership to integrate reliable SRE practices into early planning and product delivery... 
    Suggested
    Long term contract
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    22 hours ago
  • $140k - $155k per year

     ...profiles! This is a hands-on senior engineering role focused on improving production...  ...enable engineering teams to ship secure, reliable, and scalable software with confidence. You...  ...engineers on cloud-native technologies, site reliability engineering principles, and operational... 
    Suggested
    Remote job
    Permanent employment
    Full time
    Flexible hours

    Caseware

    Toronto, ON
    22 hours ago
  • $140k - $182k per year

     ...client base with operations throughout North America, Central America, Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and software development. You will support and help inform the evolution of... 
    Full time

    Movable Ink

    Toronto, ON
    22 hours ago
  • $154k - $200k per year

     ...client base with operations throughout North America, Central America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic technical leadership across infrastructure and software development... 
    Long term contract
    Full time

    Movable Ink

    Toronto, ON
    22 hours ago
  • $153.82k - $277k per year

     ...place where you can thrive, we can’t wait to meet you. Site Reliability Engineers (SREs) at Braze are responsible for keeping all internal-facing...  ...level feature requirements and helping translate them into reliable, highly scalable technology stacks. You will be responsible... 
    Permanent employment
    Full time
    Internship
    Work at office
    Local area
    Remote work
    Flexible hours
    Rotating shift

    Braze

    Toronto, ON
    22 hours ago
  •  ...of both work styles in a workplace that is intentional about belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an engineer, a builder, and a systems thinker. You’ll create middleware and platform guardrails... 
    Full time
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to friday
    Flexible hours

    Imanage

    Toronto, ON
    22 hours ago
  •  ...competitive advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in software...  ...experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly evaluate products... 
    Full time

    Serigor Inc

    Toronto, ON
    22 hours ago
  • $140k - $180k per year

     ...the shared infrastructure that all other engineering teams at Tripstack depend on. This spans...  ...boundary sits, and the deployment patterns and reliability standards that apply across both. This...  ...engineering depth - interconnects, BGP, site-to-site VPN, cross-region peering GDPR... 
    Remplacement
    Full time
    Work at office
    Immediate start
    Flexible hours

    Tripstack

    Toronto, ON
    22 hours ago
  •  ...AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work...  ...storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking... 
    Full time
    Remote work

    bosonai

    Toronto, ON
    23 days ago
  • $80 - $110 per hour

     ...Networks is headquartered in Irvine, CA USA with Asia HQ in Singapore and also operating in Denmark, Spain and Vietnam. The Site Reliability Engineer  will improve the availability, performance, scalability and recoverability of AXON Networks cloud solutions. You will... 
    Remote job
    Full time
    Contract work

    Axon-networks

    Toronto, ON
    22 hours ago
  • $92k - $118k per year

     ...Build reliable, resilient cloud platforms that keep critical financial services running. The Role We are looking for an experienced Site Reliability and DevOps Engineer to join the SRE team within our growing Corporate Action Processing group. You will bring 3+ years... 
    Internship
    Immediate start

    Capco

    Toronto, ON
    4 days ago
  •  ...customers. Cohere is a team of researchers, engineers, designers, and more, who are passionate...  ...building high-performance, scalable and reliable machine learning systems? Do you want to...  ...NLP applications? We are looking for a Site Reliability Engineer to join the Model... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    Toronto, ON
    22 hours ago
  •  ...visibility, and optimize spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s...  ...Excellence & Automation Design, develop, and automate reliable cloud infrastructure and platform services. Apply Infrastructure... 
    Full time
    Manual labor
    Local area
    Flexible hours

    Emburse

    Toronto, ON
    22 hours ago
  • $144k - $200k per year

    **The Team** Platform Engineering is the department within SRE that is responsible for a range...  ...role in developing and maintaining the reliable and globally connected multi-cloud network...  ...Overview** We are seeking a talented Site Reliability Engineer (SRE) with a strong... 
    Full time
    Work at office
    Remote work
    Worldwide
    Flexible hours

    Mongodb

    Toronto, ON
    22 hours ago
  • $130k - $180k per year

     ...Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure...  ...implement, and maintain CI/CD pipelines, enabling rapid and reliable software releases. Automate and optimize our infrastructure... 
    Full time
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Acquird.io

    Toronto, ON
    22 hours ago
  • $110k - $125k per year

     ...systems bringing #BetterGlobalHealth to patients everyday! Apply today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across... 
    Full time
    Remote work
    Flexible hours

    Smile Digital Health

    Toronto, ON
    22 hours ago
  • $116k - $235.1k per year

     ...About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission... 
    Long term contract
    Remplacement
    Full time
    Temporary work
    Work at office
    Local area
    Immediate start
    Flexible hours
    2 days per week

    Tubi - Canada

    Toronto, ON
    4 days ago
  •  ...that are particularly strong in a few areas, and have some interest and capabilities in others. About the Role: As a Site Reliability Engineer , you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS... 
    Full time

    Kong Company

    Toronto, ON
    22 hours ago
  •  ...To learn more about CIBC please visit What youll be doing As a member of CIBCs Application Reliability Engineering Platform team the Consultant Site Reliability Engineering will play a key role in CIBC Digital and Client Experience Technology to support the... 
    Full time
    Contract work
    3 days per week
    1 day per week

    Canadian Imperial Bank of Commerce

    Toronto, ON
    a month ago
  •  ...a skill on their LinkedIn profiles! This is a hands-on senior engineering role focused on improving production resilience, strengthening...  ...operational practices that enable engineering teams to ship secure, reliable, and scalable software with confidence. You will help... 
    Permanent employment
    Full time
    Remote work

    Caseware

    Toronto, ON
    a month ago
  • $140k - $165k per year

     ...transformation, transitioning from traditional service desk operations to an AI-driven, self-healing operational fabric. As Lead Site Reliability Engineer (SRE), you will lead enterprise observability, AIOps automation, and SRE governance across a global footprint. You will... 

    Randstad

    Toronto, ON
    25 days ago
  • $78.62 - $89.34 per hour

    Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering and Operations team. In this engineering-focused role, you will move beyond traditional database administration to act... 
    Long term contract
    Permanent employment
    Full time
    Contract work
    Work at office

    Randstad

    Toronto, ON
    more than 2 months ago
  • $120k - $170k per year

     ...We are seeking a highly skilled and motivated Senior DevOps Engineer to join our dynamic team and play a key role in designing, implementing...  ..., investigate incidents, and drive improvements that increase reliability and operational efficiency; Write and maintain automation,... 
    Full time
    Work at office
    Local area
    Flexible hours

    Magnet Forensics

    Toronto, ON
    22 hours ago
  •  ...Description Application Support (Windows, UNIX/Linux, OpenShift, and PostgreSQL) This role is part of the Production Support and Reliability Engineering team, responsible for ensuring the stability, availability, and performance of enterprise applications running across Windows... 
    Permanent employment
    Full time
    Local area

    Capgemini

    Toronto, ON
    a month ago
  • $82.2 - $89.34 per hour

    We are seeking a highly skilled Site Reliability Specialist IV for an enterprise-level contract opportunity based in Toronto. In this role, you will take on a premier cloud engineering, platform automation, and operational reliability capacity, specializing in designing, building... 
    Long term contract
    Permanent employment
    Contract work

    Randstad

    Toronto, ON
    21 days ago
  • $141k - $191k per year

     ...and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence. About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for: Team Leadership... 
    Work at office
    Local area
    Flexible hours
    2 days per week
    3 days per week

    Thomson Reuters

    Toronto, ON
    more than 2 months ago
  •  ...contributors of all seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next...  ...takes to run hands-on testing at manufacturing and test partner sites, including in Taiwan.   Tenstorrent offers a highly competitive... 
    Permanent employment
    Internship
    Second job

    Tenstorrent

    Toronto, ON
    4 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!