Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Full-time

Tecsys Inc.

Toronto, ON
  • Remote job

Having recognized the advantages of remote work, including employee morale, productivity, reduced commuting on employee wellbeing and the environment, we are proud to be a digital-first company. The technologies and programs in which we invested have provided a fantastic foundation to this end. Our digital-first work environment, together with our conveniently located offices and collaborative workspaces, provide our team with the freedom and flexibility to work in the way that makes our employees most productive.

About us

Tecsys is a fast-growing innovator offering supply chain solutions to industry leading healthcare systems, hospitals, and pharmacy businesses to distributors, retailers, and 3PLs. We work with industry leaders to transform their supply chains through technology. If you thrive on tackling interesting challenges with continuous learning opportunities, then Tescys could be a good fit for you!

About the Role

We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a team at the heart of platform reliability for mission-critical SaaS environments. You will help maintain, optimize, and ensure the reliability and performance of the systems that power our cloud infrastructure across AWS and Kubernetes, with a strong focus on automation, observability, and continuous improvement. This role blends reliability engineering with incident command, giving you real ownership over uptime, performance, and innovation. You will be part of a highly skilled team that values creative problem-solving, operational excellence, and continuous improvement through automation and resilience engineering.

Your responsibilities

  • Collaborate with other Engineering teams to support services before they go live through activities such as system design consulting, developing software platforms and frameworks, capacity planning and launch reviews.
  • Innovate relentlessly: Identify pain points, propose creative solutions, and drive initiatives that simplify, scale, and strengthen the platform.
  • Maintain services once they are live by measuring and monitoring availability, latency and overall system health.
  • Own observability: Enhance and expand monitoring and alerting using Datadog; define SLOs/SLIs and create actionable dashboards that drive reliability outcomes.
  • Drive automation: Develop and improve internal tooling, IaC frameworks, and pipelines (Terraform, GitLab CI/CD) to reduce manual intervention and enable self-healing systems.
  • Scale systems sustainably through automation and evolve systems by pushing for changes that improve reliability and velocity.
  • Be on-call.
  • Practice sustainable incident response and blameless postmortems. Lead post-incident reviews (RCAs) and identify long-term fixes that improve stability, reliability, and developer experience.
  • Implement monitoring, Logging, alerting, and SLA Reporting.
  • Create and maintain technical documentation.
  • Implement, maintain and mature SRE best practices.
  • Lead incidents: Act as Incident Commander for Incidents; coordinate cross-team response, manage communications, and ensure rapid service restoration.
  • Provide support for our planning and deployment teams to enable stability, predictability, and scale in our continued growth.
  • Collaborate with members of the Platform Engineering team to implement and support far-reaching strategic efforts, provide constructive feedback, and foster a collaborative environment.
  • Work cross-functionally with internal teams and vendors to manage our growth around the globe, with a strong focus on maintaining the high level of performance, availability, and reliability for our users.
  • 5+ years in Site Reliability, Cloud, or DevOps Engineering, ideally in SaaS or large-scale production environments.
  • Experience designing and deploying large scale systems, multi-vendor platforms and globally distributed infrastructure.
  • Proven experience managing cloud infrastructure in AWS (multi-account, VPC, EC2, EKS) and Kubernetes at scale.
  • Strong hands-on experience with IaC and automation (Terraform, Ansible, or similar).
  • Familiarity with CI/CD pipelines and release automation (GitLab preferred, Jenkins acceptable).
  • Deep understanding of monitoring and observability using Datadog (or equivalent), including metric design, log pipelines, alerting, and dashboards.
  • Experience with incident management, on-call participation, escalation, and structured postmortems.
  • Scripting skills in Python, Bash, Java or equivalent for automation and diagnostics.
  • Curiosity, ownership, and a bias for action; you see a problem, you solve it, and you share the lessons learned.
  • Experience with Fedramp (The Federal Risk and Authorization Management Program) compliance is a strong asset.
  • Basic knowledge of Java- or .Net-based development required.
  • Strong English communication skills, both written and spoken, are essential for effective correspondence with customers, business partners and colleagues beyond the province of Quebec.

Additional requirements:

  • Escalation on-call rotation
  • Occasional travel (quarterly offsites, conferences – less than 10%)

At Tecsys, we are committed to fostering a diverse and inclusive workplace where all employees feel valued, respected, and empowered. We believe that diversity drives innovation and strengthens our ability to deliver exceptional solutions. We welcome and encourage applicants from all backgrounds, experiences, and perspectives to join our team.

Tecsys is an equal opportunity employer. Accommodation is available for applicants selected for an interview.

NB: if you are applying to this position, you must be a Canadian Citizen or a Permanent Resident of Canada, OR , have a valid Canadian work permit.

Vacancy posted 6 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Toronto, ON vacancy
  •  ...the world and our goal is preserving uncensored Internet access and online privacy for all. Right now we are looking for a Site Reliability Engineer to help us tame DNS. This position can be hybrid at our Toronto offices or remote in Canada only but you MUST reside in the... 
    Suggested
    Full time
    Direct hire
    Work at office
    Remote work

    Funded.club

    Toronto, ON
    6 hours ago
  • $115k - $130k per year

     ..., and continuously raising the bar.  Kaseya is hiring a Site Reliability Engineer to keep our production systems healthy as we scale. You'll own...  ...enforce SLOs, SLIs, and error budgets that keep our systems reliable Lead incident response, troubleshooting, and blameless... 
    Suggested
    Long term contract
    Full time
    Worldwide

    Kaseya Careers

    Toronto, ON
    6 hours ago
  • $110k - $120k per year

     ...professional development support, discounts through Perkopolis, and recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring the availability, performance, and resilience of the... 
    Suggested
    Full time
    Temporary work
    Internship
    Work at office
    Remote work

    Momentum Financial Services Group

    Toronto, ON
    6 hours ago
  • $100k - $125k per year

     ...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,...  ...culture Tech at Tipalti  Our tech teams are the engine behind our business. Tipalti’s tech ecosystem is extremely... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Tipalti

    Toronto, ON
    6 hours ago
  • $100k per year

     ...across internal clusters and customer deployments. This role sits at the intersection of site reliability, infrastructure operations, and customer engineering, ensuring our systems are reliable, observable, and production-ready. This role is hybrid, based out of Toronto, ON;... 
    Suggested
    Permanent employment
    Full time

    Tenstorrent

    Toronto, ON
    6 hours ago
  •  ...competitive advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in software...  ...experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly evaluate products... 
    Full time

    Serigor Inc

    Toronto, ON
    6 hours ago
  •  ...of both work styles in a workplace that is intentional about belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an engineer, a builder, and a systems thinker. You’ll create middleware and platform guardrails... 
    Full time
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to friday
    Flexible hours

    Imanage

    Toronto, ON
    6 hours ago
  • $140k - $182k per year

     ...client base with operations throughout North America, Central America, Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and software development. You will support and help inform the evolution of... 
    Full time

    Movable Ink

    Toronto, ON
    6 hours ago
  •  ...Docebians around the world and help us reinvent the way people learn, because learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident response while also shaping the underlying infrastructure that... 
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    6 hours ago
  •  ...learning never stops. The Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team of engineers dedicated...  ...Product, Engineering, and Support leadership to integrate reliable SRE practices into early planning and product delivery... 
    Long term contract
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    6 hours ago
  • $154k - $200k per year

     ...client base with operations throughout North America, Central America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic technical leadership across infrastructure and software development... 
    Long term contract
    Full time

    Movable Ink

    Toronto, ON
    6 hours ago
  •  ...About the Role Fivetran is looking for a high-performance engineer to be a part of a team of Site Reliability Engineers. You will be working closely with engineering teams, product managers, as well as support and sales engineers to build the future of the Fivetran Data... 
    Full time
    Work at office
    Remote work

    Fivetran

    Toronto, ON
    6 hours ago
  •  ...and how we use AI in our recruiting process here . The Site Reliability Engineering organization at Pinterest is accountable for ensuring...  ...Demonstrated ability to write effective prompts to get high-quality, reliable outputs from LLMs ~ Demonstrated ability to use AI to... 
    Full time
    Work at office

    Pinterest

    Toronto, ON
    6 hours ago
  •  ...for keeping a large scale, multi-tenant SaaS platform running reliably for customers around the world. You'll take a hands-on lead role...  ...just symptoms, and to raise the reliability bar across the wider engineering org. What You'll Be Doing Take point on major incidents... 
    Full time
    Immediate start
    Flexible hours

    HighlightTA

    Toronto, ON
    6 hours ago
  •  ...Senior Site Reliability Engineer - Edge Location : Ottawa/Toronto, On-Site Reports to: Head of Security The Role You own the edge compute module — the standard hardware stack, OS image, and runtime that integrates with Dominion Dynamics's mesh radios, sensors... 
    Full time

    dominion%20dynamics

    Toronto, ON
    5 hours ago
  • $110k - $125k per year

     ...systems bringing #BetterGlobalHealth to patients everyday! Apply today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across... 
    Full time
    Remote work
    Flexible hours

    Smile Digital Health

    Toronto, ON
    6 hours ago
  • $243k - $297k per year

     ...more resilient businesses. As Relay continues to scale, the reliability, performance, and resilience of our platform are no longer...  ...role responsible not only for guiding a strong team of Site Reliability Engineers, but for shaping how reliability strategy influences engineering... 
    Long term contract
    Full time
    Internship
    Work at office
    Trial period
    Flexible hours

    Relay

    Toronto, ON
    5 hours ago
  • $130k - $180k per year

     ...Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure...  ...implement, and maintain CI/CD pipelines, enabling rapid and reliable software releases. Automate and optimize our infrastructure... 
    Full time
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Acquird.io

    Toronto, ON
    6 hours ago
  •  ...customers. Cohere is a team of researchers, engineers, designers, and more, who are passionate...  ...building high-performance, scalable and reliable machine learning systems? Do you want to...  ...NLP applications? We are looking for a Site Reliability Engineer to join the Model... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    Toronto, ON
    6 hours ago
  •  ...San Francisco and founded in 2014, Tubi is part of Tubi Media Group, a division of Fox Corporation. About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset... 
    Remplacement
    Full time
    Contract work
    Temporary work
    Flexible hours

    Tubi

    Toronto, ON
    6 hours ago
  • $144k - $200k per year

    **The Team** Platform Engineering is the department within SRE that is responsible for a range...  ...role in developing and maintaining the reliable and globally connected multi-cloud network...  ...Overview** We are seeking a talented Site Reliability Engineer (SRE) with a strong... 
    Full time
    Work at office
    Remote work
    Worldwide
    Flexible hours

    Mongodb

    Toronto, ON
    6 hours ago
  •  ...visibility, and optimize spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s...  ...Excellence & Automation Design, develop, and automate reliable cloud infrastructure and platform services. Apply Infrastructure... 
    Full time
    Manual labor
    Local area
    Flexible hours

    Emburse

    Toronto, ON
    6 hours ago
  • $136k - $187k per year

     ...users worldwide. Our commitment to reliability is a key foundation of our product and...  ...availability expectations is a core engineering focus. As a Senior Site Reliability Engineer, you'll join...  ...solutions that make our system more reliable by design. What you’ll do: Design... 
    Local area
    Remote work
    Worldwide

    Okta

    Toronto, ON
    6 hours ago
  • $104.24k - $143.3k per year

     ...Microsoft Azure cloud technologies in a team that adopts DevOps methodologies at an innovative and growing company. Our Lead Site Reliability Engineers will provide a stable infrastructure platform throughout our clients initial onboarding journey with SimCorp -whether that... 
    Full time
    Work at office
    Immediate start
    Remote work
    Shift work
    Weekend work
    2 days per week

    SimCorp

    Toronto, ON
    5 days ago
  •  ...that are particularly strong in a few areas, and have some interest and capabilities in others. About the Role: As a Site Reliability Engineer , you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS... 
    Full time

    Kong Company

    Toronto, ON
    6 hours ago
  • $150k - $250k per year

     ...About The Role We're looking for a Senior Site Reliability Engineer to help us run one of the most exciting GPU clusters around—our Toronto datacenter packed with NVIDIA H100 and A100 GPUs, over 20PB of Ceph storage, terabit networking, and hundreds of servers. You'll... 
    Full time

    Boson Ai

    Toronto, ON
    6 hours ago
  •  ...Platform Engineer – DevOps, Site Reliability Engineering (SRE) & Dynatrace Required Skills • Strong experience as a Platform Engineer with expertise in DevOps and Site Reliability Engineering (SRE). • Experience designing, implementing, automating, and supporting... 
    Permanent employment

    Astra North Infoteck Inc.

    Toronto, ON
    24 days ago
  • $141k - $191k per year

     ...and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence. About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for: Team Leadership... 
    Work at office
    Local area
    Flexible hours
    2 days per week
    3 days per week

    Thomson Reuters

    Toronto, ON
    more than 2 months ago
  •  ...Site Reliability Engineer Location: Toronto ON Work Model: Hybrid (2 days per week in-person at the Toronto office preferred) Required Skills Site Reliability Engineering (SRE) DevOps Dynatrace Role Summary Design implement and optimize... 
    Full time
    Work at office
    2 days per week

    Astra North Infoteck Inc.

    Toronto, ON
    a month ago
  •  ...services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins...  ...over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil... 
    Long term contract
    Full time
    Work at office
    Local area
    Immediate start
    Remote work
    Worldwide
    Shift work

    Mongodb

    Toronto, ON
    6 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!