Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

Full-time

GuruLink

Location: REMOTE / Toronto, Ontario

This job allows you to work remotely.

We are partnering with a high-growth software company that operates a globally distributed, large-scale cloud platform supporting billions of daily transactions. The organization is focused on building highly available, data-intensive systems and is seeking a Lead Site Reliability Engineer to help shape the future of its infrastructure and reliability strategy.

This role combines hands-on technical leadership with long-term architectural ownership. You will play a key role in designing, scaling, and evolving mission-critical platform services across a multi-cloud environment while helping drive operational excellence, automation, observability, and platform reliability initiatives.

What You'll Do:

- Define and champion infrastructure automation strategies that reduce operational overhead, improve system performance, and enhance overall platform stability.

- Lead the design, reliability, and long-term evolution of core platform services, ensuring they align with business objectives and scalability requirements.

- Architect and guide the strategy for centralized logging and observability systems, balancing performance, availability, retention requirements, and operational costs.

- Establish frameworks for capacity planning, performance monitoring, and system optimization, proactively identifying opportunities to improve scalability and efficiency.

- Drive cross-functional reliability initiatives in partnership with engineering teams, influencing architectural decisions and promoting resilient service design practices.

- Identify systemic risks, operational bottlenecks, and platform improvement opportunities, taking ownership of solutions with a high degree of autonomy.

- Mentor engineers across the organization on reliability engineering principles, operational excellence, and distributed systems best practices.

Why This Opportunity:

- Work on systems operating at massive scale with demanding performance and reliability requirements.

- Influence platform strategy and architecture across a globally distributed cloud environment.

- Join a collaborative engineering culture that values technical excellence, ownership, and continuous improvement.

- Lead initiatives that directly impact the scalability, availability, and future growth of the platform.

Must Have Skills:

What We're Looking For:

- Demonstrated success in Site Reliability Engineering or Software Engineering roles, with experience designing, building, and operating highly scalable and resilient distributed systems.

- Deep expertise with distributed data and messaging platforms such as Apache Kafka, Apache Pulsar, ScyllaDB, Cassandra, Grafana Loki, or similar technologies.

- Strong experience creating and scaling automation frameworks, operational tooling, and performance analysis practices for large-scale production environments.

- 6+ years of hands-on experience in Site Reliability Engineering or Software Engineering, including ownership of cloud infrastructure strategy across AWS and GCP environments.

- Experience designing and leading observability platforms, including monitoring standards, service-level objectives (SLOs), alerting frameworks, and telemetry strategies.

- Exposure to technologies such as Prometheus, Grafana, Loki, Tempo, Thanos, or similar tooling is highly desirable.

- Proven track record of improving incident management processes, monitoring strategies, operational readiness, and on-call effectiveness within engineering organizations.

- Expert-level experience with Infrastructure as Code practices and tooling, including Terraform, Chef, or comparable platforms.

- Advanced Kubernetes expertise, including cluster architecture, multi-tenant environments, workload optimization, and large-scale container orchestration.

- Strong software development skills with experience working in multiple programming languages such as Go, Node.js, Python, Ruby, and shell scripting.

- Advanced Linux systems knowledge, including performance analysis, troubleshooting, tuning, and root-cause investigation of complex infrastructure issues.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Toronto, ON vacancy
  • $154k - $200k per year

     ...global client base with operations throughout North America, Central America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic technical leadership across infrastructure and software... 
    Suggested
    Long term contract
    Full time

    Movable Ink

    Toronto, ON
    more than 2 months ago
  • $100k - $125k per year

     ...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,...  ...management, tax compliance, and treasury. Tipalti partners with leading financial institutions such as Citi, Wells Fargo, J.P.... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Tipalti

    Toronto, ON
    a month ago
  • $110k - $130k per year

     ...others, and it defines our culture. About the job As a Site Reliability Engineer II on the Serving Platforms team within Infrastructure Engineering...  ...with cross-functional engineering teams worldwide, lead greenfield infrastructure projects, resolve complex incidents... 
    Suggested
    Full time
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    Opentable

    Toronto, ON
    6 days ago
  • $110k - $120k per year

     ...behind Money Mart—Canada’s largest non-bank branch network—and a leader in financial solutions for underserved communities. From...  ...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring... 
    Suggested
    Full time
    Temporary work
    Internship
    Work at office
    Remote work

    Momentum Financial Services Group

    Toronto, ON
    more than 2 months ago
  • $140k - $155k per year

     ...This is a hands-on senior engineering role focused on improving production...  ...teams to ship secure, reliable, and scalable software with confidence...  ...across the organization. Lead response efforts for high-severity...  ...on cloud-native technologies, site reliability engineering... 
    Suggested
    Remote job
    Permanent employment
    Full time
    Flexible hours

    Caseware

    Toronto, ON
    a month ago
  • $140k - $182k per year

     ...America, Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and...  ...on Cloud platforms (AWS/GCP) ~ Experience architecting and leading large-scale observability platforms, including defining observability... 
    Full time

    Movable Ink

    Toronto, ON
    more than 2 months ago
  • $153.82k - $277k per year

     ...a place where you can thrive, we can’t wait to meet you. Site Reliability Engineers (SREs) at Braze are responsible for keeping all internal-facing...  ...NGINX and Kubernetes is important. WHAT YOU'LL DO Lead NGINX & Kubernetes Ingress Infrastructure Architect and Operate... 
    Permanent employment
    Full time
    Internship
    Work at office
    Local area
    Remote work
    Flexible hours
    Rotating shift

    Braze

    Toronto, ON
    a month ago
  •  ...belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an engineer, a builder, and a...  ..., and documentation over process. You’ll engage in and often lead architectural discussions, reduce toil, and deliver scalable,... 
    Full time
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to friday
    Flexible hours

    Imanage

    Toronto, ON
    more than 2 months ago
  •  ...competitive advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in software...  ...experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly evaluate products... 
    Full time

    Serigor Inc

    Toronto, ON
    more than 2 months ago
  • $80 - $110 per hour

     ...Singapore and also operating in Denmark, Spain and Vietnam. The Site Reliability Engineer  will improve the availability, performance, scalability and...  ...device operations and routine production changes.   ~ Lead technically during incidents, drive evidence-based learning... 
    Remote job
    Full time
    Contract work

    Axon-networks

    Toronto, ON
    a month ago
  • $110k - $125k per year

     ...healthcare stakeholders to collect and exchange data with our leading FHIR-based data liberation platform.   At its heart, the...  ...today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability... 
    Full time
    Remote work
    Flexible hours

    Smile Digital Health

    Toronto, ON
    more than 2 months ago
  • $130k - $180k per year

     ...Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure (...  ...systems remain stable and responsive even during off-hours. Lead the development, implementation, and achievement of service-... 
    Full time
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Acquird.io

    Toronto, ON
    more than 2 months ago
  •  ...visibility, and optimize spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s...  ...standards for scalability, observability, and fault tolerance. Lead cross-functional troubleshooting of complex issues spanning... 
    Full time
    Manual labor
    Local area
    Flexible hours

    Emburse

    Toronto, ON
    more than 2 months ago
  •  ...customers. Cohere is a team of researchers, engineers, designers, and more, who are passionate...  ...building high-performance, scalable and reliable machine learning systems? Do you want to...  ...NLP applications? We are looking for a Site Reliability Engineer to join the Model... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    Toronto, ON
    more than 2 months ago
  • $144k - $200k per year

    **The Team** Platform Engineering is the department within SRE that is responsible for a range...  ...role in developing and maintaining the reliable and globally connected multi-cloud network...  ...Overview** We are seeking a talented Site Reliability Engineer (SRE) with a strong... 
    Full time
    Work at office
    Remote work
    Worldwide
    Flexible hours

    Mongodb

    Toronto, ON
    more than 2 months ago
  •  ...Employment Status: Permanent Schedule: 40 hours/week – 100% remote work Job Description We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability, performance, and resilience of high-availability SaaS platforms. Working... 
    Permanent employment
    Work at office
    Local area
    Remote work

    TOTEM Recruteur de talent

    Toronto, ON
    7 days ago
  • $130k - $150k per year

     ...engagement worldwide, and we're looking for an Intermediate Site Reliability Engineer to join our Engineering team.   About the job We're looking...  ..., and repetitive manual work is part of the role. We want reliable systems and sustainable operations for the people supporting... 
    Full time
    Work at office
    Remote work
    Worldwide
    1 day per week

    ContactMonkey

    Toronto, ON
    14 hours ago
  •  ...that combines Strategy, Experience & Design, Engineering and Managed Services. We build digital solutions...  ...response and run post-incident reviews that lead to real fixes. QUALIFICATIONS ~6+ years in SRE, platform reliability or observability engineering, with strong... 
    Full time

    Appnovation Technologies

    Toronto, ON
    5 days ago
  •  ...s possible. Join us and help the world’s leading organizations unlock the value of technology...  ...is looking for a Production support Engineer to work for the Commercial Line of Business...  ...and other stakeholders to ensure system reliability and availability. Documentation: Maintain... 
    Permanent employment
    Full time
    Local area

    Capgemini

    Toronto, ON
    15 days ago
  • Locations & Program Hours # Lake Simcoe Public School (38 Thornlodge Drive, Keswick) 2:40pm to 5:30pm # Crosby Heights Public School (190 Neal Drive, Richmond Hill) 2:40pm to 5:30pm # Maple Leaf Public School (155 Longford Drive, Newmarket) 2:40pm to 5:30pm BGC...
    Full time
    Contract work
    Seasonal work
    Local area
    Monday to friday
    Afternoon shift

    Bgc St. Alban's Club

    Toronto, ON
    1 day ago
  • $120k - $170k per year

     ...Going   Magnet Forensics is a global leader in the development of digital investigative...  ...skilled and motivated Senior DevOps Engineer to join our dynamic team and play a key role...  ...incidents, and drive improvements that increase reliability and operational efficiency; Write and... 
    Full time
    Work at office
    Local area
    Flexible hours

    Magnet Forensics

    Toronto, ON
    a month ago
  •  ...programs. We are seeking enthusiastic, reliable, and motivated individuals to join our...  ...Program team for the 2026–2027 school year. Site Coordinators oversee the daily operations...  ...Development Certification. Experience leading sports, arts, recreation, educational, health... 
    Full time
    Contract work
    Seasonal work
    Local area
    Monday to friday

    Bgc St. Alban's Club

    Toronto, ON
    more than 2 months ago
  • $110k - $130k per year

     ...others, and it defines our culture. About the job As a Site Reliability Engineer II on the Serving Platforms team within Infrastructure Engineering...  ...with cross-functional engineering teams worldwide, lead greenfield infrastructure projects, resolve complex incidents... 
    Work at office
    Local area
    Worldwide
    Flexible hours
    2 days per week

    OpenTable

    Toronto, ON
    9 days ago
  • $116k - $235.1k per year

     ...About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that...  ...Participate in a 24/7 on-call rotation, acting as a key technical leader and incident commander during critical service disruptions.... 
    Long term contract
    Remplacement
    Full time
    Temporary work
    Work at office
    Local area
    Immediate start
    Flexible hours
    2 days per week

    Tubi - Canada

    Toronto, ON
    22 days ago
  •  ...AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work...  ...storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking... 
    Full time
    Remote work

    bosonai

    Toronto, ON
    a month ago
  • $115k - $125k per year

     ...Job Summary   The Fixed Plant Asset Reliability Engineer, as part of the Asset Management team, plays...  ...the asset management program. Will also lead and mentor junior engineers, and...  ...position requires 50%+ travel to remote sites overseas, ensuring that global operations... 
    Long term contract
    Temporary work
    For contractors
    Casual work
    Local area
    Immediate start
    Remote work
    Overseas

    Kinross Gold Corporation

    Toronto, ON
    5 hours ago
  •  ...Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance...  .... Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy...  ...products. A strong problem-solver who can lead root-cause investigations on failures... 
    Permanent employment
    Full time
    Internship
    Second job

    Tenstorrent

    Toronto, ON
    more than 2 months ago
  • $141k - $191k per year

     ...Reuters and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence. About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for: Team... 
    Work at office
    Local area
    Flexible hours
    2 days per week
    3 days per week

    Thomson Reuters

    Toronto, ON
    more than 2 months ago
  • $186k - $236k per year

     ...role / impact You will focus on solving engineering problems at scale, moving beyond feature...  ...ensures our platform remains robust and reliable for millions of global users. As a senior...  ...essential. You possess the ability to lead major code design decisions and... 
    Long term contract
    Full time
    Work at office
    Work from home
    Relocation
    2 days per week
    1 day per week

    Xero

    Toronto, ON
    more than 2 months ago
  • $172k - $229k per year

     ...BuildOps is looking for a Staff Software Engineer to set and drive our company-wide technical...  ...strategy for building, shipping, and operating reliable software. This is a high-impact, cross-...  ...of customer-impacting failures and lead cross-team initiatives that address root causes... 
    Long term contract
    Permanent employment
    Full time
    For contractors
    Work at office
    Local area
    Work from home
    Flexible hours

    Buildops

    Toronto, ON
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!