Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead, Site Reliability Engineer

$140k - $180k per year
Full-time

Tripstack



About Tripstack

Founded in Toronto, Canada in 2016, Tripstack has been part of Etraveli Group since 2019. It is a B2B Flights as a Service provider and a world leader in virtual interlining.

Operating from offices in Canada, India, and Poland, Tripstack is the gateway into Etraveli Group’s world leading tech platform - giving partners access to global flight content, virtual interlining, and a full suite of services including payments, fraud prevention, pricing, and customer support. As a world leader in virtual interlining technology, Tripstack connects non-partner, low-cost and full-service carriers, enabling the creation of unique and flexible itineraries through a simple, cost-effective API.

Its technology ingest over 30B price points and handles over 240 million searches daily.

Through partnerships with airlines, OTAs, and other distribution channels across the globe, Tripstack expands networks, drives new revenue streams, and offers more choice at competitive prices, all backed by robust technology and traveler protection.

 

For more information, visit:

The role

Our SRE team runs the shared infrastructure that all other engineering teams at Tripstack depend on. This spans three Kubernetes environments today - GKE on GCP as the primary platform for booking and search workloads, a multi-tenant Kubernetes-on-OpenStack cluster running Talos Linux for content acquisition and SRE tooling, and an isolated PCI-compliant environment for payment processing - plus a substantial VM footprint, self-hosted Concourse CI, and a Prometheus, Thanos, and Grafana observability stack.

We are in the middle of a major infrastructure transition. Our Toronto data centre closes in September 2026, and the first phase of migrating the Tripstack Platform to a new OpenStack and Talos-based Kubernetes environment in Gothenburg, Sweden - built in partnership with our parent company Etraveli Group under a shared responsibility model - is in progress. Over the next year we intend to further optimize workloads across our environments. The most immediate priority for this role is moving our GCP footprint from the US to the EU by the end of 2026 with a zero downtime objective.

This role has two parts. First, lead the Toronto SRE team day-to-day: on-call, incident management, delivery, and developing engineers who can communicate risk and progress clearly to leadership. Second, shape the overall strategy for application deployments across GCP and Gothenburg - which workloads move, which stay, where the PII boundary sits, and the deployment patterns and reliability standards that apply across both.

This is a hands-on leadership role on a distributed team spanning Toronto, Pune, and Kraków, working closely with engineering counterparts in Stockholm and Gothenburg.

Responsibilities

Lead the Toronto SRE team

  • Provide day-to-day technical and people leadership for SRE engineers based in Canada: priorities, delivery, code and change review, career development

  • Own on-call and incident management for shared infrastructure: sustainable rotations, current runbooks, and post-incident reviews that result in concrete improvements

  • Develop engineers who communicate well with leadership: your team should be able to present migration risks and reliability trade-offs to senior stakeholders directly

  • Coordinate closely with SRE and engineering colleagues in Pune and Kraków so that ownership and handoffs across time zones are clearly defined

Shape deployment strategy across GCP and Gothenburg

  • Own the GCP US-to-EU region migration end to end: planning, sequencing, cutover, and validation for live production traffic, with completion targeted by the end of 2026

  • Define the workload placement strategy: which workloads move to the OpenStack/Talos environment in Gothenburg, which remain on GCP, and how the requirement that PII stays in GCP EU is enforced

  • Establish a consistent deployment model across both platforms: Concourse pipelines, Helm, and GitOps patterns that work the same way whether the target is GCP or Gothenburg

  • Consolidate the two identity systems we operate today - LDAP/Keystone for the OpenStack estate and GCP IAM for cloud - into a clear and consistent access model

  • Define the standard for how application teams onboard workloads: node pools, namespaces, quotas, network policy, and secrets management, all documented and consistent

Deliver the data centre migration

  • Co-own execution of the Gothenburg migration with our Sr. Manager, SRE and ETG ITOPS counterparts, within the shared responsibility model - Tripstack owning the Kubernetes control plane and everything above the hypervisor, adopting ETG standards below it

  • Complete phase one, then plan and execute the subsequent migration waves - moving business-critical services from single-homed to fully redundant, with failover scenarios enabled and tested regularly

  • Own the network architecture of a cross-Atlantic platform: peering between Canada, Sweden, and GCP EU, and latency requirements for critical paths

  • Decommission legacy infrastructure as part of the migration: legacy Terraform, Puppet-managed VMs, and VM-based tooling that has a Kubernetes-native replacement

Raise the reliability bar

  • Build a formal SLO framework: SLIs, error budgets, and dashboards for our critical APIs, so that reliability decisions are based on data

  • Extend the Prometheus / Thanos / Grafana / alerting stack so that both platforms, and the migration itself, are fully observable

  • Make post-incident follow-through, capacity planning, and change safety standard practice across the engineering organization

Own the security posture of the infrastructure

  • Partner with our Security team on infrastructure security: least-privilege access across both identity systems, secrets management through Vault, network segmentation, and a consistent patching and vulnerability-management cadence

  • Operate our PCI-scoped environment to its required standard: change control, access reviews, audit evidence, and explicit lead approval for production changes - maintained throughout the migration

  • Enforce the data-residency requirement that PII stays in the EU region, through both policy and technical controls, with compliance treated as an integral part of migration planning

  • Ensure infrastructure leaving service is decommissioned securely: credentials rotated, access revoked, and data destruction documented

Requirements

  • Proven experience leading an SRE, platform, or infrastructure team, including running on-call rotations, acting as incident commander, managing performance, and developing engineers into senior roles

  • Deep production Kubernetes experience including self-managed or bare-metal clusters: cluster lifecycle, upgrades, networking (CNI, ingress, load balancing), and multi-tenant isolation

  • Strong GCP experience - GKE, IAM, VPC networking, Cloud SQL - ideally including a region or cross-region migration with production traffic

  • Senior-level Infrastructure as Code and GitOps experience - Terraform, Helm, and pipeline-as-code CI/CD; experience migrating away from legacy configuration management (Puppet, Ansible) is directly relevant

  • Strong observability and reliability practices - Prometheus, Grafana, SLOs, error budgets, and a track record of documentation and runbooks that prevent repeat incidents

  • Security-minded operations - least-privilege access design, secrets management (Vault or equivalent), network policy, and patching discipline, with security treated as a core part of reliability

  • Experience with a major infrastructure transition - a data centre migration, cloud migration, or platform rebuild with production traffic

  • 8+ years in production infrastructure, at least 2 in a lead or management role, with accountability for business-critical systems

  • Clear written and verbal English; based in the Greater Toronto Area and comfortable working across Toronto, Pune, Kraków, and Stockholm time zones

Additional Experience That Would Be Considered An Asset

  • OpenStack operations experience - Neutron networking, Cinder/Ceph storage, Octavia load balancing, Keystone identity

  • Talos Linux or another immutable, API-managed Kubernetes OS in production

  • Concourse CI or comparable pipelines-as-code platforms at organizational scale

  • Experience operating PCI-scoped or similarly regulated environments, including change control and audit discipline

  • Experience running infrastructure under a shared responsibility model with a parent company, partner, or major vendor, including cross-organization coordination

  • Regular use of agentic coding tools (Claude Code, Gemini, or equivalent) in your workflow, with sound judgement about validating AI-generated configuration before production use

Nice to have

  • Operational exposure to data systems our SRE team supports - Druid, Redpanda, Elasticsearch, Airflow

  • Network engineering depth - interconnects, BGP, site-to-site VPN, cross-region peering

  • GDPR data-residency, SOC 2, or ISO 27001 experience - the EU migration makes this increasingly relevant

  • Exposure to travel, flights, or large-scale search and cache systems

Compensation:
Canada - Toronto Office : 140, 000 - 180,000 CAD / Annual

Our pay ranges reflect the minimum and maximum target for new hire pay for the full-time position determined by role, level, and location.The pay range shown is based on our compensation structure in place at the time of posting and may be updated periodically based on business needs. Individual pay is based on additional factors including job-related skills, experience, and relevant education and/or training.

The targeted pay range listed reflects the base pay only and does not include bonus, or other benefits.

We use AI in our hiring process.

Benefits

We offer an opportunity to work with a young, dynamic, and a growing team composed of high-caliber professionals. We value professionalism and promote a culture where individuals are encouraged to do more and be more. If you feel you share our passion for excellence, and growth, then look no further. We have an ambitious mission, and we need a world-class team to make it a reality. Upgrade to a First Class team!

At Tripstack, we proudly believe in embracing diversity. This is true for our team, clients, communities and stakeholders. We are an equal opportunity employer and committed to creating a safe, healthy and accessible environment. We encourage applications regardless of race, colour, ancestry, religion, sex, national origin, sexual orientation, age, citizenship, marital status, disability, gender identity or any other grounds protected by law. Please let us know if you need any accommodations during any part of the recruitment process.

Tripstack thanks all applicants for their interest, however only those selected to continue in the process will be contacted.

Learn more about us at

#tripstack

Vacancy posted 20 hours ago
Similar jobs that could be interesting for youBased on the Lead, Site Reliability Engineer in Toronto, ON vacancy
  • $154k - $200k per year

     ...global client base with operations throughout North America, Central America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic technical leadership across infrastructure and software... 
    Suggested
    Long term contract
    Full time

    Movable Ink

    Toronto, ON
    20 hours ago
  • $100k - $125k per year

     ...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,...  ...management, tax compliance, and treasury. Tipalti partners with leading financial institutions such as Citi, Wells Fargo, J.P.... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Tipalti

    Toronto, ON
    20 hours ago
  • $110k - $120k per year

     ...behind Money Mart—Canada’s largest non-bank branch network—and a leader in financial solutions for underserved communities. From...  ...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring... 
    Suggested
    Full time
    Temporary work
    Internship
    Work at office
    Remote work

    Momentum Financial Services Group

    Toronto, ON
    20 hours ago
  •  ...world and help us reinvent the way people learn, because learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident response while also shaping the underlying infrastructure that supports the... 
    Suggested
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    20 hours ago
  • $153.82k - $277k per year

     ...a place where you can thrive, we can’t wait to meet you. Site Reliability Engineers (SREs) at Braze are responsible for keeping all internal-facing...  ...NGINX and Kubernetes is important. WHAT YOU'LL DO Lead NGINX & Kubernetes Ingress Infrastructure Architect and Operate... 
    Suggested
    Permanent employment
    Full time
    Internship
    Work at office
    Local area
    Remote work
    Flexible hours
    Rotating shift

    Braze

    Toronto, ON
    20 hours ago
  • $140k - $182k per year

     ...America, Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and...  ...on Cloud platforms (AWS/GCP) ~ Experience architecting and leading large-scale observability platforms, including defining observability... 
    Full time

    Movable Ink

    Toronto, ON
    20 hours ago
  • $140k - $155k per year

     ...This is a hands-on senior engineering role focused on improving production...  ...teams to ship secure, reliable, and scalable software with confidence...  ...across the organization. Lead response efforts for high-severity...  ...on cloud-native technologies, site reliability engineering... 
    Remote job
    Permanent employment
    Full time
    Flexible hours

    Caseware

    Toronto, ON
    20 hours ago
  •  ...belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an engineer, a builder, and a...  ..., and documentation over process. You’ll engage in and often lead architectural discussions, reduce toil, and deliver scalable,... 
    Full time
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to friday
    Flexible hours

    Imanage

    Toronto, ON
    20 hours ago
  •  ...help us reinvent the way people learn, because learning never stops. The Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team of engineers dedicated to safeguarding the operational health and resilience of the Docebo... 
    Long term contract
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    20 hours ago
  •  ...competitive advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in software...  ...experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly evaluate products... 
    Full time

    Serigor Inc

    Toronto, ON
    20 hours ago
  • $80 - $110 per hour

     ...Singapore and also operating in Denmark, Spain and Vietnam. The Site Reliability Engineer  will improve the availability, performance, scalability and...  ...device operations and routine production changes.   ~ Lead technically during incidents, drive evidence-based learning... 
    Remote job
    Full time
    Contract work

    Axon-networks

    Toronto, ON
    20 hours ago
  •  ...visibility, and optimize spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s...  ...standards for scalability, observability, and fault tolerance. Lead cross-functional troubleshooting of complex issues spanning... 
    Full time
    Manual labor
    Local area
    Flexible hours

    Emburse

    Toronto, ON
    20 hours ago
  • $130k - $180k per year

     ...Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure (...  ...systems remain stable and responsive even during off-hours. Lead the development, implementation, and achievement of service-... 
    Full time
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Acquird.io

    Toronto, ON
    20 hours ago
  • $110k - $125k per year

     ...healthcare stakeholders to collect and exchange data with our leading FHIR-based data liberation platform.   At its heart, the...  ...today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability... 
    Full time
    Remote work
    Flexible hours

    Smile Digital Health

    Toronto, ON
    20 hours ago
  • $144k - $200k per year

    **The Team** Platform Engineering is the department within SRE that is responsible for a range...  ...role in developing and maintaining the reliable and globally connected multi-cloud network...  ...Overview** We are seeking a talented Site Reliability Engineer (SRE) with a strong... 
    Full time
    Work at office
    Remote work
    Worldwide
    Flexible hours

    Mongodb

    Toronto, ON
    20 hours ago
  •  ...customers. Cohere is a team of researchers, engineers, designers, and more, who are passionate...  ...building high-performance, scalable and reliable machine learning systems? Do you want to...  ...NLP applications? We are looking for a Site Reliability Engineer to join the Model... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    Toronto, ON
    20 hours ago
  •  ...interest and capabilities in others. About the Role: As a Site Reliability Engineer , you’ll join the global Platform SRE team responsible for...  ...of operational playbooks and postmortem practices. Lead and contribute to scaling initiatives that improve elasticity... 
    Full time

    Kong Company

    Toronto, ON
    20 hours ago
  • $140k - $165k per year

     ...transformation, transitioning from traditional service desk operations to an AI-driven, self-healing operational fabric. As Lead Site Reliability Engineer (SRE), you will lead enterprise observability, AIOps automation, and SRE governance across a global footprint. You will... 

    Randstad

    Toronto, ON
    14 days ago
  •  ...AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work...  ...storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking... 
    Full time
    Remote work

    bosonai

    Toronto, ON
    12 days ago
  •  ...Employment Status: Permanent Schedule: 40 hours/week – 100% remote work Job Description We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability, performance, and resilience of high-availability SaaS platforms. Working... 
    Permanent employment
    Work at office
    Local area
    Remote work

    TOTEM Recruteur de talent

    Toronto, ON
    8 days ago
  • $120k - $170k per year

     ...Going   Magnet Forensics is a global leader in the development of digital investigative...  ...skilled and motivated Senior DevOps Engineer to join our dynamic team and play a key role...  ...incidents, and drive improvements that increase reliability and operational efficiency; Write and... 
    Full time
    Work at office
    Local area
    Flexible hours

    Magnet Forensics

    Toronto, ON
    20 hours ago
  •  ...programs. We are seeking enthusiastic, reliable, and motivated individuals to join our...  ...Program team for the 2026–2027 school year. Site Coordinators oversee the daily operations...  ...Development Certification. Experience leading sports, arts, recreation, educational, health... 
    Full time
    Contract work
    Seasonal work
    Local area
    Monday to friday

    Bgc St. Alban's Club

    Toronto, ON
    20 hours ago
  •  ...youll be doing As a member of CIBCs Application Reliability Engineering Platform team the Consultant Site Reliability Engineering will play a key role in...  ...before client impact. Participate in and as required lead incident response and post-mortem analysis to identify... 
    Full time
    Contract work
    3 days per week
    1 day per week

    Canadian Imperial Bank of Commerce

    Toronto, ON
    20 days ago
  •  ...Senior Site Reliability Engineer - Edge Location : Ottawa/Toronto, On-Site Reports to: Head of Security The Role We build rugged...  .... This is a build-from-scratch role. You'll define what "reliable" means for hardware that has to keep working in harsh, sometimes... 
    Full time

    dominion%20dynamics

    Toronto, ON
    1 day ago
  •  ...Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance...  .... Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy...  ...products. A strong problem-solver who can lead root-cause investigations on failures... 
    Permanent employment
    Full time
    Internship
    Second job

    Tenstorrent

    Toronto, ON
    20 hours ago
  • $100k per year

     ...Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost...  ...seniorities. Tenstorrent is looking for an experienced Reliability Engineer to drive and execute reliability testing for our products from... 
    Permanent employment
    Full time
    Internship

    Tenstorrent

    Toronto, ON
    20 hours ago
  •  ...single device. This approach allows Cerebras to deliver industry-leading training and inference speeds and empowers machine learning...  .... About The Role Join Cerebras as a Performance & Reliability Engineer within our innovative Co-Design and Next Generation Team. Our... 
    Full time

    Cerebras Systems

    Toronto, ON
    20 hours ago
  •  ...to reimagine what’s possible. Join us and help the world’s leading organizations unlock the value of technology and build a more...  ...PostgreSQL) This role is part of the Production Support and Reliability Engineering team, responsible for ensuring the stability, availability,... 
    Permanent employment
    Full time
    Local area

    Capgemini

    Toronto, ON
    28 days ago
  • $78.62 - $89.34 per hour

    Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering and Operations team. In this engineering-focused role, you will move beyond traditional database administration to act... 
    Long term contract
    Permanent employment
    Full time
    Contract work
    Work at office

    Randstad

    Toronto, ON
    a month ago
  • $172k - $229k per year

     ...BuildOps is looking for a Staff Software Engineer to set and drive our company-wide technical...  ...strategy for building, shipping, and operating reliable software. This is a high-impact, cross-...  ...of customer-impacting failures and lead cross-team initiatives that address root causes... 
    Long term contract
    Permanent employment
    Full time
    For contractors
    Work at office
    Local area
    Work from home
    Flexible hours

    Buildops

    Toronto, ON
    20 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead, Site Reliability Engineer. Be the first to apply!