Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Manager, Site Reliability Engineering

Full-time

Tubi

About Tubi:


Boldly built for every fandom, Tubi is a free streaming service that entertains over 100 million monthly active users. Tubi offers the world's largest collection of Hollywood movies and TV shows, thousands of creator-led stories and hundreds of Tubi Originals made for the most passionate fans. Headquartered in San Francisco and founded in 2014, Tubi is part of Tubi Media Group, a division of Fox Corporation.

About the Role:

Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission is to engineer resilience from the ground up, enabling our product teams to innovate rapidly while ensuring our users have a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through a culture of data-driven decision-making, blameless learning, and relentless automation.

We are seeking an experienced and visionary Senior SRE Manager to lead and grow our newly built Site Reliability Engineering team. You are more than a people manager or a tech lead; you are the strategic leader responsible for architecting our reliability roadmap. You will build and mentor a team of talented engineers, foster a culture of blameless learning and continuous improvement, and champion the engineering practices that allow us to balance rapid innovation with rock-solid stability. You will be a key influencer in our engineering leadership, partnering with peers across the organization to ensure reliability is a shared responsibility and a core tenet of our engineering culture.

What You'll Do:


  • Team Leadership & Mentorship:


    • Lead, mentor, and grow a team of Site Reliability Engineers. Foster a culture of innovation and technical excellence where engineers feel empowered to do their best work. Provide personalized coaching, create professional development plans, and guide the careers of senior and emerging talent within the team.

    • Establish equitable, sustainable on-call practices (including global coverage where applicable) that protect focus time and avoid burnout.

    • Define team rituals - runbook reviews, game days, and incident retros - that reinforce quality and learning.

  • Strategic Planning & Vision: Define and drive the multi-year technical strategy and vision for Tubi’s observability, and automation platforms. Partner with infra lead to align Tubi’s infrastructure & SRE roadmap. Partner with tech leaders to align the SRE roadmap with business objectives. Champion a data-driven approach to reliability, using Service Level Objectives (SLOs) and error budgets to facilitate productive conversations about risk and feature velocity.

  • Operational Excellence & Incident Management:  


    • Own the end-to-end availability, performance, and efficiency of our critical user-facing services. Evolve our incident response practice to reduce Mean Time to Resolution (MTTR) and Mean Time Between Failures (MTBF). Champion a rigorous, blameless, and data-driven post-mortem culture to ensure we learn from both successes and failures, driving eng teams for systemic fixes and automation to prevent the recurrence of incidents.

    • Streamline and improve our existing processes and practices, and collaborate with other teams to enhance our production release standards by improving current processes.

    • Define and tune a 24×7 on-call rotation for low noise and fast response; act as executive escalation partner during major incidents.

    • Own disaster-recovery strategy (playbooks, failover drills, recovery simulations) and track SLO gaps with time-bound remediations.

  • Financial & Vendor Management: Own the SRE budget, tooling, and headcount. Manage relationships with key third-party vendors for our observability and SRE related AI platforms, work with infra lead and finance team for contract negotiations and ensure we derive maximum value from our investments.

  • Cross-Functional Collaboration: Act as a key influencer and strategic partner to leaders in Software Engineering, Product Management, and Infra/Sec. Drive the adoption of SRE best practices and principles throughout the organization, ensuring new services are designed for reliability, scalability, and observability from day one.

Your Background:


  • 8+ years of experience in a technical field, with at least 3+ years in an engineering leadership position managing SRE, DevOps, or Production Engineering teams.

  • A deep, principled understanding of SRE tenets, including Service Level Indicators (SLIs), SLOs, error budgets, toil reduction, and capacity planning.

  • Exceptional communication, negotiation, and influencing skills, with the ability to articulate complex technical concepts and strategies to both technical and non-technical stakeholders at all levels of the organization.

  • A strong technical background as a hands-on software engineer or site reliability engineer prior to moving into management. Deep knowledge of AWS services (especially networking, IAM, EKS, ALBs/NLBs, Route 53, CloudWatch). Proven experience with Kubernetes in production (EKS preferred), including service exposure, networking, and availability engineering.

  • Hands-on familiarity with modern SRE tools and technologies, including Infrastructure as Code (e.g., Terraform, Ansible), container orchestration (Kubernetes), observability platforms (e.g., Prometheus, Grafana, Datadog, Splunk), and incident tooling (e.g., PagerDuty, FireHydrant), deployment-safety tooling (e.g., Argo Rollouts, LaunchDarkly), and observability standards (e.g., OpenTelemetry).

Preferred Qualifications (Nice-to-Haves)



  • Executive-caliber incident communication/storytelling skills (clear status, stakeholder alignment, and post-incident narratives).

  • Demonstrated success in hiring, developing, and mentoring high-performing engineers, including managing senior and principal-level talent.

  • Experience managing globally distributed teams and developing equitable and sustainable on-call rotation practices.

  • Experience in financial planning, budget management, and vendor contract negotiation for technical infrastructure and tooling.

The AI Mandate: Building the Future of Observability with AI


You will not just manage a team that uses AI; you will lead the charge in building an AI-native SRE function. This is a strategic mandate that requires a forward-thinking leader who understands both the potential and the pitfalls of integrating intelligent systems into critical operations. This includes:


  • AIOps Strategy Development: Developing and executing the strategy for integrating AIOps and machine learning into our observability stack. Your goal will be to move the team from a reactive monitoring posture to one of predictive maintenance and automated anomaly detection, fundamentally changing how we ensure reliability.

  • Accelerating Automation with AI: Championing the effective and responsible use of AI-assisted coding tools (e.g., Claude Code, Cursor) within the SRE team. You will set the standards and practices to leverage these tools to accelerate the development of automation, operational tooling, and infrastructure code.

  • Building the Business Case: Building the techno-economic case for new AI tooling, managing vendor relationships, and ensuring the cost-effective and secure implementation of these powerful systems. You must be able to articulate the ROI of these investments in terms of reduced downtime, improved operational efficiency, and faster incident resolution.

  • Fostering Critical AI Literacy: Fostering a culture that can critically evaluate, debug, and learn from the outputs of AI systems. This involves extending our blameless post-mortem philosophy to AI-driven actions and recommendations, ensuring that the team remains in control and understands the "why" behind automated decisions.

#LI-Hybrid


Tubi is a division of Fox Corporation, and the FOX Employee Benefits summarized  here , covers the majority of all US employee benefits. The following distinctions below outline the differences between the Tubi and FOX benefits:


  • For US-based non-exempt Tubi employees, the FOX Employee Benefits summary accurately captures the Vacation and Sick Time.

  • For all salaried/exempt employees, in lieu of the FOX Vacation policy, Tubi offers a Flexible Time off Policy to manage all personal matters.

  • For all full-time, regular employees, in lieu of FOX Paid Parental Leave, Tubi offers a generous Parental Leave Program, which allows parents twelve (12) weeks of paid bonding leave within the first year of birth, adoption, surrogacy, or foster placement of a child in addition to applicable government leave program(s) and FOX’s short-term disability policy. This time is 100% paid through a combination of any applicable state, city, and federal leaves and wage-replacement programs in addition to contributions made by Tubi.

  • For all full-time, regular employees, Tubi offers a monthly wellness reimbursement.

We are an equal opportunity employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, sex, national origin, gender identity, disability, protected veteran status, or any other characteristic protected by law. We will consider for employment qualified applicants with criminal histories consistent with applicable law.

Vacancy posted 7 hours ago
Similar jobs that could be interesting for youBased on the Senior Manager, Site Reliability Engineering in Toronto, ON vacancy
  • $243k - $297k per year

     ...businesses. As Relay continues to scale, the reliability, performance, and resilience of our...  ...and business success. This is a senior leadership role responsible not only for guiding a strong team of Site Reliability Engineers, but for shaping how reliability strategy... 
    Senior
    Long term contract
    Full time
    Internship
    Work at office
    Trial period
    Flexible hours

    Relay

    Toronto, ON
    7 hours ago
  •  ...learning platform that helps organizations create, deliver, and manage training all in one place. But our real mission goes deeper:...  ..., because learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident... 
    Senior
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    7 hours ago
  •  ...best of both work styles in a workplace that is intentional about belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an engineer, a builder, and a systems thinker. You’ll create middleware and platform... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to friday
    Flexible hours

    Imanage

    Toronto, ON
    7 hours ago
  •  ...About the Role Fivetran is looking for a high-performance engineer to be a part of a team of Site Reliability Engineers. You will be working closely with engineering teams, product managers, as well as support and sales engineers to build the future of the Fivetran Data... 
    Senior
    Full time
    Work at office
    Remote work

    Fivetran

    Toronto, ON
    7 hours ago
  • $140k - $182k per year

     ...global client base with operations throughout North America, Central America, Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and software development. You will support and help inform the... 
    Senior
    Full time

    Movable Ink

    Toronto, ON
    7 hours ago
  •  ...scale, multi-tenant SaaS platform running reliably for customers around the world. You'll...  ...raise the reliability bar across the wider engineering org. What You'll Be Doing Take point...  .... ~ Comfort writing code or scripts to manage infrastructure, with a genuine preference... 
    Senior
    Full time
    Immediate start
    Flexible hours

    HighlightTA

    Toronto, ON
    7 hours ago
  •  ...learning platform that helps organizations create, deliver, and manage training all in one place. But our real mission goes deeper...  ...never stops. The Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team of engineers dedicated... 
    Long term contract
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    7 hours ago
  • $130k - $180k per year

     ...to legally work in Canada (visa or sponsorship won't be provided) Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure (those with Azure will be prioritized first) About Us:... 
    Senior
    Full time
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Acquird.io

    Toronto, ON
    7 hours ago
  •  ...Senior Site Reliability Engineer - Edge Location : Ottawa/Toronto, On-Site Reports to: Head of Security The Role You own the edge compute module — the standard hardware stack, OS image, and runtime that integrates with Dominion Dynamics's mesh radios, sensors... 
    Senior
    Full time

    dominion%20dynamics

    Toronto, ON
    7 hours ago
  •  ...and online privacy for all. Right now we are looking for a Site Reliability Engineer to help us tame DNS. This position can be hybrid at our...  ...configuration and troubleshooting Working with Salt configuration management software to admin hosts Write Python scripts for... 
    Senior
    Full time
    Direct hire
    Work at office
    Remote work

    Funded.club

    Toronto, ON
    7 hours ago
  • $136k - $187k per year

     ...hundreds of millions of users worldwide. Our commitment to reliability is a key foundation of our product and our dedication to exceeding customer availability expectations is a core engineering focus. As a Senior Site Reliability Engineer, you'll join our SRE team based in... 
    Senior
    Local area
    Remote work
    Worldwide

    Okta

    Toronto, ON
    7 hours ago
  • $110k - $120k per year

     ...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring...  ...reporting, SLO performance metrics, and incident trends to senior management. What You Bring : Technical Proficiency:... 
    Senior
    Full time
    Temporary work
    Internship
    Work at office
    Remote work

    Momentum Financial Services Group

    Toronto, ON
    7 hours ago
  •  ...interest and capabilities in others. About the Role: As a Site Reliability Engineer , you’ll join the global Platform SRE team responsible for...  ...or equivalent practical experience. Proven experience managing SaaS or PaaS systems at enterprise scale (multi-region,... 
    Senior
    Full time

    Kong Company

    Toronto, ON
    7 hours ago
  •  ...good fit for you! About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a...  ...Incident Commander for Incidents; coordinate cross-team response, manage communications, and ensure rapid service restoration.... 
    Remote job
    Long term contract
    Permanent employment
    Full time

    Tecsys Inc.

    Toronto, ON
    7 hours ago
  • $115k - $130k per year

     ...Kaseya Kaseya is the leading provider of AI-powered IT management and cybersecurity software, serving Managed Service Providers...  ...and continuously raising the bar.  Kaseya is hiring a Site Reliability Engineer to keep our production systems healthy as we scale. You'll... 
    Long term contract
    Full time
    Worldwide

    Kaseya Careers

    Toronto, ON
    7 hours ago
  • $100k per year

     ...growing our team and looking for contributors of all seniorities. Tenstorrent is building large-scale AI...  ...deployments. This role sits at the intersection of site reliability, infrastructure operations, and customer engineering, ensuring our systems are reliable, observable,... 
    Permanent employment
    Full time

    Tenstorrent

    Toronto, ON
    7 hours ago
  • $100k - $125k per year

     ...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer...  ...procurement, employee expenses, corporate cards, supplier management, tax compliance, and treasury. Tipalti partners with... 
    Full time
    Work at office
    Flexible hours

    Tipalti

    Toronto, ON
    8 hours ago
  •  ...technical support for the right fineness in management utilities at any time in a firm standing...  ...The SRE Role · SREs are engineers with the right mix of knowledge and skills...  ...observation to entire systems to improve reliability, performance and operability). · We constantly... 
    Full time

    Serigor Inc

    Toronto, ON
    7 hours ago
  • $154k - $200k per year

     ...America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with...  ...optimization Establish capacity planning and performance management frameworks, proactively identifying scaling opportunities and... 
    Long term contract
    Full time

    Movable Ink

    Toronto, ON
    7 hours ago
  •  ...and how we use AI in our recruiting process here . The Site Reliability Engineering organization at Pinterest is accountable for ensuring...  ..., to minimize risk and maximize the speed of innovation Manage capacity and performance to help scale our infrastructure both... 
    Senior
    Full time
    Work at office

    Pinterest

    Toronto, ON
    7 hours ago
  • $104.24k - $143.3k per year

     ...methodologies at an innovative and growing company. Our Lead Site Reliability Engineers will provide a stable infrastructure platform throughout our...  ...Site Reliability Engineer, you will be responsible for: Manage and maintain pipelines for client onboarding and live... 
    Senior
    Full time
    Work at office
    Immediate start
    Remote work
    Shift work
    Weekend work
    2 days per week

    SimCorp

    Toronto, ON
    5 days ago
  • $110k - $125k per year

     ...the impact of our innovative health data platform and data management solutions, which are used in over 20 countries. We were #19...  ...Apply today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability,... 
    Full time
    Remote work
    Flexible hours

    Smile Digital Health

    Toronto, ON
    7 hours ago
  •  ...visibility, and optimize spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s...  ...services. Apply Infrastructure-as-Code (IaC) principles to manage large-scale distributed systems. Write and maintain... 
    Full time
    Manual labor
    Local area
    Flexible hours

    Emburse

    Toronto, ON
    7 hours ago
  •  ...Cohere is a team of researchers, engineers, designers, and more, who are...  ...-performance, scalable and reliable machine learning systems? Do you...  ...? We are looking for a Site Reliability Engineer to join the...  ...service systems that automate managing, deploying and operating services... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    Toronto, ON
    7 hours ago
  • $144k - $200k per year

    **The Team** Platform Engineering is the department within SRE that is...  ...alerting systems. The Fabric team manages the infrastructure that...  ...developing and maintaining the reliable and globally connected multi-...  ...* We are seeking a talented Site Reliability Engineer (SRE) with... 
    Full time
    Work at office
    Remote work
    Worldwide
    Flexible hours

    Mongodb

    Toronto, ON
    7 hours ago
  • $150k - $250k per year

     ...About The Role We're looking for a Senior Site Reliability Engineer to help us run one of the most exciting GPU clusters around—our Toronto datacenter packed with NVIDIA H100 and A100 GPUs, over 20PB of Ceph storage, terabit networking, and hundreds of servers. You'll... 
    Senior
    Full time

    Boson Ai

    Toronto, ON
    7 hours ago
  • $141k - $191k per year

     ...Do you have experience in Service Management, working with cloud providers, software...  ...SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational...  ...the Role: In this opportunity as Site Reliability Engineering Manager , you will be... 
    Work at office
    Local area
    Flexible hours
    2 days per week
    3 days per week

    Thomson Reuters

    Toronto, ON
    more than 2 months ago
  •  ...capacity plans, and ensure the reliability, durability, and operational...  ...Atlas. You’ll join a small, senior team of SREs as founding members...  ...are a small team of software engineers with a strong bias towards...  ...and durability requirements Managing and scaling infrastructure... 
    Senior
    Long term contract
    Full time
    Work at office
    Local area
    Immediate start
    Remote work
    Worldwide
    Shift work

    Mongodb

    Toronto, ON
    7 hours ago
  •  ...Platform Engineer – DevOps, Site Reliability Engineering (SRE) & Dynatrace Required Skills • Strong experience as a Platform Engineer with...  ...monitoring • Observability • Automation • Incident management • Strong understanding of SRE best practices and... 
    Permanent employment

    Astra North Infoteck Inc.

    Toronto, ON
    24 days ago
  •  ...part of our team! Apply now for the following position: Senior Project Manager Product Reliability. Overview: As the Senior Project Manager Product...  ...Reliability strategy, and Reporting to the Director Sustaining Engineering, this role is integral to achieving our must-win... 
    Senior
    Full time
    Flexible hours

    Sonova AG

    Toronto, ON
    3 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Manager, Site Reliability Engineering. Be the first to apply!