Inscrivez-vous pour accéder à toutes les fonctionnalités de notre service.
  • Recherche d'offres d'emploi
  • Favoris
  • Créer un CV
    Nouveau
  • Salaires
  • Souscriptions

Manager, Network Reliability and Resiliency

Intérimaire

ServiceNow

Description de l'entreprise

It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.

Join us to put AI to work for people.

 

Description du poste

Due to Government of Canada regulatory requirements, this position requires the successful completion of a Government of Canada Reliability Status screening as a condition of employment. The screening process requires 5 years of verifiable background history. This includes identity verification, education verification, a criminal record check, and a credit check. Candidates must be eligible to obtain and maintain Reliability Status, which generally requires Canadian citizenship or Canadian permanent resident status. Employment is contingent upon successful completion and maintenance of the required screening.

What you get to do in this role:

We are seeking a Manager, Network Reliability and Resiliency to lead a team responsible for the reliability and day-to-day operation of production network services supporting ServiceNow's cloud platform. This is a technical people-manager role. You will develop engineers and manage team priorities while staying actively engaged in complex troubleshooting, high-severity incidents, customer escalations, operational readiness, and reliability improvement.

You will apply SRE principles to network operations by using service indicators and objectives, error-budget thinking, observability, post-incident learning, and automation to improve availability, reduce operational toil, and make execution safer and more consistent. While this is not an individual contributor role, you must have the technical depth and judgment to guide investigations, challenge assumptions, make risk-based decisions, and help the team reach durable solutions.

Lead and develop the team

  • Manage, coach, and develop network reliability engineers through clear goals, regular feedback, performance reviews, and career development.
  • Set priorities and ownership for operational work, reliability initiatives, technical debt, and project commitments.
  • Build sustainable on-call and escalation practices and promote calm, accountable execution during high-pressure events.
  • Hire and onboard new team members and ensure they gain the technical context, operating practices, and support needed to succeed.

Provide technical and incident leadership

  • Actively engage in complex production troubleshooting and customer-impacting escalations by reviewing evidence, guiding technical hypotheses, identifying risk, and coordinating the right subject-matter experts.
  • Lead or support major incident response, including mitigation decisions, stakeholder communication, escalation management, and restoration of service.
  • Ensure post-incident reviews identify contributing factors and result in clear, prioritized, and completed preventive actions.
  • Review high-risk changes and operational plans for technical soundness, rollback readiness, monitoring coverage, and customer impact.

Improve reliability through SRE practices

  • Partner with engineering and service owners to define and use meaningful SLIs and SLOs for network services.
  • Use error budgets, incident trends, capacity signals, and operational data to balance service reliability, delivery pace, and risk.
  • Improve observability, alert quality, dashboards, runbooks, and operational readiness so the team can detect and resolve issues efficiently.
  • Track practical reliability outcomes such as availability, recurring incidents, change success, alert effectiveness, and time to detect and recover.

Embed automation in daily operations

  • Create a strong automation mindset across the team and identify repetitive, error-prone, or slow operational activities that should be eliminated or automated.
  • Prioritize automation that improves change safety, validation, triage, remediation, reporting, and operational consistency.
  • Work with engineering and automation partners to move useful tools and workflows into production with clear ownership, documentation, monitoring, and support models.
  • Measure whether automation reduces toil and operational risk rather than treating automation delivery alone as the outcome.

Partner across the organization

  • Collaborate with network engineering, SRE, security, platform, data center, customer support, and other partner teams to resolve issues and improve service reliability.
  • Represent the team's technical assessment, customer impact, risks, dependencies, and recovery plan clearly to technical and business stakeholders.
  • Ensure new technologies, services, and automations meet operational acceptance criteria before the team assumes production ownership.
  • Improve incident, change, problem-management, and escalation processes based on operational evidence and team feedback.

Qualifications



To be successful in this role you have:

  • Five or more years of relevant experience in network engineering, network reliability, cloud infrastructure, SRE, or large-scale production operations.
  • Experience managing or formally leading engineers, including prioritization, coaching, performance feedback, and delivery accountability.
  • Sufficient hands-on technical background to guide production troubleshooting across Linux-based systems and network services. You can interpret logs, metrics, alerts, and packet-level evidence and make sound operational decisions.
  • Working knowledge of networking concepts and technologies such as TCP/IP, routing, DNS, load balancing or ADCs, firewalls, cloud networking, and network observability. Deep expertise in every area is not required.
  • Experience leading or coordinating significant incidents and customer-impacting escalations in an always-on service environment.
  • Working knowledge of SRE practices, including SLIs, SLOs, error budgets, monitoring and alerting, incident management, and post-incident improvement.
  • An automation mindset and experience using scripting, workflow automation, or engineering partnerships to reduce manual operational work and improve consistency.
  • Experience working with geographically distributed teams and cross-functional partners in software, platform, infrastructure, or cloud services.
  • Strong written and verbal communication skills, sound judgment under pressure, and consistent attention to detail.
  • Experience using or evaluating AI-assisted tools to improve analysis, decision-making, automation, or team workflows, with appropriate attention to accuracy, security, and operational risk.

Preferred

qualifications

  • Experience operating networking for a global SaaS, large enterprise, cloud provider, or similarly complex production environment.
  • Familiarity with BGP or OSPF, data center fabrics, load balancers or ADCs, DDoS protection, VPNs, firewalls, or public-cloud networking.
  • Experience improving observability, change safety, capacity management, or operational readiness for production services.
  • Experience with IT service management practices, including incident, change, and problem management.
  • Relevant certifications such as CCNA, CCNP, Azure/AWS/GCP related

Informations complémentaires

 

Work Personas

We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here . To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.

Equal Opportunity Employer

ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.  

Accommodations

We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact Voir email sur jobs.smartrecruiters.com for assistance. 

Export Control Regulations

For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities. 

From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.

L'offre d'emploi a été publiée il y a 2 jours
Des emplois similaires qui pourraient vous intéresserBasé sur l'offre Manager, Network Reliability and Resiliency à Toronto, ON
  •  ...members. Together we have over $142 billion in combined assets under management and administration, with a clear mandate to drive change in...  ...for the Designing, planning and executing the bank’s Cyber Resilience Testing and offensive security program, commonly referred to as... 
    Suggéré
    Taux horaire
    Emploi permanent
    Temps plein
    Travail au bureau

    Eq Bank

    Toronto, ON
    il y a 20 heures
  • $177.1k - $221.4k par année

     ...As Marqeta’s Senior Manager, Network Operations, you will lead our Network Operations team through...  ...evolution while ensuring the reliability our customers depend on. We work Flexible...  ...beyond), setting the foundation for a resilient cloud presence Protect 99.999% uptime... 
    Réseau
    Temps plein
    Travail au bureau
    Travail à distance
    Horaires flexibles

    Marqeta

    Toronto, ON
    il y a 20 heures
  • $146k - $201.3k par année

     ...let's talk. The Engineering Opportunity Reporting to the Director of Quality & Performance, this role as Engineering Manager of Performance & Resilience will drive performance and resiliency improvements for the Auth0 product at Okta. Here you'll be working with some of... 
    Suggéré
    Temps plein
    Zone locale
    Le monde entier

    Okta

    Toronto, ON
    il y a 20 heures
  •  ...AXON Networks delivers a robust AI-driven, analytics-based orchestration...  ...give ISPs the ability to manage and troubleshoot their...  ...Network Operations to deliver reliable, predictable and continuously...  ...model, runbooks, SLOs, capacity, resilience, security, recovery, ownership... 
    Réseau
    Travail à distance
    Temps plein
    Сontrat

    Axon-networks

    Toronto, ON
    il y a 20 heures
  •  ...organizations create, deliver, and manage training all in one place....  ...As the Manager of Site Reliability Engineering (SRE), you will lead...  ...safeguarding the operational health and resilience of the Docebo platform. This...  ...leadership to integrate reliable SRE practices into early... 
    Suggéré
    Contrat Longue Durée
    Temps plein
    Pour les contractants
    Travail au bureau
    Le monde entier
    3 jours par semaine

    Docebo

    Toronto, ON
    il y a 20 heures
  • $120k - $170k par année

     ...secure, and highly available network architectures in AWS.   This...  ...responsible for building and managing complex multi-account, multi-...  ...; ~ Implement and manage secure networking solutions using...  ...improvements that increase reliability and operational efficiency;... 
    Réseau
    Temps plein
    Travail au bureau
    Zone locale
    Horaires flexibles

    Magnet Forensics

    Toronto, ON
    il y a 20 heures
  • $90k - $105k par année

     ...just be in the right place! NuORDER by Lightspeed is seeking an experienced Integrated Marketing Manager to drive adoption, engagement, and growth across our wholesale network. This role owns full-funnel marketing strategies across acquisition, onboarding, and ongoing... 
    Réseau
    Temps plein
    Travail temporaire
    Stage
    Travail à distance
    Horaires flexibles

    Lightspeed Commerce

    Toronto, ON
    il y a 20 heures
  • $110k - $120k par année

     ...Canada’s largest non-bank branch network—and a leader in financial...  ...your impact The Job: Site Reliability Engineer The Site Reliability...  ..., performance, and resilience of the organization's digital...  ...and incident trends to senior management. What You Bring : Technical... 
    Réseau
    Temps plein
    Travail temporaire
    Stage
    Travail au bureau
    Travail à distance

    Momentum Financial Services Group

    Toronto, ON
    il y a 20 heures
  •  .... We are looking for an experienced fund manager to bring the capital, the LP relationships...  ...needed to make it real. As AI Resilience Fund Manager, you will build and run this...  ...institutional investing An established network of potential LPs and co-investors — you have... 
    Réseau
    Contrat Longue Durée
    Temps plein
    Pour les contractants
    Travail à distance
    Travail posté

    Human Agency

    Toronto, ON
    il y a 18 heures
  •  ...healthcare. We're currently seeking a  Network and System Administrator to join our IT...  ...and enterprise hardware systems management software and architecture  Provide general...  ...workstations, and related infrastructure to ensure reliable clinical operations and system... 
    Réseau
    Temps plein
    Travail occasionnel
    Stage
    Travail au bureau
    Zone locale
    Travail à distance
    Le monde entier
    Horaires flexibles

    Ramsoft Inc

    Toronto, ON
    il y a 20 heures
  • $153.82k - $277k par année

     ...’t wait to meet you. Site Reliability Engineers (SREs) at Braze are...  ...ingress controller layers that manage massive, real-time API...  ...Partner with Product Teams on Resilient Architectures Translate Product...  ...administration, cluster networking, container orchestration, cluster... 
    Réseau
    Emploi permanent
    Temps plein
    Stage
    Travail au bureau
    Zone locale
    Travail à distance
    Horaires flexibles
    Poste rotatif

    Braze

    Toronto, ON
    il y a 20 heures
  •  ...organizations create, deliver, and manage training all in one place....  ...Overview As a Senior Site Reliability Engineer, you'll take a hands...  ...continuously to improve its resilience and ability to scale. Run...  ...major cloud provider and its networking layer. ~ Comfort writing... 
    Réseau
    Temps plein
    Pour les contractants
    Travail au bureau
    Le monde entier
    3 jours par semaine

    Docebo

    Toronto, ON
    il y a 20 heures
  • $140k - $155k par année

     ...focused on improving production resilience, strengthening security,...  ...engineering teams to ship secure, reliable, and scalable software with...  ..., and drive incident management and post-incident learning practices...  ...Lambda, CloudFront, S3, and cloud networking/security best practices. ~... 
    Réseau
    Travail à distance
    Emploi permanent
    Temps plein
    Horaires flexibles

    Caseware

    Toronto, ON
    il y a 20 heures
  • $96.9k - $136.8k par année

     ...this role. Job Description: The Manager Network Service Delivery plays a critical leadership role in ensuring the reliability security and continuous evolution of the network...  ...drives service performance operational resilience and lifecycle modernization across LAN WAN... 
    Réseau
    Temps plein
    Travail à domicile
    Relocation

    TD Bank

    Toronto, ON
    il y a 15 jours
  •  ...the enterprise. The Site Reliability Engineer III (SRE III) plays...  ...availability, scalability, and resilience in mind. Monitor,...  ...-as-Code (IaC) principles to manage large-scale distributed systems...  ...infrastructure, databases, and networks. Leadership & Mentorship... 
    Réseau
    Temps plein
    Travail manuel
    Zone locale
    Horaires flexibles

    Emburse

    Toronto, ON
    il y a 20 heures
  • $90 - $110 par heure

     ...fun?   We are looking for an energetic and client-orientated  Network Security Consultant  for our Toronto branch. You will  lead...  ...Strong skills in high availability network design, failover testing, resilient routing, and multi-region/multi-cloud connectivity Automation... 
    Réseau
    Temps plein
    Сontrat
    Horaires flexibles

    Long View Systems

    Toronto, ON
    il y a 20 heures
  • $80 - $110 par heure

     ...AXON Networks delivers a robust AI-driven, analytics-based orchestration...  ...give ISPs the ability to manage and troubleshoot their...  ...Spain and Vietnam. The Site Reliability Engineer  will improve the availability...  ...observable, supportable and resilient at fleet scale.   You... 
    Réseau
    Travail à distance
    Temps plein
    Сontrat

    Axon-networks

    Toronto, ON
    il y a 20 heures
  • $154k - $200k par année

     ..., and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical...  ...capacity planning and performance management frameworks, proactively identifying scaling...  ...and championing practices that ensure resilient service delivery Demonstrate a high... 
    Contrat Longue Durée
    Temps plein

    Movable Ink

    Toronto, ON
    il y a 20 heures
  • $144k - $200k par année

     ...alerting systems. The Fabric team manages the infrastructure that...  ...Their responsibilities encompass network architecture, service mesh,...  ...developing and maintaining the reliable and globally connected multi-...  ...automation to ensure our systems are resilient, scalable, and reliable. **... 
    Réseau
    Temps plein
    Travail au bureau
    Travail à distance
    Le monde entier
    Horaires flexibles

    Mongodb

    Toronto, ON
    il y a 20 heures
  •  ...of our large clients is looking for contract Sr. Network Engineer, details are as below: Hiring Manager Name: Jason Nickel, DG2458 Hiring Manager Location...  ...to work extra time or on-call: No Enhanced Reliability Status:     Experience & Skills: Minimum... 
    Réseau
    Remplacement
    Emploi permanent
    Temps plein
    Сontrat
    Temps partiel
    Pour les contractants
    Travail au bureau
    Travail posté
    3 jours par semaine

    Strategize It

    Toronto, ON
    il y a 20 heures
  • $100k - $125k par année

     ...Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site...  ...employee expenses, corporate cards, supplier management, tax compliance, and treasury. Tipalti...  ...lower levels of software frameworks and networking. Distributed monitoring experience with... 
    Réseau
    Temps plein
    Travail au bureau
    Horaires flexibles

    Tipalti

    Toronto, ON
    il y a 20 heures
  •  ...high-performance, scalable and reliable machine learning systems? Do...  ...systems that automate managing, deploying and operating services...  ...environment observability and resilience. Enable all developers to troubleshoot...  ...in compute/storage/network resource and cost management... 
    Réseau
    Temps plein
    Travail au bureau
    Travail à distance
    Horaires flexibles

    Cohere

    Toronto, ON
    il y a 20 heures
  • $75k - $80k par année

     ...Requisition Number: 105730 Partner Manager, Networking Salary: T he base salary range for this position is typically $75,000 to $80,000 CAD with additional bonus and benefits available. However, compensation decisions are dependent on the facts and circumstances of... 
    Réseau
    Travail au bureau
    Zone locale
    Recrutement immédiat

    Insight Enterprises, Inc.

    Toronto, ON
    il y a 8 jours
  • $140k - $182k par année

     ...Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across...  ...Engineering, building and maintaining scalable, resilient services. ~ Building the tooling and automation to manage those services, as well as investigating system... 
    Temps plein

    Movable Ink

    Toronto, ON
    il y a 20 heures
  •  ...accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an...  ...discussions, reduce toil, and deliver scalable, resilient platforms that support our customers and...  ...be a key voice in observability, change management, and service scalability, providing... 
    Temps plein
    Travail au bureau
    Zone locale
    Travail à distance
    Le monde entier
    Lundi au vendredi
    Horaires flexibles

    Imanage

    Toronto, ON
    il y a 20 heures
  •  ...software models, compilers, platforms, networking, and semiconductors. Our diverse team of...  ...seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability...  ...of AI hardware by helping build highly reliable systems that power tomorrow's largest AI... 
    Réseau
    Emploi permanent
    Temps plein
    Stage
    Deuxième emploi

    Tenstorrent

    Toronto, ON
    il y a 20 heures
  • $80.46k - $105.6k par année

     ...We are looking for an energetic and client-orientated Senior Network Consultant for our Toronto branch. You will be a member of our...  ...monitoring solutions (Wireshark, SolarWinds) Cisco Voice/Call Manager What Makes You Extra Awesome: Cisco Certified Network Associate... 
    Réseau
    Temps plein
    Horaires flexibles

    Long View Systems

    Toronto, ON
    il y a 20 heures
  • $155k - $269k par année

     ...Job Responsibility: Senior Manager, Network Architecture JR-152422 Hybrid Remote Toronto Warsaw Show More Information Technology Full time Who are we? Equinix is the world's digital infrastructure company®, operating over 260 data centers across the globe. Digital... 
    Réseau
    Temps plein
    Сontrat
    Zone locale
    Travail à distance

    Equinix

    Toronto, ON
    il y a 6 jours
  • $100k par année

     ...to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed...  ...seniorities. Tenstorrent is looking for an experienced Reliability Engineer to drive and execute reliability testing for our... 
    Réseau
    Emploi permanent
    Temps plein
    Stage

    Tenstorrent

    Toronto, ON
    il y a 20 heures
  • $103.37k par année

     ...division, is responsible for campus core network, campus wireless, wide area network...  ...services related to departmental network management, network, server and storage infrastructure...  ...planning, design, engineering, operational reliability, and cost efficiency. You will be... 
    Réseau
    Temps plein

    University of Toronto

    Toronto, ON
    il y a 2 jours

Voulez-vous recevoir plus d'offres d'emploi ?

S'abonner et recevoir des offres d'emploi similaires à Manager, Network Reliability and Resiliency. Soyez parmi les premiers à postuler !