Manager, Network Reliability and Resiliency
ServiceNow
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.
Join us to put AI to work for people.
Description du poste
Due to Government of Canada regulatory requirements, this position requires the successful completion of a Government of Canada Reliability Status screening as a condition of employment. The screening process requires 5 years of verifiable background history. This includes identity verification, education verification, a criminal record check, and a credit check. Candidates must be eligible to obtain and maintain Reliability Status, which generally requires Canadian citizenship or Canadian permanent resident status. Employment is contingent upon successful completion and maintenance of the required screening.
What you get to do in this role:
We are seeking a Manager, Network Reliability and Resiliency to lead a team responsible for the reliability and day-to-day operation of production network services supporting ServiceNow's cloud platform. This is a technical people-manager role. You will develop engineers and manage team priorities while staying actively engaged in complex troubleshooting, high-severity incidents, customer escalations, operational readiness, and reliability improvement.
You will apply SRE principles to network operations by using service indicators and objectives, error-budget thinking, observability, post-incident learning, and automation to improve availability, reduce operational toil, and make execution safer and more consistent. While this is not an individual contributor role, you must have the technical depth and judgment to guide investigations, challenge assumptions, make risk-based decisions, and help the team reach durable solutions.
Lead and develop the team
- Manage, coach, and develop network reliability engineers through clear goals, regular feedback, performance reviews, and career development.
- Set priorities and ownership for operational work, reliability initiatives, technical debt, and project commitments.
- Build sustainable on-call and escalation practices and promote calm, accountable execution during high-pressure events.
- Hire and onboard new team members and ensure they gain the technical context, operating practices, and support needed to succeed.
Provide technical and incident leadership
- Actively engage in complex production troubleshooting and customer-impacting escalations by reviewing evidence, guiding technical hypotheses, identifying risk, and coordinating the right subject-matter experts.
- Lead or support major incident response, including mitigation decisions, stakeholder communication, escalation management, and restoration of service.
- Ensure post-incident reviews identify contributing factors and result in clear, prioritized, and completed preventive actions.
- Review high-risk changes and operational plans for technical soundness, rollback readiness, monitoring coverage, and customer impact.
Improve reliability through SRE practices
- Partner with engineering and service owners to define and use meaningful SLIs and SLOs for network services.
- Use error budgets, incident trends, capacity signals, and operational data to balance service reliability, delivery pace, and risk.
- Improve observability, alert quality, dashboards, runbooks, and operational readiness so the team can detect and resolve issues efficiently.
- Track practical reliability outcomes such as availability, recurring incidents, change success, alert effectiveness, and time to detect and recover.
Embed automation in daily operations
- Create a strong automation mindset across the team and identify repetitive, error-prone, or slow operational activities that should be eliminated or automated.
- Prioritize automation that improves change safety, validation, triage, remediation, reporting, and operational consistency.
- Work with engineering and automation partners to move useful tools and workflows into production with clear ownership, documentation, monitoring, and support models.
- Measure whether automation reduces toil and operational risk rather than treating automation delivery alone as the outcome.
Partner across the organization
- Collaborate with network engineering, SRE, security, platform, data center, customer support, and other partner teams to resolve issues and improve service reliability.
- Represent the team's technical assessment, customer impact, risks, dependencies, and recovery plan clearly to technical and business stakeholders.
- Ensure new technologies, services, and automations meet operational acceptance criteria before the team assumes production ownership.
- Improve incident, change, problem-management, and escalation processes based on operational evidence and team feedback.
Qualifications
To be successful in this role you have:
- Five or more years of relevant experience in network engineering, network reliability, cloud infrastructure, SRE, or large-scale production operations.
- Experience managing or formally leading engineers, including prioritization, coaching, performance feedback, and delivery accountability.
- Sufficient hands-on technical background to guide production troubleshooting across Linux-based systems and network services. You can interpret logs, metrics, alerts, and packet-level evidence and make sound operational decisions.
- Working knowledge of networking concepts and technologies such as TCP/IP, routing, DNS, load balancing or ADCs, firewalls, cloud networking, and network observability. Deep expertise in every area is not required.
- Experience leading or coordinating significant incidents and customer-impacting escalations in an always-on service environment.
- Working knowledge of SRE practices, including SLIs, SLOs, error budgets, monitoring and alerting, incident management, and post-incident improvement.
- An automation mindset and experience using scripting, workflow automation, or engineering partnerships to reduce manual operational work and improve consistency.
- Experience working with geographically distributed teams and cross-functional partners in software, platform, infrastructure, or cloud services.
- Strong written and verbal communication skills, sound judgment under pressure, and consistent attention to detail.
- Experience using or evaluating AI-assisted tools to improve analysis, decision-making, automation, or team workflows, with appropriate attention to accuracy, security, and operational risk.
Preferred qualifications
- Experience operating networking for a global SaaS, large enterprise, cloud provider, or similarly complex production environment.
- Familiarity with BGP or OSPF, data center fabrics, load balancers or ADCs, DDoS protection, VPNs, firewalls, or public-cloud networking.
- Experience improving observability, change safety, capacity management, or operational readiness for production services.
- Experience with IT service management practices, including incident, change, and problem management.
- Relevant certifications such as CCNA, CCNP, Azure/AWS/GCP related
Work Personas
We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here . To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal Opportunity Employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact Voir email sur jobs.smartrecruiters.com for assistance.
Export Control Regulations
For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.
From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.
- ...members. Together we have over $142 billion in combined assets under management and administration, with a clear mandate to drive change in... ...for the Designing, planning and executing the bank’s Cyber Resilience Testing and offensive security program, commonly referred to as...SuggéréTaux horaireEmploi permanentTemps pleinTravail au bureau
$177.1k - $221.4k par année
...As Marqeta’s Senior Manager, Network Operations, you will lead our Network Operations team through... ...evolution while ensuring the reliability our customers depend on. We work Flexible... ...beyond), setting the foundation for a resilient cloud presence Protect 99.999% uptime...RéseauTemps pleinTravail au bureauTravail à distanceHoraires flexibles$146k - $201.3k par année
...let's talk. The Engineering Opportunity Reporting to the Director of Quality & Performance, this role as Engineering Manager of Performance & Resilience will drive performance and resiliency improvements for the Auth0 product at Okta. Here you'll be working with some of...SuggéréTemps pleinZone localeLe monde entier- ...AXON Networks delivers a robust AI-driven, analytics-based orchestration... ...give ISPs the ability to manage and troubleshoot their... ...Network Operations to deliver reliable, predictable and continuously... ...model, runbooks, SLOs, capacity, resilience, security, recovery, ownership...RéseauTravail à distanceTemps pleinСontrat
- ...organizations create, deliver, and manage training all in one place.... ...As the Manager of Site Reliability Engineering (SRE), you will lead... ...safeguarding the operational health and resilience of the Docebo platform. This... ...leadership to integrate reliable SRE practices into early...SuggéréContrat Longue DuréeTemps pleinPour les contractantsTravail au bureauLe monde entier3 jours par semaine
$120k - $170k par année
...secure, and highly available network architectures in AWS. This... ...responsible for building and managing complex multi-account, multi-... ...; ~ Implement and manage secure networking solutions using... ...improvements that increase reliability and operational efficiency;...RéseauTemps pleinTravail au bureauZone localeHoraires flexibles$90k - $105k par année
...just be in the right place! NuORDER by Lightspeed is seeking an experienced Integrated Marketing Manager to drive adoption, engagement, and growth across our wholesale network. This role owns full-funnel marketing strategies across acquisition, onboarding, and ongoing...RéseauTemps pleinTravail temporaireStageTravail à distanceHoraires flexibles$110k - $120k par année
...Canada’s largest non-bank branch network—and a leader in financial... ...your impact The Job: Site Reliability Engineer The Site Reliability... ..., performance, and resilience of the organization's digital... ...and incident trends to senior management. What You Bring : Technical...RéseauTemps pleinTravail temporaireStageTravail au bureauTravail à distance- .... We are looking for an experienced fund manager to bring the capital, the LP relationships... ...needed to make it real. As AI Resilience Fund Manager, you will build and run this... ...institutional investing An established network of potential LPs and co-investors — you have...RéseauContrat Longue DuréeTemps pleinPour les contractantsTravail à distanceTravail posté
- ...healthcare. We're currently seeking a Network and System Administrator to join our IT... ...and enterprise hardware systems management software and architecture Provide general... ...workstations, and related infrastructure to ensure reliable clinical operations and system...RéseauTemps pleinTravail occasionnelStageTravail au bureauZone localeTravail à distanceLe monde entierHoraires flexibles
$153.82k - $277k par année
...’t wait to meet you. Site Reliability Engineers (SREs) at Braze are... ...ingress controller layers that manage massive, real-time API... ...Partner with Product Teams on Resilient Architectures Translate Product... ...administration, cluster networking, container orchestration, cluster...RéseauEmploi permanentTemps pleinStageTravail au bureauZone localeTravail à distanceHoraires flexiblesPoste rotatif- ...organizations create, deliver, and manage training all in one place.... ...Overview As a Senior Site Reliability Engineer, you'll take a hands... ...continuously to improve its resilience and ability to scale. Run... ...major cloud provider and its networking layer. ~ Comfort writing...RéseauTemps pleinPour les contractantsTravail au bureauLe monde entier3 jours par semaine
$140k - $155k par année
...focused on improving production resilience, strengthening security,... ...engineering teams to ship secure, reliable, and scalable software with... ..., and drive incident management and post-incident learning practices... ...Lambda, CloudFront, S3, and cloud networking/security best practices. ~...RéseauTravail à distanceEmploi permanentTemps pleinHoraires flexibles$96.9k - $136.8k par année
...this role. Job Description: The Manager Network Service Delivery plays a critical leadership role in ensuring the reliability security and continuous evolution of the network... ...drives service performance operational resilience and lifecycle modernization across LAN WAN...RéseauTemps pleinTravail à domicileRelocation- ...the enterprise. The Site Reliability Engineer III (SRE III) plays... ...availability, scalability, and resilience in mind. Monitor,... ...-as-Code (IaC) principles to manage large-scale distributed systems... ...infrastructure, databases, and networks. Leadership & Mentorship...RéseauTemps pleinTravail manuelZone localeHoraires flexibles
$90 - $110 par heure
...fun? We are looking for an energetic and client-orientated Network Security Consultant for our Toronto branch. You will lead... ...Strong skills in high availability network design, failover testing, resilient routing, and multi-region/multi-cloud connectivity Automation...RéseauTemps pleinСontratHoraires flexibles$80 - $110 par heure
...AXON Networks delivers a robust AI-driven, analytics-based orchestration... ...give ISPs the ability to manage and troubleshoot their... ...Spain and Vietnam. The Site Reliability Engineer will improve the availability... ...observable, supportable and resilient at fleet scale. You...RéseauTravail à distanceTemps pleinСontrat$154k - $200k par année
..., and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical... ...capacity planning and performance management frameworks, proactively identifying scaling... ...and championing practices that ensure resilient service delivery Demonstrate a high...Contrat Longue DuréeTemps plein$144k - $200k par année
...alerting systems. The Fabric team manages the infrastructure that... ...Their responsibilities encompass network architecture, service mesh,... ...developing and maintaining the reliable and globally connected multi-... ...automation to ensure our systems are resilient, scalable, and reliable. **...RéseauTemps pleinTravail au bureauTravail à distanceLe monde entierHoraires flexibles- ...of our large clients is looking for contract Sr. Network Engineer, details are as below: Hiring Manager Name: Jason Nickel, DG2458 Hiring Manager Location... ...to work extra time or on-call: No Enhanced Reliability Status: Experience & Skills: Minimum...RéseauRemplacementEmploi permanentTemps pleinСontratTemps partielPour les contractantsTravail au bureauTravail posté3 jours par semaine
$100k - $125k par année
...Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site... ...employee expenses, corporate cards, supplier management, tax compliance, and treasury. Tipalti... ...lower levels of software frameworks and networking. Distributed monitoring experience with...RéseauTemps pleinTravail au bureauHoraires flexibles- ...high-performance, scalable and reliable machine learning systems? Do... ...systems that automate managing, deploying and operating services... ...environment observability and resilience. Enable all developers to troubleshoot... ...in compute/storage/network resource and cost management...RéseauTemps pleinTravail au bureauTravail à distanceHoraires flexibles
$75k - $80k par année
...Requisition Number: 105730 Partner Manager, Networking Salary: T he base salary range for this position is typically $75,000 to $80,000 CAD with additional bonus and benefits available. However, compensation decisions are dependent on the facts and circumstances of...RéseauTravail au bureauZone localeRecrutement immédiat$140k - $182k par année
...Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across... ...Engineering, building and maintaining scalable, resilient services. ~ Building the tooling and automation to manage those services, as well as investigating system...Temps plein- ...accomplishment. Being a Senior Site Reliability Engineer at iManage Means… You are an... ...discussions, reduce toil, and deliver scalable, resilient platforms that support our customers and... ...be a key voice in observability, change management, and service scalability, providing...Temps pleinTravail au bureauZone localeTravail à distanceLe monde entierLundi au vendrediHoraires flexibles
- ...software models, compilers, platforms, networking, and semiconductors. Our diverse team of... ...seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability... ...of AI hardware by helping build highly reliable systems that power tomorrow's largest AI...RéseauEmploi permanentTemps pleinStageDeuxième emploi
$80.46k - $105.6k par année
...We are looking for an energetic and client-orientated Senior Network Consultant for our Toronto branch. You will be a member of our... ...monitoring solutions (Wireshark, SolarWinds) Cisco Voice/Call Manager What Makes You Extra Awesome: Cisco Certified Network Associate...RéseauTemps pleinHoraires flexibles$155k - $269k par année
...Job Responsibility: Senior Manager, Network Architecture JR-152422 Hybrid Remote Toronto Warsaw Show More Information Technology Full time Who are we? Equinix is the world's digital infrastructure company®, operating over 260 data centers across the globe. Digital...RéseauTemps pleinСontratZone localeTravail à distance$100k par année
...to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed... ...seniorities. Tenstorrent is looking for an experienced Reliability Engineer to drive and execute reliability testing for our...RéseauEmploi permanentTemps pleinStage$103.37k par année
...division, is responsible for campus core network, campus wireless, wide area network... ...services related to departmental network management, network, server and storage infrastructure... ...planning, design, engineering, operational reliability, and cost efficiency. You will be...RéseauTemps plein
Voulez-vous recevoir plus d'offres d'emploi ?
S'abonner et recevoir des offres d'emploi similaires à Manager, Network Reliability and Resiliency. Soyez parmi les premiers à postuler !
- printing manager Toronto, ON
- loan manager Toronto, ON
- home building and renovation manager Toronto, ON
- volunteer manager Toronto, ON
- concession manager Toronto, ON
- patient experience manager Toronto, ON
- landscape manager Toronto, ON
- apartment manager Toronto, ON
- custodial manager Toronto, ON
- greenhouse manager Toronto, ON

