Staff Software Engineer - SRE & AIOps
$125.7k - $220k per yearServiceNow
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.
Join us to put AI to work for people.
Job Description
About the Role
ServiceNow is seeking a Staff Software Engineer – SRE & AIOps to drive infrastructure automation, operational resilience, and toil elimination across our hybrid cloud and data center operations. Embedded within the Site Reliability & Database Engineering organization, you will design and implement automation-first systems that reduce manual intervention, accelerate incident remediation, and enable our global engineering teams to operate reliably at scale.
This role combines strong hands-on technical expertise in Kubernetes, cloud platforms, and DevOps practices with technical leadership influence across infrastructure teams. You will architect SRE tooling, develop auto-remediation capabilities, and establish patterns that allow ServiceNow's cloud platform to maintain high reliability while minimizing operational toil across follow-the-sun global teams.What you get to do in this role:
- Design, deploy, and operate enterprise-scale Kubernetes clusters across hybrid and multi-cloud environments, establishing governance, scaling policies, and operational practices that support high-velocity application deployments at 99.99%+ availability targets.
- Architect and implement closed-loop auto-remediation systems that detect, classify, and resolve transient infrastructure failures without human intervention, leveraging agentic AI and machine learning frameworks to predict failures, trigger preventive actions, and continuously reduce MTTR and on-call burden.
- Design and evolve the SRE tooling stack, including monitoring platforms, incident management systems, log aggregation, and observability integrations, that support global follow-the-sun on-call operations and enable data-driven incident response.
- Establish SLO frameworks, error budgets, and alerting policies that balance rapid incident response with alert fatigue management, while developing automated runbooks and playbooks that empower on-call engineers to resolve issues autonomously.
- Design and maintain Infrastructure-as-Code frameworks and GitOps pipelines that enable reproducible, auditable infrastructure deployments across hybrid and multi-cloud environments with consistent security and compliance guardrails.
- Architect hybrid cloud and data center operations, spanning on-premises infrastructure, public cloud environments, and edge computing, including workload migration strategies, disaster recovery patterns, and cost optimization practices across multi-region deployments.
- Drive adoption of containerization, microservices, and DevOps patterns across engineering teams, establishing CI/CD best practices, service mesh architectures, and network security controls that enable rapid, safe release cycles.
- Design on-call rotation schedules, escalation policies, and incident command systems that span across different time zones, ensuring 24/7 incident response while driving post-incident review processes that capture learning and drive systemic improvements.
- Mentor and guide junior SRE engineers and infrastructure teams on reliability patterns, incident investigation techniques, automation best practices, and agentic AI applications for infrastructure operations.
- Champion a culture of blameless incident analysis, data-driven decision-making, continuous improvement, and experimentation across engineering teams, establishing knowledge-sharing practices and technical documentation standards.
- Reduce operational toil through systematic automation of repetitive tasks, from infrastructure provisioning to incident response to cost optimization, directly improving team capacity and job satisfaction across globally distributed operations.
Qualifications
To Be Successful in This Role You Have
- Kubernetes Mastery: Strong hands-on expertise operating production Kubernetes clusters at scale, including cluster design, node management, pod orchestration, resource quotas, network policies, security controls, and troubleshooting complex runtime issues.
- Incident Auto-Remediation Expertise: Proven experience designing and implementing closed-loop automated remediation systems, including anomaly detection, alert correlation, runbook automation, and self-healing mechanisms, that measurably reduce MTTR and on-call burden.
- Cloud Platform Experience: Extensive hands-on experience with AWS (EKS, EC2, RDS, Lambda), Azure (AKS, VMs, CosmosDB), and GCP (GKE, Compute Engine, Cloud SQL), capable of architecting multi-region solutions.
- DevOps & IaC Proficiency: Strong experience with Infrastructure-as-Code tools and GitOps platforms to drive reproducible, auditable infrastructure deployments.
- SRE Tooling Fluency: Strong working knowledge of observability platforms, incident management systems, and log aggregation.
- Distributed Systems Thinking: Solid understanding of distributed system challenges, eventual consistency, cascading failures, network partitions, and proven ability to design systems resilient to these conditions.
- On-Call Operations: Experience operating in follow-the-sun, 24/7 on-call models; ability to design escalation policies, runbooks, and communication patterns that balance responsiveness with operator well-being.
- Data Center & Hybrid Cloud Operations: Hands-on experience managing both on-premises infrastructure and public cloud environments, including hybrid networking, disaster recovery, and workload migration strategies.
- AI/ML Integration: Demonstrated ability to apply machine learning and AI-driven insights to infrastructure operations, including anomaly detection, predictive alerting, and intelligent remediation.
- Technical Leadership: Proven ability to drive technical decisions across teams and mentor engineers on reliability practices through credibility and technical depth.
Qualifications
- Experience in leveraging or critically thinking about how to integrate AI into work processes, decision-making, or problem-solving. This may include using AI-powered tools, automating workflows, analyzing AI-driven insights, or exploring AI's potential impact on the function or industry.
- 8+ years in software engineering or infrastructure operations, with 5+ years in SRE, DevOps, or cloud platform engineering roles managing large-scale distributed systems with a Bachelor's degree; or 6 years and a Master's degree; or a PhD with 3 years experience; or equivalent experience.
- 4+ years hands-on experience designing, deploying, and operating production Kubernetes clusters at scale.
- Proficiency in Infrastructure-as-Code: Terraform, CloudFormation, or equivalent tools used to manage infrastructure at scale.
- Public Cloud Expertise: Demonstrable experience across 2+ of the following: AWS, Azure, GCP, with solid knowledge of services relevant to SRE operations (compute, networking, storage, observability).
- On-Call Operations: Experience operating or designing components of 24/7 follow-the-sun on-call models for distributed teams, including runbook development and incident response.
- Incident Auto-Remediation: Proven ability to design and implement automated remediation systems that measurably reduce manual toil.
- Linux & Systems Programming: Strong foundation in Linux system administration, performance troubleshooting, and scripting (Python, Go, Bash).
- SRE Mindset: Demonstrated commitment to reliability through engineering, favoring durable automation over heroics, and data-driven decision-making.
- Bachelor's degree in computer science, Computer Engineering, or related field (or equivalent professional experience).
Preferred:
- Kubernetes certification (CKA, CKAD, or equivalent).
- Experience with service mesh platforms or advanced networking in Kubernetes environments.
- Background in migrating workloads from on-premises data centers to public cloud environments.
- Experience with cost optimization practices in hybrid cloud environments (reserved instances, spot instances, resource right-sizing).
- Track record of mentoring infrastructure engineering teams.
Why This Role?
This role offers the opportunity to eliminate operational toil at scale and build the automation-first infrastructure practices that define modern cloud operations. You will architect systems that allow ServiceNow's global teams to operate reliably with confidence in automated remediation systems actively preventing and resolving failures. Your work will directly shape how the organization scales reliability as it grows, establishing patterns that benefit teams across the platform. This is a role for an engineer who wants technical depth, team influence, and the satisfaction of watching systems recover from failure automatically.
For positions in this location, we offer a base pay of C $125,700 - $220,000 , plus equity (when applicable), variable/incentive compensation and benefits. Sales positions generally offer a competitive On Target Earnings (OTE) incentive compensation structure. Please note that the base pay shown is a guideline, and individual total compensation will vary based on factors such as
qualifications
, skill level, competencies, and work location. We also offer health plans, including flexible spending accounts, a 401(k) Plan with company match, ESPP, matching donations, a flexible time away plan and family leave programs. Compensation is based on the geographic location in which the role is located and is subject to change based on work location.Additional Information
Work Personas
We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here . To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal Opportunity Employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact View email address on jobs.smartrecruiters.com for assistance.
Export Control Regulations
For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.
From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.
$124k - $179k per year
...Come join us!\n\nJob Summary:\n\nAre you a Software Engineer interested in joining our core R\u0026D... ...our state-of-the-art rendering engine, shading system, and scene representation... ...environment and collaborate with artistic staff.\n\n * Real-time rendering experience.\n...SuggestedFull timeWork at office3 days per week$95k - $145k per year
...artificial intelligence, and software-defined networking to provide... ...prestigious awards, such as Best Engineering Team, Best Company for... ...CloudVision-as-a-Service (CVaaS) global SRE team. SREs at Arista... ...Platform) and GKE (Google Kubernetes Engine) is preferred. Our technical...SuggestedFull time$122k - $163k per year
...providers and application connectivity. We are looking for a Staff Software Engineer who can turn ideas into production capabilities and work... ...Partner closely with Product Management, Architecture, UX, SRE, and peer engineering teams to shape solutions from discovery...SuggestedLocal areaWorldwideFlexible hours- ...with brands like Tinder, Hinge, Match,\u00a0and OkCupid, is looking for a talented and motivated\u00a0\u003cstrong\u003eSenior Software Engineer \u003c/strong\u003eto join our team. As part of our Match Group Central Services team, you\u2019ll design, build, and maintain critical...SuggestedFull timeInternshipWork at office3 days per week
$95k - $145k per year
...advancements in cloud computing, artificial intelligence, and software-defined networking to provide our clients with a competitive edge... ...has earned us several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation, and Work-Life Balance...SuggestedFull time$120k - $195k per year
...computing, artificial intelligence, and software-defined networking to provide our clients... ...several prestigious awards, such as Best Engineering Team, Best Company for Diversity, Compensation... ..., test, and debug packet forwarding engine and a hardware component’s vendor provided...Remote jobFull time$220k - $359.21k per year
...high-quality results more efficiently. We're looking for a Staff Fullstack Engineer to join the Core Objects team as a hands-on technical leader... ...: lead by example and mentor engineers on best practices in software design, coding, and operations Collaborate Cross-...InternshipManual laborWork at officeLocal areaFlexible hours- ...fraud, errors and mistakes. Our specialised Accounts Payable software integrates with leading business ERP systems like SYSPRO,... ...and motivated individual to join our team as a Software Support Engineer. In this role, you will be the critical link between our development...Remote jobLong term contractFull time
$90k - $125k per year
...for our customers so they can spend more time doing what they love. The Role: Push is looking for an experienced Full Stack Software Engineer (back end focus preferred) who loves tracking down tricky bugs, untangling thorny refactors, and uncovering performance...Full timeWork at officeRemote work$132k - $198k per year
...you ready to be a part of a transformative journey from our legacy tech stack to cutting-edge Stack 2.0? We are seeking a Senior Software Engineer to drive this migration and usher in a new era of geolocation services. You’ll be at the forefront, migrating our legacy...Remote jobFull timeLive InWork at officeLocal areaWork from homeShift work$120k - $210k per year
...As Lead Software Engineer, you would lead a team of engineers to write and maintain the tools necessary to support VFX workflows with a focus on production. Our ideal candidate is able to collaborate with non-technical stakeholders to define and document requirements, and...Full timeWorldwide$216k - $252k per year
...seek individuals who embody our core traits: Scrappy, Curious, Optimistic, Persistent, and Empathetic . Your role As a Staff Software Engineer, you’ll own the platformization roadmap for shared services, architecting a plan to decouple existing functionality into...Work at office$173.4k - $238.35k per year
...can use deep data insights to improve their business. Founded by engineers — and customer obsessed — we leap at every opportunity to solve... ...virtual machines. And we're only getting started. As a Fullstack software engineer, you will work with your team and product management to...Summer workWorldwide$173.4k - $238.35k per year
...world's best data and AI infrastructure platform so our customers can use deep data insights to improve their business. Founded by engineers — and customer obsessed — we leap at every opportunity to solve technical challenges, from designing next-gen UI/UX for interfacing...Summer workWorldwide$80k - $100k per year
...execution. You aren’t just managing a roadmap; you are building the engine that allows our entire team to scale. You will work side-by-side... ...of the "Total Life Internal OS," ensuring non-technical staff have the digital tools they need to provide world-class care....Full timeRemote workFlexible hours- ...jump in to solve problems. Our actions reflect our values of honesty, reliability, openness, and humility. Your Role: Staff Software Engineers at Treasure AI are pod leads — player-coaches who remain active contributors while setting the technical and process...Contract workWork at officeNight shift
- ...right next to the Canada Line. We have a satellite office in Cloverdale, should you need somewhere in the ‘burbs. We have effective staff meetings and fun activities to open our communication channels and foster teamwork. We are passionate about mentorship and the...Full timeWork at office
$145k - $155k per year
...real-world conditions. We’re hiring two Software Developers to join the Casper team: one... ...and respond, Casper is adapting too. The engineering work behind it is becoming more complex,... ...’s already doing a lot. This is not a staff-level architecture role. You’ll contribute...Remote jobLong term contractFull timeShift workNight shift$110k - $155k per year
...employees. We are building a talent pool for future Rendering Engineers (intermediate or senior) to help build and optimize a visually... ...senior levels; title/level will match experience).Strong Unreal Engine 5 (or UE4 plus meaningful UE5 experience) and modern C++ skills....Remote jobLong term contractPermanent employmentFull time- ...through this video primer!: Our interview process for Senior Software Developers respects the growth you’ve experienced while... ...various internal programs for refreshing and strengthening your engineering and customer facing skills, and work with your manager to make...Full timeLive out
- ...We’re a growing SaaS company with startup energy, industry respect, and global reach. The Opportunity We are seeking a Staff Software Engineer to serve as the technical anchor for our team as we pivot to an AI-driven, agentic development workflow. Much of our code is...Local areaWorldwide
$147k - $174k per year
...Your role We are seeking a talented, experienced Full-Stack Engineer passionate about building high-quality, scalable web applications... ...Skills you’ll bring ~5+ years of experience in Full-stack software engineering. ~ Strong experience with writing high-performance...Work at office$98.52k - $123.16k per year
...Opportunity We are seeking a talented and motivated Full Stack Software Developer to join our dynamic team. The ideal candidate will... ...stack web development Bachelor’s degree in computer science, Engineering, or a related field, or equivalent experience Proven...Full timeManual laborWork at office- ...innovation, above and beyond fleeting trends, Marvell is a place to thrive, learn, and lead. Your Team, Your Impact As an Analog Layout Staff Engineer with Marvell, you'll be a member of the Central Engineering business group. If you picture Marvell as a wheel, Central Engineering...Full time
- ...specifically for electrical contractors and distributors. Founded by an engineer to solve field inefficiencies, our award-winning, SOC 2 Type II–... ...innovation. ABOUT THE ROLE We are seeking a Team Lead, Software Engineering to guide a focused, high-performing agile pod of...For contractorsWork at office
$98.52k - $123.16k per year
...Opportunity The Insurance Council of BC, the regulatory authority for the insurance industry in British Columbia, is seeking a skilled Software Developer specializing in Dynamics 365 Customer Service to join our technology team. This role is responsible for designing,...Full timeWork at office- ...reliability, openness, and humility. Your Role: As a Senior Software Engineer on the Realtime/Personalization team, you will be a core... ...building Treasure AI's AI-native Personalization Studio and RT2.0 engine. Operating at the intersection of deep engineering and product...Long term contractContract workWork at officeNight shift
- ...high-growth, well-funded SaaS company that helps answer questions software development teams have about their applications. This allows... ...built. About You You’re an experienced software development engineer with a track record of building and shipping products that customers...Long term contractFull timeWork at officeFlexible hours
$82k - $110k per year
...sustainable business connections through digital and physical channels. Job Description As an Intermediate Quality Software Engineer on our Integration team, you'll own the quality of the integrations that connect our Accounts Payable platform with the different...Full timeWork at officeLocal areaRemote workFlexible hoursShift work$110.8k - $157.89k per year
...and professional excellence. At LMI our staff work passionately toward the common goal... ...What will you do as a Senior Vision Software Developer? LMI is seeking a Software Developer... ...in a multi-disciplinary, multi-platform, engineering team (software, electrical, mechanical/optical...
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Staff Software Engineer - SRE & AIOps. Be the first to apply!
- embedded software Vancouver, BC
- junior software developer .net Vancouver, BC
- spécialiste assurance qualité logiciel Vancouver, BC
- airline software Vancouver, BC
- software Vancouver, BC
- software intern Vancouver, BC
- software implementation project manager Vancouver, BC
- software asset management analyst Vancouver, BC
- software technical support Vancouver, BC
- software support Vancouver, BC


