Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Senior Site Reliability Engineer

$145k - $185k per year
Full-time

Parallel Domain

Remote
  • Remote job


About the Role


Before an autonomous vehicle navigates a busy intersection, before a robot learns to pick and place in a warehouse, before any Physical AI system is trusted in the real world, it has to prove itself in ours. Parallel Domain builds the platform that validates the next generation of autonomous systems in high-fidelity virtual environments, and the infrastructure underneath that platform is what makes simulation at scale possible.

We're hiring a Senior Site Reliability Engineer to help build and operate that infrastructure. This role sits at the core of how we run large-scale, distributed simulation workloads for autonomous-systems testing and validation. You'll work across multi-region AWS infrastructure, operate Kubernetes at scale, and contribute directly to reliability, security, and deployment systems that the rest of the engineering org depends on.

This is a hands-on role with the broad ownership typical of a startup. You'll partner closely with platform, simulation, and ML teams to keep the system running smoothly and evolving. We're growing the team—two of these roles are open—and the work is substantive: multi-region GPU scheduling, Windows workloads on Kubernetes, large-scale batch simulation, and an enterprise product direction that will require rethinking parts of how we deploy and operate.

Responsibilities



  • Infrastructure ownership and cloud operations. Design, build, and maintain multi-region AWS infrastructure using Terraform. Operate and scale EKS clusters across production regions: autoscaling, node lifecycle, workload health. Manage networking across environments: VPC design, DNS, load balancing, and cross-region connectivity. Support infrastructure changes, migrations, and expansions into new regions. Contribute to and improve GitOps-based deployment workflows using GitHub Actions, Helm, and Kustomize.



  • Reliability engineering and incident response. Help build and run incident management processes: severity definitions, escalation paths, on-call practices. Lead incident response, debugging, and root-cause analysis. Write postmortems and drive systemic reliability improvements from what they surface. Improve observability across metrics, logging, tracing, and dashboards. Support GPU and batch workloads running on Kubernetes.



  • Security and access management. Provide security-conscious feedback on platform architecture decisions. Own cloud IAM governance: roles, policies, and access boundaries across accounts and services. Lead compliance-adjacent work including audit-readiness, partner certification requirements, and supporting responses to customer security questionnaires.


  • Platform tooling and developer experience. Improve CI/CD pipelines and infrastructure validation. Support engineers with infrastructure debugging, environment setup, and performance issues. Contribute to tooling and automation in Python and Bash. Take on adjacent responsibilities as needed in a startup environment.

 

Required Qualifications



  • Experience. 5+ years in SRE, DevOps, or infrastructure engineering roles, with a track record of operating production systems across multiple regions.



  • Terraform. Modules, state management, and multi-environment patterns.



  • AWS depth. Solid experience across VPC, IAM, EKS, S3, and CloudWatch.



  • Kubernetes expertise. Cluster operations, autoscaling, RBAC, and Helm.



  • CI/CD and GitOps. Experience with GitHub Actions, ArgoCD, or similar workflows.



  • Networking fundamentals. CIDR, DNS, load balancing, VPN, and cross-region connectivity.



  • Observability. Experience with tooling such as Prometheus and Grafana.



  • Scripting. Comfort with Python and Bash for tooling and automation.



  • Cross-platform familiarity. Working knowledge of both Linux and Windows environments. Operational experience supporting Windows-based workloads is a meaningful advantage.


  • Pragmatism and ownership. Comfortable in a fast-moving startup with evolving priorities. You take ownership of systems while collaborating closely with other teams, and you're pragmatic about tradeoffs between speed, reliability, and complexity.

 

Preferred Qualifications



  • Windows on Kubernetes. Experience with Windows node pools, Windows AMIs, and GPU-adjacent components on K8s.



  • GPU scheduling. Familiarity with GPU scheduling on Kubernetes, including NVIDIA device plugin configuration.



  • Domain workloads. Experience supporting simulation, ML, or rendering workloads in cloud infrastructure.



  • AWS extras. Exposure to AWS Storage Gateway, Active Directory integrations, or AWS Transfer Family.



  • Service mesh. Familiarity with service proxy or service mesh patterns.



  • Container OS. Experience with container-optimized OS images (e.g., Bottlerocket, Packer).


  • Cost optimization. Cloud cost optimization at scale.

 

Core Tools

Terraform · AWS · Kubernetes · Helm · Kustomize · ArgoCD · GitHub Actions · Prometheus · Grafana · Docker · Python · Bash

What Makes a Great Candidate

You think in failure modes and proactively surface issues. You hold a principled view on security and push back constructively when designs introduce unnecessary risk. You communicate clearly across engineering, product, and customer-facing teams, flagging issues with urgency proportional to customer impact. You take end-to-end ownership of complex efforts and know when to push for the clean solution versus the pragmatic one.

Base salary range of CAD $145,000–$185,000, depending on skills, qualifications, and experience, plus equity, full health/dental/vision coverage, learning stipend, and generous vacation. This role is remote-friendly across Canada and the US Pacific Northwest.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

Vacancy posted 6 hours ago
Similar jobs that could be interesting for youBased on the Senior Site Reliability Engineer in Remote vacancy
  • $110k - $160k per year

     ...person to join our team working towards this goal, we would love to hear from you!  Role Overview  We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for... 
    Senior
    Full time
    Internship
    Work at office
    Local area
    Flexible hours

    Magnet Forensics

    Remote
    6 hours ago
  • $123k - $160k per year

     ...learning, and shaping the future of customer growth, you’ll find your place here. We are seeking a highly experienced Senior Site Reliability Engineer to own the reliability, performance, and operational excellence of our large-scale, distributed infrastructure. You will... 
    Senior
    Remote job
    Long term contract
    Full time
    Relocation

    Branch Metrics

    Remote
    6 hours ago
  •  ...We are seeking a Senior DevOps & Site Reliability Engineer to own the reliability, scalability, performance, and operational excellence of Medeloop’s platform. This role blends deep DevOps engineering—CI/CD pipelines, infrastructure as code, and cloud architecture—with SRE... 
    Senior
    Hourly pay
    Full time

    Medeloop

    Remote
    6 hours ago
  • $140k - $180k per year

     ...and online privacy for all. Right now we are looking for a Site Reliability Engineer to help us tame DNS. About the Position Linux system...  ...Thank you for considering this opportunity. Funded.club Senior Recruiters partner exclusively with Startups and are in direct... 
    Senior
    Full time
    Direct hire
    Work at office
    Remote work

    Funded.club

    Remote
    6 hours ago
  • $197.5k - $225k per year

     ...investors including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/CD... 
    Senior
    Full time

    Securityscorecard

    Remote
    6 hours ago
  •  ...in 2024 – but we're just getting started.      As a Sr. Site Reliability Engineer, you'll be the guardian of our platform's reliability and performance...  .... What You Bring to the Team: Design and implement reliable and scalable AWS architecture to meet the needs of the... 
    Senior
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Home office
    Weekend work

    Third-Party Job Posts

    Remote
    6 hours ago
  •  ...corporate culture by MSP Today, G2, and Great Place to Work™, to name a few.  About the role We’re looking for a Senior Site Reliability Engineer (SRE) to help strengthen and scale our multi-cloud platform and developer experience. This is a hands-on senior individual... 
    Senior
    Full time
    Internship
    Remote work
    Work from home

    Scalepad

    Remote
    6 hours ago
  •  ...a part of our journey! About the role We are committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building and leading processes to ensure the reliability,... 
    Senior
    Full time
    Local area
    Remote work
    Home office
    Flexible hours

    Clickhouse

    Remote
    6 hours ago
  • $153k - $187k per year

     ...stress into a clear signal owners can use to run stronger, more resilient businesses. We’re looking for an incredible Senior Site Reliability Engineer to join our SRE team. We aim to make reliability, security, and speed reinforce one another so that the platform becomes... 
    Senior
    Full time
    Internship

    Relay

    Remote
    6 hours ago
  •  ...trillion in AUM and 22 global investment banks. For more information, please visit .     The Role     CMG is looking for a Site Reliability Engineer (SRE) with a strong focus on monitoring, observability, and alerting to ensure the reliability, performance, and... 
    Remote job
    Full time
    Local area

    Capital Markets Gateway

    Remote
    6 hours ago
  • $101.2k - $136.9k per year

     ...strategy across digital banking, core banking, data platforms and member facing services. We are looking for a highly skilled Site Reliability Engineer (SRE) who will help build, operate and continuously improve the reliability, performance, security and automation of our... 
    Permanent employment
    Full time
    Internship
    Work at office
    Immediate start
    Home office
    Flexible hours
    2 days per week

    Vancity

    Remote
    6 hours ago
  •  ...like an environment that you believe could work for you then read on to find out more. The role: We’re looking for a Site Reliability Engineer to manage, maintain, improve and provide support on our platform. You will be curious by nature, always looking for ways to... 
    Full time
    Remote work

    Tyk Technologies Limited

    Remote
    6 hours ago
  • $20 per day

     ...our careers page to see how you can grow with us! As a Site Reliability Engineer at Hiive, you will be responsible for ensuring the...  ...performance and system behavior, and ensuring these services are reliable, scalable, and cost-efficient in production. In this role... 
    Full time
    Summer holiday
    Relocation

    Hiive

    Remote
    6 hours ago
  •  ...environments. We are a modern, IoT-enabled, cloud-based tool for reliability, safety, and operations of physical equipment and...  ...valuing the company at $2.5 billion. We’re looking for a Site Reliability Engineer (SRE) to help advance MaintainX’s reliability, observability... 
    Full time

    Maintainx

    Remote
    6 hours ago
  • $120k - $200k per year

     ...and many more.   ABOUT THE ROLE At LayerZero, our Site Reliability Engineering (SRE) team is at the intersection of software and systems engineering...  ...internal systems to those external users interact with—are reliable, meet the uptime expectations of our users, and... 
    Full time

    Layer Zero Labs Llc

    Remote
    6 hours ago
  •  ...challenges with continuous learning opportunities, then Tescys could be a good fit for you! About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a team at the heart of platform reliability for mission-critical... 
    Long term contract
    Permanent employment
    Full time
    Work at office
    Remote work

    Tecsys Inc.

    Remote
    6 hours ago
  •  ...Canada say it is a great place to work, compared to 60% at a typical company. Role Summary We are seeking an experienced Site Reliability Engineer to help design, build, and operate the infrastructure that underpins the build pipelines that allow our companies to produce... 
    Senior
    Full time
    Contract work
    Local area

    Rivian and Volkswagen Group Technologies

    Remote
    6 hours ago
  • $107k - $161k per year

     ...office to meet with your team for events or meetings. Join Our Team GoDaddy is seeking a highly skilled and motivated Senior Site Reliability Engineer to join our Database Infrastructure team. This role focuses on designing, developing, and deploying automated solutions... 
    Senior
    Full time
    Second job
    Work at office
    Local area
    Remote work
    Work from home

    GoDaddy

    Remote
    6 hours ago
  •  ...might just be in the right place! We’re looking for a Staff Site Reliability Engineer to join our Data team in Canada. As a Staff Data SRE,...  .... ~10% – Mentorship & Technical Review: Actively mentor Senior and Intermediate SREs. You set the bar for IaC quality, observability... 
    Senior
    Full time
    Work at office
    Remote work
    Flexible hours
    Shift work

    Lightspeed Commerce

    Remote
    6 hours ago
  •  ...visualizing relationships between entities in the system. As a Site Reliability Engineer you will be responsible for the availability, latency,...  ...managers. Finally we will ask you to meet with a number of our senior leaders to make sure that you are making the most informed... 
    Senior
    Full time
    Flexible hours

    Behavox

    Remote
    6 hours ago
  •  ...matter. Your Impact As a senior contributor in the APX SRE...  ...relentless about the high quality, reliability, and security our customers...  ...will reach the entire engineering organization to enable product...  ...confidence. Exemplify cloud-native site reliability best practices... 
    Senior
    Long term contract
    Full time
    Remote work
    Flexible hours

    Axon

    Remote
    6 hours ago
  •  ...The Site Reliability Engineering organization at Pinterest is accountable for ensuring overall Pinterest availability as well as enhancing Engineering teams’ capability to design, build and operate robust systems at scale. Pinterest’s applications and infrastructure that... 
    Senior
    Full time
    Work at office
    Relocation
    Relocation package

    Pinterest

    Remote
    6 hours ago
  •  ...made, and we are not tiptoeing into it. We are rebuilding our engineering culture around a simple belief: AI changes everything. How teams...  ...team builds on — and this role is at the center of keeping it reliable, fast, and scalable. As a Staff SRE, you'll own the infrastructure... 
    Remote job
    Full time
    Internship
    Work at office
    Local area
    Flexible hours
    Shift work
    Weekend work

    Babylist, Inc

    Remote
    6 hours ago
  • $120k - $160k per year

     ...ABOUT YOU We are looking for a Site Reliability Engineer (Monetization) who is pragmatic, product-minded, and equally comfortable writing...  ...developers ship features and what infrastructure needs to stay reliable - and to bring the reliability lens into design decisions... 
    Long term contract
    Full time

    Xsolla

    Remote
    6 hours ago
  •  ...client base with operations throughout North America, Central America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic technical leadership across infrastructure and software development... 
    Long term contract
    Full time

    Movable Ink

    Remote
    6 hours ago
  •  ...Gauss Labs is seeking a highly skilled Site Reliability Engineer to join our team in Vancouver. As an SRE at Gauss Labs, you will play a critical role in ensuring our industrial AI platform's reliability, performance, and scalability. You will be responsible for building and... 
    Full time

    Gauss Labs

    Remote
    6 hours ago
  • $180.4k - $230.4k per year

     ...Coalition. About the role We are looking for a Staff Site Reliability Engineer to lead AI enablement across our engineering organization....  ...and tooling infrastructure to ensure AI-generated output is reliable, secure, and production-worthy. This role owns that layer.... 
    Full time
    Remote work
    Home office
    Flexible hours
    Shift work

    Coalition

    Remote
    6 hours ago
  • $69k - $90k per year

     ...UN-endorsed Zero Project.  About the role  As a Junior Site Reliability Engineer (SRE) at Fable, you will help support the reliability,...  ...build more accessible digital experiences, and maintaining reliable, high-performing systems is critical to delivering that impact... 
    Full time
    Internship

    Fable

    Remote
    6 hours ago
  • $260k - $275k per year

     ...enterprises • Solve complex reliability challenges at scale • Influence architecture and engineering culture at a company level •...  ...will focus on creating reusable, reliable, and scalable solutions that abstract...  ..., Platform Engineering, or Site Reliability Engineering role,... 
    Full time

    Saviynt

    Remote
    6 hours ago
  • $150k - $240k per year

     ...obsessed about achieving the high quality and reliability our customers demand. You will work...  ...technical deliverables will reach the entire engineering organization to enable product teams to...  ...-effective. Exemplify cloud-native site reliability best practices. Write code... 
    Long term contract
    Full time
    Remote work

    Axon

    Remote
    6 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Senior Site Reliability Engineer. Be the first to apply!