Senior Site Reliability Engineer (Kubernetes)
Mirantis
About Mirantis
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment—on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy.
Job Description
We are looking for a highly experienced and driven Senior Site Reliability Engineer to join our forward-thinking cloud development and operations team. In this role, you will contribute to the design, development, and operation of sophisticated cloud-based AI solutions built on the CNCF ecosystem including Kubernetes, running on cutting-edge hardware from leading vendors. This role focuses mainly on deploying AI infrastructure built on NVIDIA-certified hardware, following architecture and implementation designs produced by our engineering team. You will play a pivotal role in ensuring the reliability, security, and performance of container infrastructure, while mentoring team members and Mirantis customers to deliver high-quality software and services. As a senior engineer, you will work closely with stakeholders to define technical strategies, solve complex challenges, and ensure the seamless integration of cloud and software services. This is an excellent opportunity to make a significant impact while driving innovation in a rapidly evolving cloud ecosystem.
Main Responsibilities:
Work with geographically distributed international teams on technical challenges and process improvements.
Develop, implement, maintain, and troubleshoot cloud and AI infrastructure solutions based on open source software.
Collaborate with stakeholders to gather and refine technical requirements.
Optimize system performance, reliability, and scalability.
Troubleshoot, debug, and resolve complex technical issues.
Participate in code reviews to maintain high quality standards.
Stay up to date with industry trends and best practices in cloud operations and development.
Design and implement AI-driven automation across the DevOps lifecycle, including code development and maintenance.
Facilitate knowledge transfer to customers during the delivery phases.
Qualifications
5+ years of professional experience in DevOps, with a strong focus on Cloud, infrastructure technologies and Kubernetes
Experience with high-performance data center processing, networking, and storage
Exposure to Golang and working knowledge of other programming languages (Python, JavaScript).
Strong knowledge of distributed systems, microservices architecture, and CI/CD pipelines.
Exceptional problem-solving and debugging skills across networking and storage (hardware and software), Linux, and Kubernetes, with attention to performance optimization and security.
Demonstrated ability to lead technical tasks and collaborate effectively with diverse teams.
Comfortable making independent judgment calls when working directly with customers, often with limited day-to-day oversight.
Excellent written and spoken English.
Excellent customer-facing communication skills.
A commitment to innovation, continuous learning, and delivering high-quality results.
Ability to travel up to 25% if needed, including internationally.
Nice to have
Extensive experience in network and/or storage architecture.
Experience with high-performance computing or GPU infrastructure: GPU scheduling, MIG/vGPU, RDMA/RoCE or InfiniBand fabrics, NVLink, DCGM health-checking, GPU driver/firmware lifecycle or NVIDIA AI Enterprise.
Working experience with Openstack
Presence in the open source community including upstream contribution and conference presentations.
Prior experience with commercial container and virtual compute infrastructure platforms such as Rancher, Openshift, and VMware.
Education and Experience:
Bachelor's degree in Computer Science or a related field, or equivalent experience.
At least 5 years of DevOps or Software Development experience or in a similar role.
Additional Information
What does Mirantis offer you?
- Work with an established Silicon Valley leader in the cloud infrastructure industry;- Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
- Be a part of cutting-edge, open-source innovation;
- Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
- Professional development and training;
- Attend conferences and working groups;
- Company outings, happy hours, hackathons, and tech talks;
- Receive a competitive compensation package with a strong benefits plan.
It is understood that Mirantis, Inc. may use automated decision-making technology (ADMT) for specific employment-related decisions. Opting out of ADMT use is requested for decisions about evaluation and review connected with the specific employment decision for the position applied for. You also have the right to appeal any decisions made by ADMT by sending your request to View email address on jobs.smartrecruiters.com
We are a Leader for Container Management in G2 (#2 after AWS)!
- ...Senior Site Reliability Engineer - Edge Location : Ottawa/Toronto, On-Site Reports to: Head of Security The Role You own the edge compute module — the standard hardware stack, OS image, and runtime that integrates with Dominion Dynamics's mesh radios, sensors...SeniorFull time
$130k - $180k per year
...to legally work in Canada (visa or sponsorship won't be provided) Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure (those with Azure will be prioritized first) About Us:...SeniorFull timeRemote workVisa sponsorshipWork visaFlexible hours$78.62 - $89.34 per hour
Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering and Operations team. In this engineering-focused role, you will move beyond traditional database administration to act as...SeniorLong term contractPermanent employmentFull timeContract workWork at office$164.6k - $235.1k per year
...About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that... ...automation. We are seeking an experienced and visionary Senior SRE Manager to lead and grow our newly built Site Reliability...SeniorLong term contractRemplacementFull timeContract workTemporary workLocal areaFlexible hours$110k - $120k per year
...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring... ...reporting, SLO performance metrics, and incident trends to senior management. What You Bring : Technical Proficiency:...SeniorTemporary workInternshipWork at officeRemote work$100k - $125k per year
...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,... ...culture Tech at Tipalti Our tech teams are the engine behind our business. Tipalti’s tech ecosystem is extremely...Full timeWork at officeFlexible hours- ...best of both work styles in a workplace that is intentional about belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a systems thinker. You’ll create middleware and platform...SeniorFull timeWork at officeLocal areaRemote workWorldwideMonday to fridayFlexible hours
$78.62 - $89.34 per hour
Our client, is seeking a talented and proactive Site Reliability Specialist / Senior Cloud Platform Engineer to join their core Cloud Engineering division. In this engineering-focused role, you will move far beyond basic operational support to act as a principal architect...SeniorLong term contractPermanent employmentContract workWork at office- ...Platform Engineer – DevOps, Site Reliability Engineering (SRE) & Dynatrace Required Skills • Strong experience as a Platform Engineer with expertise in DevOps and Site Reliability Engineering (SRE). • Experience designing, implementing, automating, and supporting...Permanent employment
- ...Site Reliability Engineer Location: Toronto ON Work Model: Hybrid (2 days per week in-person at the Toronto office preferred) Required... ...and distributed tracing strategies. Work with Docker Kubernetes ECS and AKS environments. Design fault-tolerant highly...Full timeWork at office2 days per week
- ...Site Reliability Engineer – APM, Dynatrace, Observability Role Description • Deep application and system-level knowledge across complex end-to-end environments, including tightly integrated on-prem and cloud-native services supporting large-scale, multi-tier transaction...Contract work
$62.87k - $147.5k per year
...At Capgemini Engineering, the world leader in engineering services, we bring together a global... ...the job you’re considering As a Kubernetes/DevOps Engineer, you will work on one of... ...licenses, Relevant experience and skills, Seniority and performance, Market and business consideration...SeniorPermanent employmentFull timeLocal areaRemote work- ...Months Experience: 6–8 Years Required Skills Kubernetes & Containers (Kubernetes 1.24+, CRDs, Operators) Prometheus,... ...Collaborate with SRE, DevOps, and application teams to improve platform reliability and reduce MTTR. Preferred Skills AI/ML for...Contract work
$141k - $191k per year
...and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence. About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for: Team Leadership...Work at officeLocal areaFlexible hours2 days per week3 days per week$140k - $160k per year
...organizations in Canada/US. We are currently hiring Senior Forward Deployed Engineer - AI & Kubernetes for our client in the Toronto area, which specializes... ...mindset: you care about production rollout, adoption, reliability, operational handoff, and measurable impact. Willingness...SeniorPermanent employmentFull time$166k - $195k per year
...developer velocity and increases system reliability by building the foundational... ...platforms and tools that power Robinhood engineering. Within this group, the Kubernetes Compute team focuses on building... ...of technical growth! As a Senior Software Develope r, you will...SeniorWork at officeFlexible hoursShift work3 days per week- ...Application Release Engineer Top 3 Required Skills: • Azure, Kubernetes • CI/CD, DevOps • Python, Bash Years of Experience Required: 7-10+ years Job Description: Drives the automation of all parts of the application deployment pipeline including...Contract work
$160k - $220k per year
...on this mission. If you are too, let's talk. Staff Software Reliability Engineer - Data Platform About the Team The Data Platform team... ...Databricks, Hadoop ~ Coordinators and schedulers like the ones in Kubernetes, Hadoop, Mesos Experience in developing and tuning...Local areaWorldwideFlexible hours- ...We are seeking a Senior QA Engineer to ensure the quality, reliability, and performance of our Edge Kubernetes Platform , supporting both connected and air-gapped environments. This is a hands-on role where you will design and implement automated testing frameworks...Senior
- ...Core Skills: · Amazon Web Service (AWS) Cloud Computing · Kubernetes Essential Skills: · More than 5 years of... ...· Microsoft Certified Azure Administrator (AZ-104) or DevOps Engineer Expert (AZ-400) preferred. · Scripting skills in Bash, Python...Contract work
- ...We are seeking an Intermediate Kubernetes Platform Engineer to join our platform engineering team and contribute to the delivery of our Edge Kubernetes... ...environments. In this role, you will work closely with senior engineers to deploy, operate, and optimize Kubernetes-based...Senior
- ...Deploy and manage services on Kubernetes-based platforms such as Amazon EKS and Google Kubernetes Engine (GKE). ~ Provision and manage... ...solutions to enhance reliability and efficiency. ~ Implement... ...years of experience in a DevOps, Site Reliability Engineering, or Cloud...SeniorWorldwide
$150k - $170k per year
...General Information: Job Title: Senior Engineering Manager Location: Toronto, ON (Onsite... ...scalability, security, maintainability, reliability, and cost efficiency. Partner with... ...and orchestration using Docker and Kubernetes. Knowledge of cloud security, identity...SeniorLong term contractFull timeInternshipWork at officeImmediate startRemote workFlexible hours$170k - $220k per year
...startup. They are currently seeking a Senior Infrastructure Engineer to design and build an internal... ...to deploy software with confidence, reliability, and speed. This individual will bring... ...-based Pulumi codebase and Kubernetes-based runtime environments. Engineering...SeniorFull timeInternshipRemote work$136k - $160k per year
...developer velocity and increases system reliability by building the foundational... ...platforms and tools that power Robinhood engineering. Within this group, the Kubernetes Compute team focuses on building... ...platforms. Working alongside senior engineers, you will write code to...SeniorWork at officeFlexible hoursShift work3 days per week$250k per year
...Role: Kubernetes & HPC Platform Engineer – Trading Client: Elite FinTech Compensation: $120,000 - $250,000 CAD + Bonus Location: Toronto Overview My client are seeking an Engineer to join a highly technical team focused on the scheduling side of technology...Permanent employmentImmediate start- ...teamwork, and continuous improvement. As a Senior Project Superintendent , you will lead... ...role is responsible for coordinating site activities, mentoring project teams, managing... ...degree in Construction Management, Civil Engineering, or a related field (or an equivalent...SeniorRemote jobLong term contractFull timeContract workFor subcontractor
$65k - $75k per year
...Faster, more frequent, and reliable access to rapid transit with more... ...Job Description The Site Administrator is a key member... ...as providing daily updates to senior leadership on project progress... ...Construction Management, Civil Engineering, Business Administration, or related...SeniorFull timeContract workFor subcontractorWork at office- ...future of how work gets done. AS A SENIOR SOFTWARE ENGINEER YOU WILL: Drive high-impact initiatives... .... Extend the product to operate reliably in regulated and air-gapped... ...containers and orchestration (Docker, Kubernetes) and cloud infrastructure is a strong...SeniorFull time
$91k - $99k per year
...for all to use. Overview: As a Senior DevOps Engineer, you will join a collaborative team dedicated... ...that empower our engineering team to reliably deploy, monitor, and scale their... ...containerized applications using Docker/Podman, Kubernetes, and Helm. Develop, maintain, and...SeniorPermanent employmentFull timeWork at officeLocal area2 days per week1 day per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Senior Site Reliability Engineer (Kubernetes). Be the first to apply!
- senior site reliability engineer Toronto, ON
- site reliability engineer intern Toronto, ON
- site reliability engineer Toronto, ON
- site reliability engineer remote Toronto, ON
- senior solution architect Toronto, ON
- senior pattern maker Toronto, ON
- senior digital marketing manager Toronto, ON
- senior administrative officer Toronto, ON
- sénior service a la clientèle Toronto, ON
- senior mechanical designer Toronto, ON
