Site Reliability Engineer - Kubernetes
Astra North Infoteck Inc.
Enterprise Kubernetes SRE (Python, GitOps, API, Container, Cloud,, MongoDB, Postgres)
Toronto, ON - Hybrid (4 Days WFO)
12 months
We are seeking an experienced Site Reliability Engineer to join our Enterprise
Kubernetes Platform team at a leading financial services organization. You'll
be responsible for ensuring the reliability, performance, and scalability of
our enterprise-grade Kubernetes platform that powers mission-critical
applications across the organization.
This role places a strong emphasis on automation, with the expectation that
the successful candidate will continuously identify and eliminate manual toil
through intelligent tooling, self-healing systems, and AI-assisted operational
workflows. You will work alongside platform engineers, DevOps teams, and
application developers to build and maintain a world-class container
orchestration platform.
======================================================================
WHAT YOU'LL DO
======================================================================
PLATFORM RELIABILITY & OPERATIONS
--------------------------------------------------------------------------------
• Ensure 99.9% availability SLA for platform services across 60+ Kubernetes
clusters spanning production, DR, UAT, QA, and development environments
• Manage and operate enterprise Kubernetes distributions across on-premises
and cloud-hosted environments
• Implement and maintain disaster recovery patterns across multi-AZ
architectures and geographically distributed data centres
• Design and execute capacity planning, resource optimization, and cluster
scaling strategies
• Support full cluster lifecycle operations including provisioning, upgrades,
patching, and decommissioning
• Automate cluster health checks and validation workflows for continuous
reliability assurance
• Manage multi-tenant cluster environments with strict isolation and RBAC
enforcement
INCIDENT MANAGEMENT & ON-CALL
--------------------------------------------------------------------------------
• Participate in on-call rotation for platform infrastructure support with
sub-15-minute MTTR targets
• Lead incident response, troubleshooting, and root cause analysis for
platform issues
• Conduct blameless post-incident reviews and implement preventive measures
• Develop and maintain runbooks, troubleshooting guides, and operational
playbooks
• Coordinate with application teams during incidents affecting workloads
• Build automated incident detection and response systems to reduce manual
intervention
• Integrate AI-assisted triage tools for faster incident classification and
resolution
$78.62 - $89.34 per hour
Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering... ...with Linux administration and container orchestration via Kubernetes. Functional exposure to broader cloud ecosystems, such...SuggestedLong term contractPermanent employmentFull timeContract workWork at office$110k - $120k per year
...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for... ...platforms (AWS, Azure). Containers and orchestration (Docker, Kubernetes). Scripting languages (Python, Bash). Infrastructure...SuggestedTemporary workInternshipWork at officeRemote work$164.6k - $235.1k per year
...About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that... ...ALBs/NLBs, Route 53, CloudWatch). Proven experience with Kubernetes in production (EKS preferred), including service exposure,...SuggestedLong term contractRemplacementFull timeContract workTemporary workLocal areaFlexible hours- ...Senior Site Reliability Engineer - Edge Location : Ottawa/Toronto, On-Site Reports to: Head of Security The Role You own the edge compute module — the standard hardware stack, OS image, and runtime that integrates with Dominion Dynamics's mesh radios, sensors...SuggestedFull time
$130k - $180k per year
...Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure... ...implement, and maintain CI/CD pipelines, enabling rapid and reliable software releases. Automate and optimize our infrastructure...SuggestedFull timeRemote workVisa sponsorshipWork visaFlexible hours- ...belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a... ...by default. Scaling cloud infrastructure to support our Kubernetes-based ecosystem. Maintaining the freshness and utility...Full timeWork at officeLocal areaRemote workWorldwideMonday to fridayFlexible hours
- ...Site Reliability Engineer – APM, Dynatrace, Observability Location • Toronto, ON • Hybrid – 2 days in office per week Required Skills • Strong Day 1 expertise in Dynatrace, including: • DQL • Gen3 Dashboards • Traces / Grail...Contract workWork at office2 days per week
- ...Site Reliability Engineer Location: Toronto, ON Work Model: Hybrid (2 days per week in-person at the Toronto office preferred) Required... ...distributed tracing strategies. • Work with Docker, Kubernetes, ECS, and AKS environments. • Design fault-tolerant,...Contract workWork at office2 days per week
- ...Platform Engineer – DevOps, Site Reliability Engineering (SRE) & Dynatrace Required Skills • Strong experience as a Platform Engineer with expertise in DevOps and Site Reliability Engineering (SRE). • Experience designing, implementing, automating, and supporting...Permanent employment
$78.62 - $89.34 per hour
Our client, is seeking a talented and proactive Site Reliability Specialist / Senior Cloud Platform Engineer to join their core Cloud Engineering division. In this engineering-focused role, you will move far beyond basic operational support to act as a principal architect...Long term contractPermanent employmentContract workWork at office$141k - $191k per year
...and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence. About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for: Team Leadership...Work at officeLocal areaFlexible hours2 days per week3 days per week$62.87k - $147.5k per year
...At Capgemini Engineering, the world leader in engineering services, we bring together a global team of engineers, scientists, and architects... ...in the Canada About the job you’re considering As a Kubernetes/DevOps Engineer, you will work on one of the world's largest social...Permanent employmentFull timeLocal areaRemote work- ...We are looking for a Database Reliability Engineer to join our team. This is not a traditional... ...connection pooling configuration and tuning Kubernetes — understanding how workloads connect... ...? Take a look at our careers site and you’ll find everything you’d expect...Permanent employmentFull timeInternshipRemote workWorldwide
- ...Core Skills: · Amazon Web Service (AWS) Cloud Computing · Kubernetes Essential Skills: · More than 5 years of... ...· Microsoft Certified Azure Administrator (AZ-104) or DevOps Engineer Expert (AZ-400) preferred. · Scripting skills in Bash, Python...Contract work
$140k - $160k per year
...currently hiring Senior Forward Deployed Engineer - AI & Kubernetes for our client in the Toronto area,... ...care about production rollout, adoption, reliability, operational handoff, and measurable... ...Willingness to travel to customer sites as needed. A Plus Experience in...Permanent employmentFull time$136k - $160k per year
...developer velocity and increases system reliability by building the foundational... ...platforms and tools that power Robinhood engineering. Within this group, the Kubernetes Compute team focuses on building... ...seamlessly supports reliable application deployments, integrates...Work at officeFlexible hoursShift work3 days per week- ...We are seeking an Intermediate Kubernetes Platform Engineer to join our platform engineering team and contribute to the delivery of our Edge Kubernetes Platform , which supports both connected and air-gapped environments. In this role, you will work closely with senior...
$250k per year
...Role: Kubernetes & HPC Platform Engineer – Trading Client: Elite FinTech Compensation: $120,000 - $250,000 CAD + Bonus Location: Toronto Overview My client are seeking an Engineer to join a highly technical team focused on the scheduling side of technology...Permanent employmentImmediate start$166k - $195k per year
...developer velocity and increases system reliability by building the foundational... ...platforms and tools that power Robinhood engineering. Within this group, the Kubernetes Compute team focuses on building... ...seamlessly supports reliable application deployments, integrates...Work at officeFlexible hoursShift work3 days per week- ...We are seeking a Senior QA Engineer to ensure the quality, reliability, and performance of our Edge Kubernetes Platform , supporting both connected and air-gapped environments... ...into CI/CD pipelines, ensuring rapid, reliable quality feedback. Platform & Edge Validation...
$65k - $75k per year
...to Ontario Faster, more frequent, and reliable access to rapid transit with more than 227... ...option. Job Description The Site Administrator is a key member of the construction... ...degree in Construction Management, Civil Engineering, Business Administration, or related...Full timeContract workFor subcontractorWork at office- ...Job Title: Mechanical Engineer – Offshore Reliability Experience: Minimum 12 Years Qualification: Bachelor’s Degree in Mechanical Engineering Industry: Oil & Gas / Refinery (Offshore) Work Location : Saudi Arab Job Description: The Mechanical Engineer...Permanent employmentFull time
- ...Job Title: Mechanical Engineer – Onshore Reliability Experience: Minimum 12 Years Qualification: Bachelor’s Degree in Mechanical Engineering... ..., and optimize maintenance activities to ensure safe, reliable, and efficient plant operations. # Develop and implement...Permanent employmentFull time
- ...Job Title: Rotating Engineer – Offshore Reliability Experience: Minimum 12 Years Qualification: Bachelor’s Degree in Mechanical Engineering Industry: Oil & Gas / Refinery (Offshore) Work Location : Saudi Arab Job Description: The Rotating Engineer – Offshore...Permanent employmentFull time
$25 - $30 per hour
...Role Overview The Site Hand will be responsible for keeping project sites clean, safe, and stocked. This role involves traveling between... ...handling, and general site upkeep. The ideal candidate is reliable, hardworking, and physically capable of handling labor-intensive...Full timeEarly shift- ...companies! About the Opportunity We are seeking a Junior Site Supervisor to join our Site Operations team in Toronto. This role... ...building operations, clients, and internal teams. ~ Organized, reliable, and able to follow up on tasks until they are complete. ~...Permanent employmentFull timeContract workFor subcontractorManual laborWork at office
- ...Job Title: Rotating Engineer – Onshore Reliability Experience: Minimum 12 Years Qualification: Bachelor’s Degree in Mechanical Engineering... ...ensure safe and efficient plant operations. # Ensure reliable operation and optimal performance of rotating equipment including...Permanent employmentFull time
- ...Python Developer – GCP, Kubernetes, REST APIs & CI/CD Location: Hybrid – 3 Days per Week On-Site Role Summary • Experienced Python Developer with GCP and Kubernetes... ...scalable solutions • Ensure application reliability, security, monitoring, and performance...Contract work3 days per week
- ...Protecnium is an international firm specializing in engineering and technical services (). We are currently looking for a Site Engineer to join our team for a Major Underground Infrastructure Project in Toronto, Canada (on-site position). -Project: Underground Transit...Full timeContract workTemporary workLocal areaMonday to fridayShift workNight shiftDay shift2 days per week
$20 - $20.95 per hour
...today! Allied Universal is seeking Security Guard - Corporate Site in Downtown Toronto, Ontario. Job Title : Security Guard... ...Friday, 1400 - 2200 Overview : We are currently seeking a reliable and dedicated Security Guard for a Corporate Site in Downtown...Hourly payFull timeWork at officeImmediate startMonday to fridayShift workNight shift
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer - Kubernetes. Be the first to apply!
- site reliability engineer intern Toronto, ON
- site reliability engineer remote Toronto, ON
- site reliability engineer Toronto, ON
- senior site reliability engineer Toronto, ON
- site maintenance Toronto, ON
- website developer Toronto, ON
- site safety Toronto, ON
- site carpenter Toronto, ON
- site reliability engineer intern
- site reliability engineer remote
