Cloud Site Reliability Engineer II
$100k - $120k per yearBarracuda Networks
Come join our passionate team! Barracuda is a leading cybersecurity company providing complete protection against complex threats. Our platform protects email, data, applications, and networks with innovative solutions, and a managed XDR service, to strengthen cyber resilience. Hundreds of thousands of IT professionals and managed service providers worldwide trust us to protect and support them with solutions that are easy to buy, deploy, and use.
We are committed to a candidate selection process and work environment that is inclusive and barrier free. To ensure candidates are assessed in a fair and equitable manner, accommodations will be provided to prospective employees in accordance with the Accessibility for Ontarians with Disabilities Act (AODA) and the Ontario Human Rights Code.
Envision yourself at Barracuda
As a Cloud Site Reliability Engineer II on the CloudOps Platform team, you will focus on the observability, monitoring, dashboards, and automation systems that power Barracuda’s next-generation multi-Tenant Kubernetes platform. While your primary mission centers on delivering deep operational visibility, reliable telemetry pipelines, and proactive alerting, you will also play a key role in improving the broader Kubernetes platform running across AWS and Azure.
Our team values collaborative knowledge sharing, automated reliability, and modern engineering practices. We actively embrace AI-assisted workflows (Claude Code, OpenCode, Codex CLI) to accelerate development, diagnostics, and routine platform maintenance. In this role, you will build and operate centralized observability stacks (Grafana, Loki, Mimir, Tempo), create actionable dashboards, automate telemetry via GitOps and Terragrunt, and partner with internal engineering teams to optimize application reliability in production.
Tech Stack Exposure:
- Observability & Monitoring: Grafana, Prometheus / Mimir, Loki, Tempo, OpenTelemetry / Grafana Alloy, Sloth (SLOs), Alertmanager
- Orchestration & Compute: Kubernetes (AWS EKS, Azure AKS), Helm, Kustomize
- IaC & Cloud Provisioning: Terragrunt, Terraform
- GitOps & CI/CD: ArgoCD, GitHub Actions
- Public Clouds: Amazon Web Services (AWS), Microsoft Azure
- Languages & Scripting: Python, Bash (Go is a plus)
- AI Developer Tooling: Claude Code, OpenCode, Codex CLI, GitHub Copilot
What you’ll be working on
- Operating, scaling, and automating our centralized LGTM telemetry infrastructure (Loki for logs, Mimir/Prometheus for metrics, Tempo for distributed tracing, and Grafana for unified visualization).
- Designing intuitive, high-impact Grafana dashboards and executive health overviews for platform services, Kubernetes clusters, and tenant workloads.
- Establishing reliable alerting strategies, SLO/SLI tracking (via Sloth), and notification routing to detect and resolve degradation before it impacts production systems.
- Automating the deployment of log collectors, metric exporters, and monitoring agents across multi-cluster EKS and AKS environments using GitOps (ArgoCD) and Terragrunt.
- Directly contributing to core MTK Kubernetes platform health, performance tuning, and infrastructure modernization.
- Leveraging modern AI coding tools (Claude Code, OpenCode, Codex CLI) to build automation, diagnostic tooling, and operational scripts.
- Collaborating closely with internal product and tenant teams to assist with observability onboarding, distributed tracing instrumentation, and performance troubleshooting.
What you bring to the role
- 2–4 years of experience working with public cloud infrastructure (AWS and/or Azure) with a strong passion for observability, monitoring, and systems reliability.
- 1–2+ years of hands-on experience deploying, operating, or troubleshooting containerized workloads in Kubernetes (EKS/AKS).
- Practical experience configuring, operating, or building dashboards in modern observability stacks (Grafana, ELK, Splunk, etc).
- Working knowledge of Infrastructure as Code using Terraform and/or Terragrunt, and GitOps delivery workflows (ArgoCD or Flux).
- Solid scripting skills in Python or Bash for system automation, telemetry pipelines, and operational tooling (Go is a plus).
- Curiosity and eagerness to leverage AI coding agents (Claude Code, OpenCode, Codex CLI, GitHub Copilot) in everyday engineering workflows.
- Strong analytical troubleshooting instincts, clear communication skills, and a collaborative team mindset.
What you’ll get from us
A team where you can voice your opinion, make an impact, and where you and your experience are valued. Internal mobility – there are opportunities for cross training and the ability to attain your next career step within Barracuda.
- Equity, in the form of non-qualifying options
- High-quality health benefits
- Retirement Plan with employer match
- Career-growth opportunities
- Flexible Time Off and Paid Time Off benefits
- Volunteer opportunities
What you’ll get from us:
A team where you can voice your opinion, make an impact, and where you and your experience are valued. Internal mobility – there are opportunities for cross training and the ability to attain your next career step within Barracuda. In addition, you will receive equity, in the form of non-qualifying options.
The anticipated salary range for this role is $100,000 CAD to $120,000 CAD. Actual compensation offered will be dependent upon the individual's skills, experience, and qualifications as they directly relate to the requirements of the position, the budget for the position, and applicable employment law.
Location: Ottawa, ON
#LI-hybrid
27-0459(a)
- ...based in Ottawa, Ontario. We are looking for a Manager, Site Reliability Engineering to lead a team responsible for the reliability, availability... ...includes products such as Email Security Gateway and Cloud Email Archiving. What You’ll be Working on: Lead, coach...SuggestedFull timeRemote workWorldwideFlexible hours
- ...leader who thrives at the intersection of reliability, platform engineering, and operational excellence. You enjoy... ...Define and execute the platform and site reliability strategy, aligning... ...continuous improvement of scalable, resilient, cloud-native platforms using public cloud...SuggestedLong term contractFull timeWork at officeRemote work
$95.2k - $119k per year
...for We are looking for an experienced database engineer to help SurveyMonkey evolve its data stores toward a cloud-native, globally distributed architecture. You... ...design, implement, and optimize data systems that are reliable, scalable, and secure. What you’ll be...SuggestedFull timeFlexible hours- ...leading global satellite operator, providing reliable and secure satellite-delivered... ...for over 55 years. Backed by a legacy of engineering excellence, reliability and industry-leading... ...Telesat on X and LinkedIn or visit The Cloud Network Engineer will assist with the planning...SuggestedFull timeInternshipLive InWork at officeWorldwideRelocationRelocation package
$120k - $170k per year
.... At General Dynamics Mission Systems–Canada, we’re not just engineering technology — we’re shaping the future of defence and security. Our... ...General Dynamics Mission Systems–Canada is seeking a Principal, Cloud Engineering AWS Manager to own the design, build, and operation...SuggestedFull timeInternshipFlexible hours$74k - $98k per year
...shape the quality strategy behind BarracudaONE, our next-generation cloud platform that unifies the customer experience across Barracuda's... ..., Claude Code, and Playwright MCP to improve test coverage, engineering efficiency, and release confidence. We work in a supportive, high...Full timeWorldwide- ...Description The Role: The Solution Test Engineer will work to bridge the gap between... ...responsible for testing and validating innovative cloud solutions and delivering implementation... ...VMWare, ESX/ESXi. Linux/Microsoft OS, IIS, Apache, Oracle RAC, File Sharing Systems,...Permanent employmentFull time
- ...expertise across all areas of IT including cloud technology, software solutions, and... ...Implementation Specialist - Senior Linux and Cloud Engineer Reports to: Technical Services Manager... ...potential for weekly travel to customer sites. About the Role We are seeking a...Full time
$105k - $115k per year
...industrial automation. We're looking for a Cloud DevOps Developer to join our team... ...systems teams, and cloud specialists to build reliable applications that run in modern cloud... ...Science, Software Development, Computer Engineering, or a related field. ~5+ years of relevant...Full timeManual labor- ...individual for our team. Reporting to the General Superintendent, the Site Superintendent is responsible for the leading all daily field... ...Collaborate with project managers, architects, engineers and subcontractors to understand project goals, plans, and specifications...Long term contractFull timeFor contractorsFor subcontractorWork at officeLocal area
$122k - $152.5k per year
...At General Dynamics Mission Systems–Canada, we’re not just engineering technology — we’re shaping the future of defence and security.... ...impact programs that matter. Job Description The Manager, Reliability, Maintainability, Testability & Safety (RMT&S) leads a team of...Full timeInternshipFlexible hours- ...Plant Running. Improve Whats Next. Were partnering confidentially with a leading Canadian food manufacturer seeking a Mechanical Reliability Engineer to play a key role in improving equipment reliability maintenance strategy and plant performance within a highly automated...Long term contractPermanent employmentFull timeFor contractorsRelocation package
$105k - $125k per year
...high-performance user-plane functions on hyperscale cloud platforms. Working alongside experienced engineers and technical experts, you will help accelerate... ...plane capabilities, improve system performance and reliability, and solve complex networking and scalability challenges...Full time- ...EDA vendors, but we're taking it to the cloud. We're leading the way and our future keeps... ...bigger and brighter! The Metrics engineering team is based in Ottawa, Canada with additional... ...on our platform team by ensuring the reliability, scalability, and performance of our...Full time
$115k - $130k per year
...place! We're looking for a Product Manager II to join our Wholesale Network team. You... ...direction. You will work closely with engineering, design, data, and go-to-market partners to... ...exceptional customer experiences. Our cloud commerce solution transforms and unifies online...Full timeManual laborImmediate startRemote workFlexible hours$90k - $120k per year
...About the Role At Solace, reliability is a feature we build deliberately... ...it. As an Intermediate Chaos Engineer on our R&D team, you’ll design... ...Simulator (FIS), and other cloud-provider fault injection services... ...in chaos engineering, site reliability engineering, or a...Full timeInternshipWorldwide2 days per week- ...Contracting Inc. (MECI) is seeking a Superintendent to manage all site activities performed on their designated project in British... ...Requirements: Grade 12 diploma or equivalent Civil/Environmental Engineering Degree or Civil Engineering Technologist or Technician Diploma...Full timeContract workFor subcontractorShift work
- ...Job Responsibility: Senior Azure Cloud Engineer Position Overview: We are seeking a highly skilled and experienced Senior Azure Cloud... ..., and Recovery Services. Ensure high availability and reliability of cloud environments through proper configuration and monitoring...Full time
- ...Description Location: Ottawa, ON (Hybrid/On-site as required) Client: Federal Government... ...(ATO) Specialist to support a private cloud environment. This role is focused on... ...based platforms. This is not a hands-on engineering or deployment role. Instead, the successful...Full time
- ...Superintendent for our Lansdowne 2.0 Project in Ottawa As the Superintendent, you will be the senior field leader responsible for all on-site construction activities. You will oversee day-to-day site operations, manage field staff and trade partners, and ensure the project is...Full timeFor subcontractorWork at office
- ...seasoned developer who has moved beyond just writing code to architecting solutions. You’ve shipped production-grade applications, navigated cloud environments, and possess a deep curiosity for how AI can augment the development lifecycle. You’re comfortable shifting from deep-...Permanent employmentFull timeRemote workHome officeFlexible hoursShift work
- ...Job Responsibility: Converge Technology Solutions is a leading Cloud Partner in AWS, Azure, and GCP. We enable companies, both large and... ...computing. We are looking to hire an experienced Public Cloud Delivery Engineer with deep expertise to support our customers with cloud adoption...Full time
$185k - $215k per year
...client is seeking a Director of Engineering to lead the software... ...delivery of high-quality, secure, reliable, maintainable, and scalable software... ...evaluation and adoption of cloud, microservices, automation, artificial... ..., and qualifications. On-site ~ Ottawa , Ontario ,...Long term contractPermanent employmentFull timeFor contractors$90k - $105k per year
...our people—and we invest in their development every step of the way. About the Role Reporting to the Field Manager, the Assistant Site Superintendent plays a crucial role in supporting the successful execution of construction project assisting the Site Superintendent in...Long term contractFull timeContract workFor contractorsFor subcontractorWork at office- ...worlds top brands offering comprehensive engineering supply chain and manufacturing solutions.... ...industries and a vast network of over 100 sites worldwide Jabil combines global reach with... ...Summary The Electrical Design Engineer II supports the design development testing and...Full timeInternshipWork at officeLocal areaWorldwide
- ...life. We are looking for a dependable, organized, and proactive site coordinator to support the day-to-day operations of our... ...initiative and solve problems independently Valid driver's licence and reliable transportation Positive attitude with a willingness to be...Full timeFor subcontractor
- ...Is this you? Are you a Network Engineer with experience supporting cloud and managed infrastructure environments... ...connectivity solutions across customer sites and cloud environments, including... ...procedures, and enhancing network reliability and security What do you know?...Full timeCasual workManual laborWork at officeLocal areaRemote workFlexible hours
$57.02k - $67.25k per year
...Career Opportunity Position Title: Site Supervisor, Headstart Child Care Classification: Early Childhood Development Worker, level 3 Reporting to: Manager of Children and Youth Services Job Type: Permanent, 1.0 FTE (35 hours per week) Department: Family...Hourly payPermanent employmentReliefWork at office$70k - $100k per year
...Job Location: Greater Ottawa Region, Ontario Work Mode: On-Site Salary: $70,000-$100,000 annually Regular Hours: Monday-Friday... ...problems independently without requiring constant direction is reliable, practical, and follows through on commitments can adapt quickly...RemplacementFull timeSeasonal workMonday to friday$90k - $120k per year
...grow online revenue. By unifying site monitoring, experience... ...we're looking for a Solutions Engineer to help us win. You'll be embedded... ...: conversion rate, site reliability, and revenue Field and resolve... ...Shopify, Salesforce Commerce Cloud, Magento are a plus! Strong...Full timeFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Cloud Site Reliability Engineer II. Be the first to apply!
- cloud infrastructure architect Ottawa, ON
- cloud engineer Ottawa, ON
- associate cloud engineer Ottawa, ON
- cloud engineer remote Ottawa, ON
- cloud network engineer Ottawa, ON
- cloud operations engineer Ottawa, ON
- junior cloud engineer Ottawa, ON
- entry level cloud engineer Ottawa, ON
- site reliability engineer Ottawa, ON
- site reliability engineer intern Ottawa, ON

