Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer - Imunify Reliability Platform

This position is listed on behalf of a partner company, who manages all applications and next steps. Our partner is looking for a Lead Site Reliability Engineer - Imunify Reliability Platform based in Canada.

This is a greenfield SRE leadership opportunity within a large-scale, security-focused product environment. You will define what “healthy” means across roughly 70 components spanning cloud services and customer-hosted agents. You will establish SLIs, SLOs, error budgets, monitoring standards, alerting, and escalation practices from the ground up. Your work will directly improve the ability to detect silent security-control degradation before it becomes a widespread customer-impacting issue. You will collaborate closely with engineering leads and senior engineers while building the telemetry and reliability platform yourself. The environment is remote-first, async, technically demanding, and focused on measurable outcomes rather than dashboards for their own sake. This is an opportunity to shape the reliability culture and foundations of a major security product.

Accountabilities
  • Define and establish meaningful SLIs for approximately 70 product components, working with squad leads and senior engineers to agree on ownership, measurement, tiering, SLOs, and error budgets.
  • Develop a reliability taxonomy covering service availability and latency, fleet reachability and configuration convergence, security-control efficacy, artifact delivery, and telemetry pipeline health.
  • Ensure reliability indicators are independently measurable and cannot be disabled by the same failure they are intended to detect.
  • Design and build the telemetry collection pipeline for customer-hosted agents and cloud services, balancing push-based collection, sampling, privacy constraints, data quality, and cardinality.
  • Extend instrumentation across Python, Go, and Rust components in collaboration with product engineering teams.
  • Consolidate existing dashboards, queries, and reporting mechanisms into a smaller, more reliable observability platform, retiring tooling that does not provide meaningful operational value.
  • Implement symptom-based, SLO-driven alerting with multi-window burn-rate principles and clear page, ticket, and dashboard classifications.
  • Ensure every production alert has a defined owner, documented failure mode, and actionable runbook.
  • Establish ongoing alert-quality practices, including periodic reviews, measurable actionable-alert rates, and deliberate removal of unnecessary alerts.
  • Build a machine-readable ownership and escalation model that routes incidents to the appropriate engineering squads.
  • Establish severity definitions, acknowledgement expectations, follow-the-sun escalation practices, and clean handoff procedures across multiple time zones.
  • Strengthen incident command and blameless postmortem practices, including reliable timelines, ownership, and follow-through on corrective actions.
  • Coach engineering squads to own their own operational responsibilities and paging rather than becoming a centralized buffer for other teams' alerts.
  • Deliver measurable reliability outcomes over the first year, including complete SLI ownership, production telemetry, tiered alerting, squad on-call adoption, and a significant reduction in the time required to detect silent security-control degradation.
  • Requirements

    • Substantial production engineering or SRE experience, including experience defining and implementing an SLO framework rather than simply operating within an existing one.
    • Strong Python skills and the ability to read and modify Go or Rust code when implementing instrumentation and reliability improvements.
    • Strong hands-on experience with time-series and event telemetry at scale, including Prometheus/OpenMetrics, Grafana, Alertmanager-class routing systems, and columnar or high-cardinality data stores such as ClickHouse or equivalent technologies.
    • Experience debugging distributed systems running on bare metal and long-lived hosts; this role requires more than Kubernetes-centric operational experience.
    • Practical experience with production-scale configuration management and CI/CD tooling such as Ansible, GitLab CI, Jenkins, or comparable technologies.
    • Strong understanding of telemetry for systems that cannot be directly scraped or fully controlled, including push-based collection, sampling, clock skew, partial reporting, and privacy considerations on customer-managed infrastructure.
    • Excellent written and asynchronous communication skills, with the ability to align multiple engineering teams around measurable definitions of system health.
    • Strong engineering judgment and a pragmatic approach to observability, reliability, alerting, and operational ownership.
    • Experience with security products such as WAF, EDR, antivirus, or vulnerability-management platforms is a strong advantage, particularly an understanding that security-control reliability must measure effective enforcement rather than simple uptime.
    • Familiarity with monitoring and continuous-monitoring requirements related to SOC 2, ISO 27001, NIST SP 800-137, or similar frameworks is valuable.
    • Exposure to OpenTelemetry, eBPF, Sentry, cost-aware telemetry, or cardinality-management techniques is beneficial.
    • Experience working effectively with AI-assisted development tools and modern agentic engineering workflows is a plus.
    • Kubernetes experience is useful for supporting the smaller portion of the platform that runs in Kubernetes.
    • This role is focused on SRE and reliability engineering rather than DevOps ticket management, build-system ownership, cloud cost management, or acting as the on-call team for other engineering squads.
    • Benefits

      • Fully remote work with flexible working hours, enabling you to work from anywhere worldwide.
      • 24 paid vacation days per year.
      • 10 paid national holidays.
      • Unlimited sick leave.
      • Compensation toward private medical insurance.
      • Co-working space reimbursement.
      • Gym and sports reimbursement.
      • Professional development opportunities through challenging technical projects, learning opportunities, mentoring, and knowledge-sharing programs.
      • Opportunity to receive a reward for an innovative idea that can be patented.
      • Remote-first, asynchronous working environment spanning multiple time zones.
      • Opportunity to define an SRE function, reliability standards, and operational culture from the ground up.
Vacancy posted more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer - Imunify Reliability Platform. Be the first to apply!