Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliable Engineer (Canada - Remote)

$80 - $110 per hour
Full-time

Axon-networks

Toronto, ON
  • Remote job




AXON Networks delivers a robust AI-driven, analytics-based orchestration platform and a wide portfolio of next-gen high-speed routers that leverage the newest Wi-Fi technologies. Together, these technologies give ISPs the ability to manage and troubleshoot their networks in real time, and to deliver an outstanding customer experience.

AXON Networks is a trusted strategic partner for its customers, helping them evaluate their current technologies and business models, and creating and executing strategies that enable them to innovate faster, accelerate their digital transformations, and strengthen their relationships with consumers.

AXON Networks is headquartered in Irvine, CA USA with Asia HQ in Singapore and also operating in Denmark, Spain and Vietnam.

The Site Reliability Engineer  will improve the availability, performance, scalability and recoverability of AXON Networks cloud solutions. You will combine software engineering with hands-on NOC operations to make the complete cloud-to-device service path observable, supportable and resilient at fleet scale.  


You will help establish practical SRE capabilities inside the NOC while partnering closely with Support, Operations, cloud and DevOps Engineering. You will participate in a sustainable on-call rotation and improve the NOC’s ability to diagnose customer-impacting issues.

 


Role mandate  






  • Own reliability outcomes for assigned cloud services  





  • Improve observability, capacity, resilience and recovery  





  • Define and operationalize service-level indicators, service-level objectives and actionable alerting.  





  • Automate repetitive NOC work and create safe, testable mechanisms for diagnosis, recovery, device operations and routine production changes.  





  • Lead technically during incidents, drive evidence-based learning and ensure high-value corrective actions are completed.  



What you will own  






  • Establish reliability baselines, SLIs, SLOs and error budgets for cloud services and critical device-management workflows such as onboarding, provisioning, configuration, telemetry collection, command execution and firmware delivery.  





  • Trace failures across the end-to-end service path: cloud APIs and microservices, Kubernetes and infrastructure, databases and messaging, internet and access-network dependencies, device-management protocols and the devices  





  • Identify fleet-wide and customer-specific failure patterns involving device reachability, session stability, configuration drift, command latency, telemetry gaps, firmware behavior and cloud capacity.  





  • Contribute operability requirements and production evidence during design and readiness reviews  





  • Maintain NOC dashboards for service health, device reachability, provisioning success, command and telemetry performance, firmware adoption and customer impact.  





  • Participate in the NOC production on-call rotation and serve as a technical incident lead or senior troubleshooter when appropriate .  





  • Diagnose complex failures across applications, cloud infrastructure, Kubernetes, APIs, networking, DNS/TLS, databases, messaging platforms, device-management sessions and CPE behavior.  





  • Coordinate evidence gathering and technical escalation with service-provider customers, Engineering, firmware, DevOps and vendors while maintaining clear mitigation, recovery and handoff.  





  • Lead or contribute to post-incident reviews; convert recurring device, platform and process failures into prioritized and measurable corrective actions.  





  • Develop production-grade software, scripts and workflows for diagnosis, remediation, deployment safety, fleet analysis, scaling, maintenance and recovery.  






  • Improve CI/CD and GitOps practices for operational software and infrastructure, including automated testing, release validation, progressive delivery and rollback readiness.  





  • Manage or contribute to infrastructure as code, configuration as code and reusable self-service patterns for cloud and NOC operations.  





  • Measure NOC toil and partner with Automation & Tools Engineers to prioritize durable platform capabilities instead of fragmented one-off scripts.  





  • Develop capacity models for service-provider growth, managed-device populations, telemetry volume, messaging throughput, API demand and rollout events.  





  • Create and maintain runbooks, troubleshooting decision trees, service maps, device and cloud dependency records, known-error guidance and operational knowledge.  





  • Coach NOC and Support personnel on diagnosis, safe mitigation, evidence capture and escalation across cloud, network and CPE layers.  





  • Build self-service diagnostic views and tools that help the NOC determine scope, affected customers, device cohorts, likely fault domain and next action.  





  • Share reliability insights with Engineering and Product and contribute to reliability reviews, operational-readiness reviews and continuous-improvement priorities.  



Required qualifications  






  • 5+ years of experience in site reliability engineering, production engineering, DevOps, cloud infrastructure, systems engineering or a closely related role.  





  • Strong software or automation skills in Python, Go, Java, Bash or a comparable language, with experience producing maintainable operational code.  





  • Hands-on experience operating distributed production systems in a public cloud environment and troubleshooting across application, infrastructure, network and device-integration layers.  





  • Experience with Google Cloud Platform, Oracle Cloud Infrastructure and production Kubernetes environments.  





  • Experience with infrastructure as code and delivery tooling such as Terraform, Helm, Git-based CI/ CD and policy-as-code.  





  • Strong Linux, containers and Kubernetes fundamentals, including deployment behavior, resource management, networking and failure diagnosis.  





  • Strong troubleshooting & debugging skills in Kubernetes platforms.  





  • Experience with modern observability practices and tools across metrics, logs, traces, alerting, dashboards and synthetic monitoring.  





  • Familiarity with Prometheus, Grafana, OpenTelemetry or equivalent observability ecosystems.  





  • Familiarity with Apache Pulsar or similar distributed messaging and streaming platforms handling requests from millions of devices.  





  • Experience participating in an on-call rotation and responding effectively to high-severity, customer-impacting production incidents.  






  • Working knowledge of SLOs, error budgets, capacity planning, resilience engineering, change safety and blameless incident learning.  





  • Strong networking knowledge, including TCP/IP, DNS, DHCP, TLS, routing, NAT, load balancing and systematic packet- or session-level troubleshooting.  





  • Clear communication, disciplined documentation and the ability to collaborate across NOC, cloud, DevOps, firmware and service-provider teams.  





  • Bachelor’s degree in computer science, engineering or equivalent practical experience.  



Preferred qualifications  






  • Experience supporting cloud-managed CPEs such as broadband gateways, routers, ONTs, Wi-Fi/mesh systems or similar edge devices in a service-provider environment.  





  • Familiarity with TR-069/CWMP, TR-369/USP, TR-181 data models, ACS or USP controller platforms, device telemetry and remote lifecycle management.  





  • Experience supporting messaging and streaming platforms such as Apache Pulsar or Kafka, APIs and highly available databases used in device-management control planes.  





  • Understanding of access technologies such as GPON/XGS-PON, DOCSIS, Ethernet or fixed wireless and how CPE, ONTs and provider networks interact.  





  • Experience with firmware rollout automation, canary or cohort deployments, fleet health analysis and safe rollback practices.  





  • Experience building auto-remediation, safe self-service operations or internal reliability platforms.  





  • Experience supporting multiple service-provider customers in a 24×7 telecommunications, broadband or managed-network environment.  


Type: Contract

Compensation: 80 CAD/hr - 110 CAD/hr

 

Join AXON Networks!

At AXON Networks, we promote equal opportunities in all our recruitment processes, ensuring non-discrimination on the basis of gender, age, origin, disability, or any other personal circumstances. We assess talent based on objective criteria and foster an inclusive and diverse working environment.

Vacancy posted 13 hours ago
Similar jobs that could be interesting for youBased on the Site Reliable Engineer (Canada - Remote) in Toronto, ON vacancy
  • $110k - $120k per year

     ...experience, we’re the team behind Money Mart—Canada’s largest non-bank branch network—and a...  ...Environment – Flexibility to balance remote and in-office collaboration; enjoy our...  ...that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer... 
    Remote work
    Full time
    Temporary work
    Internship
    Work at office

    Momentum Financial Services Group

    Toronto, ON
    21 hours ago
  • $100k - $125k per year

     ...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,...  ...culture Tech at Tipalti  Our tech teams are the engine behind our business. Tipalti’s tech ecosystem is extremely... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Tipalti

    Toronto, ON
    21 hours ago
  • $153.82k - $277k per year

     ...a place where you can thrive, we can’t wait to meet you. Site Reliability Engineers (SREs) at Braze are responsible for keeping all internal-facing...  ...a high comfort level collaborating asynchronously across remote-first, global engineering teams Preferred / Nice to Have:... 
    Remote work
    Permanent employment
    Full time
    Internship
    Work at office
    Local area
    Flexible hours
    Rotating shift

    Braze

    Toronto, ON
    21 hours ago
  • $140k - $155k per year

     ...Caseware is one of Canada's original Fintech companies, having led...  ...This is a hands-on senior engineering role focused on improving production...  ...teams to ship secure, reliable, and scalable software with confidence...  ...Location:   This is a remote location open to candidates legally... 
    Remote job
    Permanent employment
    Full time
    Flexible hours

    Caseware

    Toronto, ON
    21 hours ago
  •  ...belonging at iManage. Mondays and Fridays are reserved for (remote-friendly) focus time to get things done. Have the best of both...  ..., collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an engineer, a builder, and a systems... 
    Remote work
    Full time
    Work at office
    Local area
    Worldwide
    Monday to friday
    Flexible hours

    Imanage

    Toronto, ON
    21 hours ago
  • $130k - $180k per year

     ...based out of North America Role is 95% remote in Toronto (we meetup 1x a month). Must be able to legally work in Canada (visa or sponsorship won't be provided) Our...  ...and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main... 
    Remote work
    Full time
    Visa sponsorship
    Work visa
    Flexible hours

    Acquird.io

    Toronto, ON
    21 hours ago
  • $140k - $180k per year

     ...Tripstack Founded in Toronto, Canada in 2016, Tripstack has been...  ...that all other engineering teams at Tripstack depend on....  ...the deployment patterns and reliability standards that apply across both...  ...depth - interconnects, BGP, site-to-site VPN, cross-region peering... 
    Remplacement
    Full time
    Work at office
    Immediate start
    Flexible hours

    Tripstack

    Toronto, ON
    21 hours ago
  •  ...learning never stops. The Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team of engineers dedicated...  ...Product, Engineering, and Support leadership to integrate reliable SRE practices into early planning and product delivery... 
    Long term contract
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    21 hours ago
  • $144k - $200k per year

    **The Team** Platform Engineering is the department within SRE that is...  ...developing and maintaining the reliable and globally connected multi-...  ...or Vancouver offices, or fully remote from anywhere in North America...  ...* We are seeking a talented Site Reliability Engineer (SRE)... 
    Remote work
    Full time
    Work at office
    Worldwide
    Flexible hours

    Mongodb

    Toronto, ON
    21 hours ago
  •  ...Docebians around the world and help us reinvent the way people learn, because learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident response while also shaping the underlying infrastructure that... 
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    21 hours ago
  • $140k - $182k per year

     ...client base with operations throughout North America, Central America, Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and software development. You will support and help inform the evolution of... 
    Full time

    Movable Ink

    Toronto, ON
    21 hours ago
  • $154k - $200k per year

     ...client base with operations throughout North America, Central America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic technical leadership across infrastructure and software development... 
    Long term contract
    Full time

    Movable Ink

    Toronto, ON
    21 hours ago
  • $110k - $125k per year

     ...Apply today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability,...  ...operational excellence. Some of the benefits we offer: * Remote Work Environment * Flexible Time Away From Work Policy including... 
    Remote work
    Full time
    Flexible hours

    Smile Digital Health

    Toronto, ON
    21 hours ago
  •  ...competitive advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in software...  ...experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly evaluate products... 
    Full time

    Serigor Inc

    Toronto, ON
    21 hours ago
  •  ...Cohere is a team of researchers, engineers, designers, and more, who are...  ...-performance, scalable and reliable machine learning systems? Do you...  ...? We are looking for a Site Reliability Engineer to join the...  ...and workspace improvement Remote-flexible, offices in Toronto,... 
    Remote work
    Full time
    Work at office
    Flexible hours

    Cohere

    Toronto, ON
    21 hours ago
  •  ...visibility, and optimize spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s...  ...Excellence & Automation Design, develop, and automate reliable cloud infrastructure and platform services. Apply Infrastructure... 
    Full time
    Manual labor
    Local area
    Flexible hours

    Emburse

    Toronto, ON
    21 hours ago
  •  ...that are particularly strong in a few areas, and have some interest and capabilities in others. About the Role: As a Site Reliability Engineer , you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS... 
    Full time

    Kong Company

    Toronto, ON
    21 hours ago
  • $120k - $170k per year

     ...We are seeking a highly skilled and motivated Senior DevOps Engineer to join our dynamic team and play a key role in designing, implementing...  ..., investigate incidents, and drive improvements that increase reliability and operational efficiency; Write and maintain automation,... 
    Full time
    Work at office
    Local area
    Flexible hours

    Magnet Forensics

    Toronto, ON
    21 hours ago
  • $85k - $100k per year

     ...best practices. Collaborate cross-functionally with Product, Engineering, Support, and other stakeholders to design, develop, test, implement...  ...$85,000 - $100,000 a year Some of the benefits we offer: * Remote Work Environment * Flexible Time Away From Work Policy... 
    Remote job
    Long term contract
    Full time
    Flexible hours

    Smile Digital Health

    Toronto, ON
    21 hours ago
  • $130k - $145k per year

     ...reasons to SMILE! The Cloud Security Engineer is responsible for designing, automating...  ...internal tools, improve system health and reliability. Accountable for ensuring that all working...  ...Some of the benefits we offer: * Remote Work Environment * Flexible Time Away From... 
    Remote job
    Full time
    Flexible hours

    Smile Digital Health

    Toronto, ON
    21 hours ago
  •  ...seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability...  ...role is   hybrid, based out of Toronto, Canada. We welcome candidates at various experience...  ...of AI hardware by helping build highly reliable systems that power tomorrow's largest... 
    Permanent employment
    Full time
    Internship
    Second job

    Tenstorrent

    Toronto, ON
    21 hours ago
  •  ...emerging threats. About the Role As the Strategic Sales Engineer for the Canadian territory at VulnCheck, you will play a...  ...strengthening their cybersecurity posture. This is a 100% remote role based in Canada, though we are primarily looking for candidates in the... 
    Remote work
    Long term contract
    Full time
    Temporary work

    Vulncheck

    Toronto, ON
    21 hours ago
  • $110k - $130k per year

     ...excellence.  If you are motivated by impact, growth, and purpose, you will find a strong sense of belonging at Quince. THE ROLE Senior Site Merchandiser We are seeking a Senior Site Merchandiser t o join our growing team. The Senior Site Merchandiser will play a... 
    Full time
    Seasonal work
    Work at office
    Local area

    Quince

    Toronto, ON
    21 hours ago
  •  ...our high-performing team of Solution Engineers/Field Engineers in Canada. You will collaborate closely with regional...  ...and will require travel to customer sites where appropriate. It can be...  ...product demonstrations, both in-person and remotely, that engage a range of stakeholders... 
    Remote work
    Long term contract
    Full time
    Internship

    Quantexa

    Toronto, ON
    21 hours ago
  • $80 - $120 per hour

     ..., Larry Summers , and Jack Dorsey . Position: Incident management / reliability / SRE Evaluator Type: Contract Compensation: $80–$120/hour Location: Remote Role Responsibilities Evaluate AI-generated artifacts against domain... 
    Remote work
    Full time
    Contract work
    Summer work
    Work at office

    Mercor

    Toronto, ON
    21 hours ago
  •  ...iteration and increasing intelligence via additional agentic computation. About The Role Join Cerebras as a Performance & Reliability Engineer within our innovative Co-Design and Next Generation Team. Our groundbreaking CS-3 system has set new benchmarks in high-... 
    Full time

    Cerebras Systems

    Toronto, ON
    21 hours ago
  • $100k per year

     ...problems. We are growing our team and looking for contributors of all seniorities. Tenstorrent is looking for an experienced Reliability Engineer to drive and execute reliability testing for our products from wafer to board level.    As a Principal Reliability Engineer... 
    Permanent employment
    Full time
    Internship

    Tenstorrent

    Toronto, ON
    21 hours ago
  •  ...Employment Status: Permanent Schedule: 40 hours/week – 100% remote work Job Description We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability, performance, and resilience of high-availability SaaS platforms. Working... 
    Remote work
    Permanent employment
    Work at office
    Local area

    TOTEM Recruteur de talent

    Toronto, ON
    3 days ago
  • $170k - $185k per year

     ...Interested in joining one of Canada’s top-performing asset managers? We are seeking a Head of Platform Engineering, Reliability & Control to lead our horizontal engineering function...  ...strategies that enable frequent, reliable, and low-risk software releases.   ~... 
    Full time
    Work at office
    3 days per week

    Connor, Clark

    Toronto, ON
    21 hours ago
  • $163k - $185k per year

     ...are seeking a Manager, Data Engineering with a strong technical background...  ...next move for you! This is a remote role overseeing a team of 2...  ...data quality and reliability. Define success metrics:...  ...driving continuous improvement. Canada Perks & Benefits: ~ Annual... 
    Remote work
    Full time
    Manual labor
    Work from home
    Flexible hours

    Zenni Optical

    Toronto, ON
    21 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliable Engineer (Canada - Remote). Be the first to apply!