Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE) - Dynatrace & AI Observability

Full-time

Astra North Infoteck Inc.

Site Reliability Engineer

Location: Toronto ON

Work Model: Hybrid (2 days per week in-person at the Toronto office preferred)

Required Skills

Site Reliability Engineering (SRE)

DevOps

Dynatrace

Role Summary

Design implement and optimize Site Reliability Engineering (SRE) and DevOps practices to ensure high system availability performance and reliability across distributed environments.

Leverage Dynatrace Davis AI automation and cloud technologies to enable proactive monitoring intelligent automation and operational excellence.

Role Description

Dynatrace & AI-Driven Observability

Lead the implementation and optimization of the Dynatrace platform across applications and infrastructure.

Utilize Dynatrace Davis AI for automated root cause analysis anomaly detection event correlation predictive performance insights and alert noise reduction.

Configure and manage OneAgent deployments Smartscape topology mapping service flow and distributed tracing.

Define and monitor SLIs SLOs and user experience metrics.

Build custom dashboards alerts and observability pipelines.

Integrate Dynatrace with CI/CD pipelines for release validation and performance gating.

Integrate Dynatrace with incident management tools such as PagerDuty and ServiceNow.

Enable self-healing automation using Dynatrace event triggers and AI-driven insights.

Automation & Configuration Management

Design and implement automation solutions using Ansible.

Automate configuration management application deployments and environment provisioning.

Develop reusable Ansible playbooks and roles for scalable operations.

Automate operational tasks patching compliance processes and remediation workflows.

Integrate Ansible with CI/CD pipelines and monitoring systems.

Cloud & DevOps

Design and manage cloud-native solutions on AWS with exposure to Azure.

Develop infrastructure using Terraform CloudFormation or AWS CDK.

Build and manage CI/CD pipelines using GitHub Actions Jenkins or GitLab CI.

Develop and deploy serverless solutions using AWS Lambda API Gateway and Step Functions.

Automate DevOps and operational workflows using Python (boto3) and Bash scripting.

Deploy and maintain production environments through automated pipelines.

Optimize cloud infrastructure for cost performance and scalability.

Monitoring & Reliability Engineering

Monitor and manage AWS CloudWatch and Azure Monitor/Log Analytics.

Design unified observability across multi-cloud environments.

Implement logging and distributed tracing strategies.

Work with Docker Kubernetes ECS and AKS environments.

Design fault-tolerant highly available and disaster recovery solutions.

Support incident response on-call activities and root cause analysis (RCA).

Required Qualifications

Proven experience with Dynatrace APM Real User Monitoring (RUM) and infrastructure monitoring.

Strong hands-on experience with Dynatrace Davis AI capabilities.

Experience with Ansible for automation and configuration management.

Deep knowledge of AWS services and cloud-native architectures.

Experience with Infrastructure as Code using Terraform CloudFormation or AWS CDK.

Proficiency in Python (boto3) and Bash scripting.

Experience supporting production-scale environments.

Business Analyst experience.

Scrum Master experience.

Nice to Have

Dynatrace Associate or Professional certification.

Experience with Dynatrace APIs and automation.

Experience building self-healing systems using AI-driven triggers.

Familiarity with Prometheus Grafana and the ELK Stack.

Azure cloud experience and certifications.

Experience with GitOps and Platform Engineering.

Required Skills:

Top 3 Required Skills: 1. IBM Financial transaction 2. Payment flow 3. Support Modernization Detailed Job Description: Design develop and maintain applications built on IBM Financial Transaction Manager (FTM) to support core payments processing. Contribute to the development of payment flows supporting transaction processing. Build and support integrations between FTM and upstream/downstream systems using enterprise integration patterns. Participate in the design development testing deployment and production support. Troubleshoot and resolve application and integration issues in a complex regulated environment. Collaborate with architecture QA and operations teams to ensure platform stability scalability and performance. Support modernization initiatives and enhancements to existing payment hub capabilities. Produce clear technical documentation and participate in code reviews and knowledge sharing.

Vacancy posted more than 2 months ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) - Dynatrace & AI Observability in Toronto, ON vacancy
  •  ...competitive advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in...  ...scientific principles of experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly... 
    Suggested
    Full time

    Serigor Inc

    Toronto, ON
    1 day ago
  •  ...drives business value and inspires extraordinary results.  Our AI-powered platform helps organizations modernize financial...  ...visibility, and optimize spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s systems... 
    Suggested
    Full time
    Manual labor
    Local area
    Flexible hours

    Emburse

    Toronto, ON
    1 day ago
  • $116k - $235.1k per year

     ...About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission... 
    Suggested
    Long term contract
    Remplacement
    Full time
    Temporary work
    Work at office
    Local area
    Immediate start
    Flexible hours
    2 days per week

    Tubi - Canada

    Toronto, ON
    4 days ago
  • $130k - $180k per year

     ...legally work in Canada (visa or sponsorship won't be provided) Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure (those with Azure will be prioritized first) About Us: We're... 
    Suggested
    Full time
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Acquird.io

    Toronto, ON
    1 day ago
  •  ...Job Role: Senior Platform Engineer – DevOps, SRE & Dynatrace Observability Location : Toronto - Hybrid (3 Days Work from Office) Job Summary We are seeking a highly skilled and experienced Platform Engineer with strong DevOps and Site Reliability Engineering (... 
    Suggested
    Full time
    Work at office

    Astra North Infoteck Inc.

    Toronto, ON
    a month ago
  • $110k - $120k per year

     ...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for...  ...UAT, production), ensuring changes are safe, repeatable, and observable. Design and maintain automated CI/CD pipelines and enforce... 
    Full time
    Temporary work
    Internship
    Work at office
    Remote work

    Momentum Financial Services Group

    Toronto, ON
    1 day ago
  • $140k - $155k per year

     ...This is a hands-on senior engineering role focused on improving production...  ...teams to ship secure, reliable, and scalable software with confidence...  ...objectives (SLOs), improve observability, automate operational...  ...on cloud-native technologies, site reliability engineering principles... 
    Remote job
    Permanent employment
    Full time
    Flexible hours

    Caseware

    Toronto, ON
    1 day ago
  • $140k - $182k per year

     ...for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable...  ..., Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and software... 
    Full time

    Movable Ink

    Toronto, ON
    1 day ago
  • $141k - $191k per year

     ...at Thomson Reuters and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence...  .... About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for:... 
    Work at office
    Local area
    Flexible hours
    2 days per week
    3 days per week

    Thomson Reuters

    Toronto, ON
    more than 2 months ago
  •  ...Artificial Intelligence. Actual Impact. At Docebo, we’re using AI to change how people learn at work—and we mean actually...  ...learning never stops. The Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team of engineers dedicated... 
    Long term contract
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    1 day ago
  • $140k - $180k per year

     ...visit: The role Our SRE team runs the shared infrastructure that all other engineering teams at Tripstack depend on....  ...Prometheus, Thanos, and Grafana observability stack. We are in the middle...  ...the deployment patterns and reliability standards that apply across both... 
    Remplacement
    Full time
    Work at office
    Immediate start
    Flexible hours

    Tripstack

    Toronto, ON
    1 day ago
  •  ...Artificial Intelligence. Actual Impact. At Docebo, we’re using AI to change how people learn at work—and we mean actually...  ...because learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident response... 
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    1 day ago
  • $154k - $200k per year

     ...for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable...  ...America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic... 
    Long term contract
    Full time

    Movable Ink

    Toronto, ON
    1 day ago
  • $100k - $125k per year

     ...Position Summary: We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,  you will play a crucial role in enhancing the reliability, performance, and scalability of our systems... 
    Full time
    Work at office
    Flexible hours

    Tipalti

    Toronto, ON
    1 day ago
  • $153.82k - $277k per year

     ...a place where you can thrive, we can’t wait to meet you. Site Reliability Engineers (SREs) at Braze are responsible for keeping all internal-facing...  ...messages to our customers' end-users daily. This Senior SRE role is specifically focused on supporting the Braze Ruby on... 
    Permanent employment
    Full time
    Internship
    Work at office
    Local area
    Remote work
    Flexible hours
    Rotating shift

    Braze

    Toronto, ON
    1 day ago
  •  ...SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe...  ...belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an engineer, a builder, and a... 
    Full time
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to friday
    Flexible hours

    Imanage

    Toronto, ON
    1 day ago
  • $80 - $110 per hour

     ...AXON Networks delivers a robust AI-driven, analytics-based orchestration platform and a wide portfolio of next-gen high...  ...Singapore and also operating in Denmark, Spain and Vietnam. The Site Reliability Engineer  will improve the availability, performance, scalability and... 
    Remote job
    Full time
    Contract work

    Axon-networks

    Toronto, ON
    1 day ago
  • $92k - $118k per year

     ...Build reliable, resilient cloud platforms that keep critical financial services running. The Role We are looking for an experienced Site Reliability and DevOps Engineer to join the SRE team within our growing Corporate Action Processing group. You will bring 3+ years... 
    Internship
    Immediate start

    Capco

    Toronto, ON
    4 days ago
  • $144k - $200k per year

    **The Team** Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure...  ..., deployment machinery, and observability and alerting systems. The Fabric...  ...in developing and maintaining the reliable and globally connected multi-cloud... 
    Full time
    Work at office
    Remote work
    Worldwide
    Flexible hours

    Mongodb

    Toronto, ON
    1 day ago
  • $110k - $125k per year

     ...systems bringing #BetterGlobalHealth to patients everyday! Apply today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across... 
    Full time
    Remote work
    Flexible hours

    Smile Digital Health

    Toronto, ON
    1 day ago
  •  ...About The Role   Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work.   Based in Toronto or remote... 
    Full time
    Remote work

    bosonai

    Toronto, ON
    23 days ago
  •  ...work Job Description We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability,...  ...improve scalability and resilience. Define and evolve observability practices, including dashboards, alerts, SLOs, and SLIs.... 
    Permanent employment
    Work at office
    Local area
    Remote work

    TOTEM Recruteur de talent

    Toronto, ON
    5 days ago
  •  ...and enterprises who are building AI systems to power magical...  ...Cohere is a team of researchers, engineers, designers, and more, who are passionate...  ...high-performance, scalable and reliable machine learning systems? Do...  ...? We are looking for a Site Reliability Engineer to join the... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    Toronto, ON
    1 day ago
  •  ...strong in a few areas, and have some interest and capabilities in others. About the Role: As a Site Reliability Engineer , you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS platform that powers the... 
    Full time

    Kong Company

    Toronto, ON
    1 day ago
  •  ...youll be doing As a member of CIBCs Application Reliability Engineering Platform team the Consultant Site Reliability Engineering will play a key role in...  ...scalability and efficiency. Define and maintain observability strategies to ensure visibility into application behaviour... 
    Full time
    Contract work
    3 days per week
    1 day per week

    Canadian Imperial Bank of Commerce

    Toronto, ON
    a month ago
  •  ...is part of the Production Support and Reliability Engineering team, responsible for ensuring the stability...  ...Support, Production Support, SRE, or Platform Operations roles. Strong...  ...reliability. Knowledge of monitoring and observability tools for proactive system management.... 
    Permanent employment
    Full time
    Local area

    Capgemini

    Toronto, ON
    a month ago
  • $140k - $165k per year

     ...leader driving a major IT transformation, transitioning from traditional service desk operations to an AI-driven, self-healing operational fabric. As Lead Site Reliability Engineer (SRE), you will lead enterprise observability, AIOps automation, and SRE governance across a... 

    Randstad

    Toronto, ON
    26 days ago
  •  ...LinkedIn profiles! This is a hands-on senior engineering role focused on improving production...  ...engineering teams to ship secure, reliable, and scalable software with confidence....  ...service level objectives (SLOs), improve observability, automate operational processes, and drive... 
    Permanent employment
    Full time
    Remote work

    Caseware

    Toronto, ON
    a month ago
  • $78.62 - $89.34 per hour

    Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering and Operations team. In this engineering-focused role, you will move beyond traditional database administration to act as... 
    Long term contract
    Permanent employment
    Full time
    Contract work
    Work at office

    Randstad

    Toronto, ON
    more than 2 months ago
  • $170k - $185k per year

     ...managers? We are seeking a Head of Platform Engineering, Reliability & Control to lead our horizontal...  ...role, you will shape a more connected, observable, and automated technology ecosystem spanning...  ...for command-and-control capabilities, SRE practices, process orchestration,... 
    Full time
    Work at office
    3 days per week

    Connor, Clark

    Toronto, ON
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - Dynatrace & AI Observability. Be the first to apply!