Site Reliability Engineer (SRE) - Dynatrace & AI Observability
Astra North Infoteck Inc.
Site Reliability Engineer
Location: Toronto ON
Work Model: Hybrid (2 days per week in-person at the Toronto office preferred)
Required Skills
Site Reliability Engineering (SRE)
DevOps
Dynatrace
Role Summary
Design implement and optimize Site Reliability Engineering (SRE) and DevOps practices to ensure high system availability performance and reliability across distributed environments.
Leverage Dynatrace Davis AI automation and cloud technologies to enable proactive monitoring intelligent automation and operational excellence.
Role Description
Dynatrace & AI-Driven Observability
Lead the implementation and optimization of the Dynatrace platform across applications and infrastructure.
Utilize Dynatrace Davis AI for automated root cause analysis anomaly detection event correlation predictive performance insights and alert noise reduction.
Configure and manage OneAgent deployments Smartscape topology mapping service flow and distributed tracing.
Define and monitor SLIs SLOs and user experience metrics.
Build custom dashboards alerts and observability pipelines.
Integrate Dynatrace with CI/CD pipelines for release validation and performance gating.
Integrate Dynatrace with incident management tools such as PagerDuty and ServiceNow.
Enable self-healing automation using Dynatrace event triggers and AI-driven insights.
Automation & Configuration Management
Design and implement automation solutions using Ansible.
Automate configuration management application deployments and environment provisioning.
Develop reusable Ansible playbooks and roles for scalable operations.
Automate operational tasks patching compliance processes and remediation workflows.
Integrate Ansible with CI/CD pipelines and monitoring systems.
Cloud & DevOps
Design and manage cloud-native solutions on AWS with exposure to Azure.
Develop infrastructure using Terraform CloudFormation or AWS CDK.
Build and manage CI/CD pipelines using GitHub Actions Jenkins or GitLab CI.
Develop and deploy serverless solutions using AWS Lambda API Gateway and Step Functions.
Automate DevOps and operational workflows using Python (boto3) and Bash scripting.
Deploy and maintain production environments through automated pipelines.
Optimize cloud infrastructure for cost performance and scalability.
Monitoring & Reliability Engineering
Monitor and manage AWS CloudWatch and Azure Monitor/Log Analytics.
Design unified observability across multi-cloud environments.
Implement logging and distributed tracing strategies.
Work with Docker Kubernetes ECS and AKS environments.
Design fault-tolerant highly available and disaster recovery solutions.
Support incident response on-call activities and root cause analysis (RCA).
Required Qualifications
Proven experience with Dynatrace APM Real User Monitoring (RUM) and infrastructure monitoring.
Strong hands-on experience with Dynatrace Davis AI capabilities.
Experience with Ansible for automation and configuration management.
Deep knowledge of AWS services and cloud-native architectures.
Experience with Infrastructure as Code using Terraform CloudFormation or AWS CDK.
Proficiency in Python (boto3) and Bash scripting.
Experience supporting production-scale environments.
Business Analyst experience.
Scrum Master experience.
Nice to Have
Dynatrace Associate or Professional certification.
Experience with Dynatrace APIs and automation.
Experience building self-healing systems using AI-driven triggers.
Familiarity with Prometheus Grafana and the ELK Stack.
Azure cloud experience and certifications.
Experience with GitOps and Platform Engineering.
Required Skills:
Top 3 Required Skills: 1. IBM Financial transaction 2. Payment flow 3. Support Modernization Detailed Job Description: Design develop and maintain applications built on IBM Financial Transaction Manager (FTM) to support core payments processing. Contribute to the development of payment flows supporting transaction processing. Build and support integrations between FTM and upstream/downstream systems using enterprise integration patterns. Participate in the design development testing deployment and production support. Troubleshoot and resolve application and integration issues in a complex regulated environment. Collaborate with architecture QA and operations teams to ensure platform stability scalability and performance. Support modernization initiatives and enhancements to existing payment hub capabilities. Produce clear technical documentation and participate in code reviews and knowledge sharing.
- ...competitive advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in... ...scientific principles of experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly...SuggestedFull time
- ...drives business value and inspires extraordinary results. Our AI-powered platform helps organizations modernize financial... ...visibility, and optimize spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s systems...SuggestedFull timeManual laborLocal areaFlexible hours
$116k - $235.1k per year
...About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission...SuggestedLong term contractRemplacementFull timeTemporary workWork at officeLocal areaImmediate startFlexible hours2 days per week$130k - $180k per year
...legally work in Canada (visa or sponsorship won't be provided) Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure (those with Azure will be prioritized first) About Us: We're...SuggestedFull timeRemote workVisa sponsorshipWork visaFlexible hours- ...Job Role: Senior Platform Engineer – DevOps, SRE & Dynatrace Observability Location : Toronto - Hybrid (3 Days Work from Office) Job Summary We are seeking a highly skilled and experienced Platform Engineer with strong DevOps and Site Reliability Engineering (...SuggestedFull timeWork at office
$110k - $120k per year
...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for... ...UAT, production), ensuring changes are safe, repeatable, and observable. Design and maintain automated CI/CD pipelines and enforce...Full timeTemporary workInternshipWork at officeRemote work$140k - $155k per year
...This is a hands-on senior engineering role focused on improving production... ...teams to ship secure, reliable, and scalable software with confidence... ...objectives (SLOs), improve observability, automate operational... ...on cloud-native technologies, site reliability engineering principles...Remote jobPermanent employmentFull timeFlexible hours$140k - $182k per year
...for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable... ..., Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and software...Full time$141k - $191k per year
...at Thomson Reuters and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence... .... About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for:...Work at officeLocal areaFlexible hours2 days per week3 days per week- ...Artificial Intelligence. Actual Impact. At Docebo, we’re using AI to change how people learn at work—and we mean actually... ...learning never stops. The Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team of engineers dedicated...Long term contractFull timeFor contractorsWork at officeWorldwide3 days per week
$140k - $180k per year
...visit: The role Our SRE team runs the shared infrastructure that all other engineering teams at Tripstack depend on.... ...Prometheus, Thanos, and Grafana observability stack. We are in the middle... ...the deployment patterns and reliability standards that apply across both...RemplacementFull timeWork at officeImmediate startFlexible hours- ...Artificial Intelligence. Actual Impact. At Docebo, we’re using AI to change how people learn at work—and we mean actually... ...because learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident response...Full timeFor contractorsWork at officeWorldwide3 days per week
$154k - $200k per year
...for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable... ...America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic...Long term contractFull time$100k - $125k per year
...Position Summary: We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer, you will play a crucial role in enhancing the reliability, performance, and scalability of our systems...Full timeWork at officeFlexible hours$153.82k - $277k per year
...a place where you can thrive, we can’t wait to meet you. Site Reliability Engineers (SREs) at Braze are responsible for keeping all internal-facing... ...messages to our customers' end-users daily. This Senior SRE role is specifically focused on supporting the Braze Ruby on...Permanent employmentFull timeInternshipWork at officeLocal areaRemote workFlexible hoursRotating shift- ...SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe... ...belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a...Full timeWork at officeLocal areaRemote workWorldwideMonday to fridayFlexible hours
$80 - $110 per hour
...AXON Networks delivers a robust AI-driven, analytics-based orchestration platform and a wide portfolio of next-gen high... ...Singapore and also operating in Denmark, Spain and Vietnam. The Site Reliability Engineer will improve the availability, performance, scalability and...Remote jobFull timeContract work$92k - $118k per year
...Build reliable, resilient cloud platforms that keep critical financial services running. The Role We are looking for an experienced Site Reliability and DevOps Engineer to join the SRE team within our growing Corporate Action Processing group. You will bring 3+ years...InternshipImmediate start$144k - $200k per year
**The Team** Platform Engineering is the department within SRE that is responsible for a range of critical infrastructure... ..., deployment machinery, and observability and alerting systems. The Fabric... ...in developing and maintaining the reliable and globally connected multi-cloud...Full timeWork at officeRemote workWorldwideFlexible hours$110k - $125k per year
...systems bringing #BetterGlobalHealth to patients everyday! Apply today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across...Full timeRemote workFlexible hours- ...About The Role Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work. Based in Toronto or remote...Full timeRemote work
- ...work Job Description We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability,... ...improve scalability and resilience. Define and evolve observability practices, including dashboards, alerts, SLOs, and SLIs....Permanent employmentWork at officeLocal areaRemote work
- ...and enterprises who are building AI systems to power magical... ...Cohere is a team of researchers, engineers, designers, and more, who are passionate... ...high-performance, scalable and reliable machine learning systems? Do... ...? We are looking for a Site Reliability Engineer to join the...Full timeWork at officeRemote workFlexible hours
- ...strong in a few areas, and have some interest and capabilities in others. About the Role: As a Site Reliability Engineer , you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS platform that powers the...Full time
- ...youll be doing As a member of CIBCs Application Reliability Engineering Platform team the Consultant Site Reliability Engineering will play a key role in... ...scalability and efficiency. Define and maintain observability strategies to ensure visibility into application behaviour...Full timeContract work3 days per week1 day per week
- ...is part of the Production Support and Reliability Engineering team, responsible for ensuring the stability... ...Support, Production Support, SRE, or Platform Operations roles. Strong... ...reliability. Knowledge of monitoring and observability tools for proactive system management....Permanent employmentFull timeLocal area
$140k - $165k per year
...leader driving a major IT transformation, transitioning from traditional service desk operations to an AI-driven, self-healing operational fabric. As Lead Site Reliability Engineer (SRE), you will lead enterprise observability, AIOps automation, and SRE governance across a...- ...LinkedIn profiles! This is a hands-on senior engineering role focused on improving production... ...engineering teams to ship secure, reliable, and scalable software with confidence.... ...service level objectives (SLOs), improve observability, automate operational processes, and drive...Permanent employmentFull timeRemote work
$78.62 - $89.34 per hour
Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering and Operations team. In this engineering-focused role, you will move beyond traditional database administration to act as...Long term contractPermanent employmentFull timeContract workWork at office$170k - $185k per year
...managers? We are seeking a Head of Platform Engineering, Reliability & Control to lead our horizontal... ...role, you will shape a more connected, observable, and automated technology ecosystem spanning... ...for command-and-control capabilities, SRE practices, process orchestration,...Full timeWork at office3 days per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - Dynatrace & AI Observability. Be the first to apply!
- senior site reliability engineer Toronto, ON
- site reliability engineer Toronto, ON
- site reliability engineer intern Toronto, ON
- site safety Toronto, ON
- site maintenance Toronto, ON
- site carpenter Toronto, ON
- website developer Toronto, ON
- senior site reliability engineer
- site reliability engineer sre
- site reliability engineer

