Site Reliability Engineer (SRE) - Dynatrace & AI Observability
Astra North Infoteck Inc.
Site Reliability Engineer
Location: Toronto ON
Work Model: Hybrid (2 days per week in-person at the Toronto office preferred)
Required Skills
Site Reliability Engineering (SRE)
DevOps
Dynatrace
Role Summary
Design implement and optimize Site Reliability Engineering (SRE) and DevOps practices to ensure high system availability performance and reliability across distributed environments.
Leverage Dynatrace Davis AI automation and cloud technologies to enable proactive monitoring intelligent automation and operational excellence.
Role Description
Dynatrace & AI-Driven Observability
Lead the implementation and optimization of the Dynatrace platform across applications and infrastructure.
Utilize Dynatrace Davis AI for automated root cause analysis anomaly detection event correlation predictive performance insights and alert noise reduction.
Configure and manage OneAgent deployments Smartscape topology mapping service flow and distributed tracing.
Define and monitor SLIs SLOs and user experience metrics.
Build custom dashboards alerts and observability pipelines.
Integrate Dynatrace with CI/CD pipelines for release validation and performance gating.
Integrate Dynatrace with incident management tools such as PagerDuty and ServiceNow.
Enable self-healing automation using Dynatrace event triggers and AI-driven insights.
Automation & Configuration Management
Design and implement automation solutions using Ansible.
Automate configuration management application deployments and environment provisioning.
Develop reusable Ansible playbooks and roles for scalable operations.
Automate operational tasks patching compliance processes and remediation workflows.
Integrate Ansible with CI/CD pipelines and monitoring systems.
Cloud & DevOps
Design and manage cloud-native solutions on AWS with exposure to Azure.
Develop infrastructure using Terraform CloudFormation or AWS CDK.
Build and manage CI/CD pipelines using GitHub Actions Jenkins or GitLab CI.
Develop and deploy serverless solutions using AWS Lambda API Gateway and Step Functions.
Automate DevOps and operational workflows using Python (boto3) and Bash scripting.
Deploy and maintain production environments through automated pipelines.
Optimize cloud infrastructure for cost performance and scalability.
Monitoring & Reliability Engineering
Monitor and manage AWS CloudWatch and Azure Monitor/Log Analytics.
Design unified observability across multi-cloud environments.
Implement logging and distributed tracing strategies.
Work with Docker Kubernetes ECS and AKS environments.
Design fault-tolerant highly available and disaster recovery solutions.
Support incident response on-call activities and root cause analysis (RCA).
Required Qualifications
Proven experience with Dynatrace APM Real User Monitoring (RUM) and infrastructure monitoring.
Strong hands-on experience with Dynatrace Davis AI capabilities.
Experience with Ansible for automation and configuration management.
Deep knowledge of AWS services and cloud-native architectures.
Experience with Infrastructure as Code using Terraform CloudFormation or AWS CDK.
Proficiency in Python (boto3) and Bash scripting.
Experience supporting production-scale environments.
Business Analyst experience.
Scrum Master experience.
Nice to Have
Dynatrace Associate or Professional certification.
Experience with Dynatrace APIs and automation.
Experience building self-healing systems using AI-driven triggers.
Familiarity with Prometheus Grafana and the ELK Stack.
Azure cloud experience and certifications.
Experience with GitOps and Platform Engineering.
Required Skills:
Top 3 Required Skills: 1. IBM Financial transaction 2. Payment flow 3. Support Modernization Detailed Job Description: Design develop and maintain applications built on IBM Financial Transaction Manager (FTM) to support core payments processing. Contribute to the development of payment flows supporting transaction processing. Build and support integrations between FTM and upstream/downstream systems using enterprise integration patterns. Participate in the design development testing deployment and production support. Troubleshoot and resolve application and integration issues in a complex regulated environment. Collaborate with architecture QA and operations teams to ensure platform stability scalability and performance. Support modernization initiatives and enhancements to existing payment hub capabilities. Produce clear technical documentation and participate in code reviews and knowledge sharing.
- ...Platform Engineer – DevOps, Site Reliability Engineering (SRE) & Dynatrace Required Skills • Strong experience as a Platform Engineer with expertise in DevOps and Site Reliability Engineering (SRE). • Experience designing, implementing, automating, and supporting...SuggestedPermanent employment
- ...Site Reliability Engineer – APM, Dynatrace, Observability Location • Toronto, ON • Hybrid – 2 days in office per week Required Skills • Strong Day 1 expertise in Dynatrace, including: • DQL • Gen3 Dashboards • Traces / Grail...SuggestedContract workWork at office2 days per week
- ...Design, deploy, and manage enterprise observability platforms across Kubernetes... ...management tools. Collaborate with SRE, DevOps, and application teams to improve platform reliability and reduce MTTR. Preferred Skills AI/ML for Observability, AIOps, LLM integration...SuggestedContract work
- ...Job Role: Senior Platform Engineer – DevOps, SRE & Dynatrace Observability Location : Toronto - Hybrid (3 Days Work from Office) Job Summary We are seeking a highly skilled and experienced Platform Engineer with strong DevOps and Site Reliability Engineering (...SuggestedContract workWork at office
$130k - $180k per year
...legally work in Canada (visa or sponsorship won't be provided) Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure (those with Azure will be prioritized first) About Us: We're...SuggestedFull timeRemote workVisa sponsorshipWork visaFlexible hours$141k - $191k per year
...at Thomson Reuters and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence... .... About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for:...Work at officeLocal areaFlexible hours2 days per week3 days per week$110k - $120k per year
...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for... ...UAT, production), ensuring changes are safe, repeatable, and observable. Design and maintain automated CI/CD pipelines and enforce...Temporary workInternshipWork at officeRemote work$78.62 - $89.34 per hour
Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering and Operations team. In this engineering-focused role, you will move beyond traditional database administration to act as...Long term contractPermanent employmentFull timeContract workWork at office- ...Enterprise Kubernetes SRE (Python, GitOps, API, Container, Cloud,, MongoDB, Postgres) Toronto, ON - Hybrid (4 Days WFO) 12 months We are seeking an experienced Site Reliability Engineer to join our Enterprise Kubernetes Platform team at a leading financial...Contract work
$164.6k - $235.1k per year
...About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset and toolkit to the challenges of building and running large-scale, distributed systems. Our mission...Long term contractRemplacementFull timeContract workTemporary workLocal areaFlexible hours- ...Waabi, founded by AI visionary Raquel Urtasun, is the leader in Physical AI. With a... ...and development of Waabi’s monitoring and observability stack, used to monitor the health and performance... ...Qualifications: - 5+ years software engineering or systems/performance engineering...Full time
- ...SRE is part of a global organization that leverages the latest technology to communicate with our colleagues across the globe... ...belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a...Full timeWork at officeLocal areaRemote workWorldwideMonday to fridayFlexible hours
$78.62 - $89.34 per hour
Our client, is seeking a talented and proactive Site Reliability Specialist / Senior Cloud Platform Engineer to join their core Cloud Engineering division. In this engineering-focused role, you will move far beyond basic operational support to act as a principal architect...Long term contractPermanent employmentContract workWork at office- ...Observability Engineer – Kubernetes, Prometheus, Grafana & Cloud Monitoring Role Overview • Experienced Observability Engineer to join... ...self-healing infrastructure using modern observability tools and AI/ML capabilities Key Responsibilities • Design,...Long term contractPermanent employment
$95k - $145k per year
GCP Cloud Engineering Consultant (Dynatrace) Position Description CGI is seeking a highly skilled Dynatrace... ..., and optimizing Dynatrace-based observability across cloud-native applications,... ...requires deep expertise in Dynatrace, Site Reliability Engineering (SRE) practices...Work at office3 days per week$135k - $170k per year
...General Information: Job Title: AI Engineer Location: Toronto, ON (Onsite/Hybrid)... ...Optimize systems for cost, latency, and reliability Collaborate across teams where needed... ...experience (vLLM, TGI, llama.cpp) Observability tooling (Langfuse, LangSmith) Prior...Long term contractFull timeInternshipWork at officeImmediate startRemote workFlexible hours$130k - $165k per year
...What We Need We’re looking for a Senior AI Engineer to design and build production-grade... ...and RAG systems that power intelligent, reliable automation across our platform. This role... ...architecture and evaluation to scalability, observability, and reliability in production. The...Work at office1 day per week- ...Performance Test Engineer – LoadRunner, Performance Testing, Dynatrace & Splunk Required Skills: • Minimum 6+ years of experience in Performance Testing. • Proficiency in LoadRunner with hands-on scripting experience. • Experience leading end-to-end Performance...Permanent employment
- ...Big Viking Games is hiring a Head of Engineering & AI to drive AI-first transformation across our... ...live game services Ensure live-service reliability, security posture, anti-cheat and fraud... ...practices Elevate DevOps and SRE maturity including build pipelines, deployment...Long term contractFull timeShift work
$80k - $180k per year
...Zafin is an AI platform company helping regulated institutions modernize... ...? The AI Evaluation Engineer ensures AI agent solutions are accurate, reliable, safe, and production-ready within... ...CI/CD, automated evaluation, AI observability, and engineering delivery practices...Full time$135k - $210k per year
...Overview: Guidepoint seeks an experienced AI Engineer as an integral member of the Toronto-... ...like LLM Judges or MLflow, AI observability, and system monitoring. Evaluate and... ...# Live Coding & System Design Test (On-site, 2 hours) # Technical Leadership Interview...Full timeWorldwide$250k per year
...Role: Observability Engineer – Trading Client: Elite FinTech Compensation: $120,000 - $250,000 CAD + Bonus Location: Toronto... ...Working with multiple technical teams to ensure visibility and reliability across systems. Key Responsibilities Monitoring Tools...Permanent employmentImmediate start$160k - $220k per year
...Secure Every Identity, from AI to Human Identity is the key to unlocking the... ...you are too, let's talk. Staff Software Reliability Engineer - Data Platform About the Team The... ...in areas such as data quality, data observability and incident management Nice to have...Local areaWorldwideFlexible hours- ...organization, apply now. We are currently seeking a Agentic AI Engineer to join our team in toronto, Ontario (CA-ON), Canada (CA). The... ...Whenever possible, we hire locally to NTT DATA offices or client sites. This ensures we can provide timely and effective support...Work at officeRemote workFlexible hours
$72k - $138k per year
...and on the job coaching -- We are looking for a passionate AI Research Engineer to join our team. You will work at the intersection of... ...evaluating, and deploying generative AI (GenAI) systems that are both reliable and impactful. This role blends fundamental research, model...Permanent employmentFlexible hours- ...that makes this real: a unified AI agent that sells, supports,... ...team move faster and more reliably. About the Role You’ll... ...platforms and tooling used by AI and engineering teams Reduce manual... ...LLMs and agents. Continuous Observability: Take ownership of the...Long term contractRemplacementFull timeInternshipWork at officeWork from homeShift work
$140.6k - $190.6k per year
...The Lead AI Forward Engineer designs and guides the delivery of AI-powered solutions that reduce... ...to production, ensuring solutions meet reliability, security, and compliance expectations... ...requirements. Integrate AI observability tooling into CI/CD so new models, prompts...Full timeManual laborWork at officeLocal areaFlexible hours2 days per week3 days per week- ...Job Title: Performance Engineer – LoadRunner, JMeter, Dynatrace (Payments Domain) Work Location: Toronto, ON – Hybrid (4 Days WFO) Required Experience: • 10+ years of hands-on experience with LoadRunner and JMeter. • Hands-on experience in Resiliency...Permanent employment
$135k - $150k per year
...The Opportunity: We’re looking for an AI Engineer, AI Platform & ML Engineering to join... ...solutions to become repeatable, governed, observable, and production ready. As one aviso... ...Help monitor AI cost, performance, reliability, usage, and operational risk, while contributing...Full timeInternship- ...Continuum Continuum is an enterprise AI agent execution and control... ...help teams build, run, and deploy reliable agentic applications. It provides agent... ...governance, guardrails, evaluation, and observability. We are looking for an AI Engineer Intern interested in building...Internship
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - Dynatrace & AI Observability. Be the first to apply!
- senior site reliability engineer Toronto, ON
- site reliability engineer intern Toronto, ON
- site reliability engineer Toronto, ON
- site reliability engineer remote Toronto, ON
- site safety Toronto, ON
- website developer Toronto, ON
- site maintenance Toronto, ON
- senior site reliability engineer
- site reliability engineer sre
- site reliability engineer intern
