Site Reliability Engineer (SRE) - Dynatrace & AI Observability
Astra North Infoteck Inc.
Site Reliability Engineer
Location: Toronto ON
Work Model: Hybrid (2 days per week in-person at the Toronto office preferred)
Required Skills
Site Reliability Engineering (SRE)
DevOps
Dynatrace
Role Summary
Design implement and optimize Site Reliability Engineering (SRE) and DevOps practices to ensure high system availability performance and reliability across distributed environments.
Leverage Dynatrace Davis AI automation and cloud technologies to enable proactive monitoring intelligent automation and operational excellence.
Role Description
Dynatrace & AI-Driven Observability
Lead the implementation and optimization of the Dynatrace platform across applications and infrastructure.
Utilize Dynatrace Davis AI for automated root cause analysis anomaly detection event correlation predictive performance insights and alert noise reduction.
Configure and manage OneAgent deployments Smartscape topology mapping service flow and distributed tracing.
Define and monitor SLIs SLOs and user experience metrics.
Build custom dashboards alerts and observability pipelines.
Integrate Dynatrace with CI/CD pipelines for release validation and performance gating.
Integrate Dynatrace with incident management tools such as PagerDuty and ServiceNow.
Enable self-healing automation using Dynatrace event triggers and AI-driven insights.
Automation & Configuration Management
Design and implement automation solutions using Ansible.
Automate configuration management application deployments and environment provisioning.
Develop reusable Ansible playbooks and roles for scalable operations.
Automate operational tasks patching compliance processes and remediation workflows.
Integrate Ansible with CI/CD pipelines and monitoring systems.
Cloud & DevOps
Design and manage cloud-native solutions on AWS with exposure to Azure.
Develop infrastructure using Terraform CloudFormation or AWS CDK.
Build and manage CI/CD pipelines using GitHub Actions Jenkins or GitLab CI.
Develop and deploy serverless solutions using AWS Lambda API Gateway and Step Functions.
Automate DevOps and operational workflows using Python (boto3) and Bash scripting.
Deploy and maintain production environments through automated pipelines.
Optimize cloud infrastructure for cost performance and scalability.
Monitoring & Reliability Engineering
Monitor and manage AWS CloudWatch and Azure Monitor/Log Analytics.
Design unified observability across multi-cloud environments.
Implement logging and distributed tracing strategies.
Work with Docker Kubernetes ECS and AKS environments.
Design fault-tolerant highly available and disaster recovery solutions.
Support incident response on-call activities and root cause analysis (RCA).
Required Qualifications
Proven experience with Dynatrace APM Real User Monitoring (RUM) and infrastructure monitoring.
Strong hands-on experience with Dynatrace Davis AI capabilities.
Experience with Ansible for automation and configuration management.
Deep knowledge of AWS services and cloud-native architectures.
Experience with Infrastructure as Code using Terraform CloudFormation or AWS CDK.
Proficiency in Python (boto3) and Bash scripting.
Experience supporting production-scale environments.
Business Analyst experience.
Scrum Master experience.
Nice to Have
Dynatrace Associate or Professional certification.
Experience with Dynatrace APIs and automation.
Experience building self-healing systems using AI-driven triggers.
Familiarity with Prometheus Grafana and the ELK Stack.
Azure cloud experience and certifications.
Experience with GitOps and Platform Engineering.
Required Skills:
Top 3 Required Skills: 1. IBM Financial transaction 2. Payment flow 3. Support Modernization Detailed Job Description: Design develop and maintain applications built on IBM Financial Transaction Manager (FTM) to support core payments processing. Contribute to the development of payment flows supporting transaction processing. Build and support integrations between FTM and upstream/downstream systems using enterprise integration patterns. Participate in the design development testing deployment and production support. Troubleshoot and resolve application and integration issues in a complex regulated environment. Collaborate with architecture QA and operations teams to ensure platform stability scalability and performance. Support modernization initiatives and enhancements to existing payment hub capabilities. Produce clear technical documentation and participate in code reviews and knowledge sharing.
- ...Platform Engineer – DevOps, Site Reliability Engineering (SRE) & Dynatrace Required Skills • Strong experience as a Platform Engineer with expertise in DevOps and Site Reliability Engineering (SRE). • Experience designing, implementing, automating, and supporting...SuggestedPermanent employment
- ...Job Role: Senior Platform Engineer DevOps SRE & Dynatrace Observability Location : Toronto - Hybrid (3 Days Work from Office) Job Summary We are seeking a highly skilled and experienced Platform Engineer with strong DevOps and Site Reliability Engineering (SRE...SuggestedFull timeWork at office
$141k - $191k per year
...at Thomson Reuters and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence... .... About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for:...SuggestedWork at officeLocal areaFlexible hours2 days per week3 days per week- ...work Job Description We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability,... ...improve scalability and resilience. Define and evolve observability practices, including dashboards, alerts, SLOs, and SLIs....SuggestedPermanent employmentWork at officeLocal areaRemote work
$100k - $125k per year
...Position Summary: We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer, you will play a crucial role in enhancing the reliability, performance, and scalability of our systems...SuggestedFull timeWork at officeFlexible hours$153.82k - $277k per year
...like a place where you can thrive we cant wait to meet you. Site Reliability Engineers (SREs) at Braze are responsible for keeping all internal-... ...of messages to our customers end-users daily. This Senior SRE role is specifically focused on supporting the Braze Ruby on...Permanent employmentFull timeInternshipWork at officeLocal areaRemote workFlexible hoursRotating shift- ...About The Role Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work. Based in Toronto or remote...Full timeRemote work
- ...is part of the Production Support and Reliability Engineering team, responsible for ensuring the stability... ...Support, Production Support, SRE, or Platform Operations roles. Strong... ...reliability. Knowledge of monitoring and observability tools for proactive system management....Permanent employmentFull timeLocal area
- ...youll be doing As a member of CIBCs Application Reliability Engineering Platform team the Consultant Site Reliability Engineering will play a key role in... ...scalability and efficiency. Define and maintain observability strategies to ensure visibility into application behaviour...Full timeContract work3 days per week1 day per week
$140k - $165k per year
...leader driving a major IT transformation, transitioning from traditional service desk operations to an AI-driven, self-healing operational fabric. As Lead Site Reliability Engineer (SRE), you will lead enterprise observability, AIOps automation, and SRE governance across a...$110k - $120k per year
...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for... ...UAT, production), ensuring changes are safe, repeatable, and observable. Design and maintain automated CI/CD pipelines and enforce...Temporary workInternshipWork at officeRemote work$78.62 - $89.34 per hour
Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering and Operations team. In this engineering-focused role, you will move beyond traditional database administration to act as...Long term contractPermanent employmentFull timeContract workWork at office$115k - $125k per year
...symbol: KGC). Job Summary The Reliability Engineer, as part of the Asset Management team, plays... ...position requires 50%+ travel to remote sites overseas, ensuring that global... ...performance, and long-term development. Use of AI in Our Hiring Process We use AI-...Long term contractTemporary workFor contractorsCasual workLocal areaImmediate startRemote workOverseas$108k - $135k per year
...ideas for the benefit of others. As an Observability team member, you are responsible for the... ...distributed systems. We count on the reliability of our infrastructure to empower Lyft... ...are seeking experienced Infrastructure Engineer to ensure that as our Infrastructure continues...Hourly payWork at officeFlexible hours3 days per week- ...Number : R2868511 Position title : AI Engineer Department: Commercial Data Science... ...across our markets, R&D and manufacturing sites. The team is located in major hubs in... ...improve AI application performance and reliability Develop and maintain production-ready...Work at officeWork from homeHome officeFlexible hours
$130k - $165k per year
...What We Need We’re looking for a Senior AI Engineer to design and build production-grade... ...and RAG systems that power intelligent, reliable automation across our platform. This role... ...architecture and evaluation to scalability, observability, and reliability in production. The...Work at office1 day per week$135k - $170k per year
...General Information: Job Title: AI Engineer Location: Toronto, ON (Onsite/Hybrid)... ...Optimize systems for cost, latency, and reliability Collaborate across teams where needed... ...experience (vLLM, TGI, llama.cpp) Observability tooling (Langfuse, LangSmith) Prior...Long term contractFull timeInternshipWork at officeImmediate startRemote workFlexible hours- ...Palona’s AI agents operate in real restaurant environments: noisy... ...for an applied AI Modeling Engineer to improve the intelligence, accuracy... ...hypothesis to experiment to reliable deployment. What you’ll... ...work is reproducible, tested, observable, and usable by other engineers...Long term contractFull timeTemporary workInternshipImmediate start
- ...Palona’s AI agents operate continuously in production, handle real-time guest... ...part of the product: latency, reliability, deployment safety, observability, security, and cost directly shape the... ...We are looking for an Infrastructure Engineer who combines cloud and reliability depth...Long term contractFull timeTemporary workInternship
- ...is leading the industry on cutting-edge AI technology, revolutionizing performance... ...seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy... ...at manufacturing and test partner sites, including in Taiwan. Tenstorrent...Permanent employmentInternshipSecond job
$250k per year
...Role: Observability Engineer – Trading Client: Elite FinTech Compensation: $120,000 - $250,000 CAD + Bonus Location: Toronto... ...Working with multiple technical teams to ensure visibility and reliability across systems. Key Responsibilities Monitoring Tools...Permanent employmentImmediate start$140.6k - $190.6k per year
...The Lead AI Forward Engineer designs and guides the delivery of AI-powered solutions that reduce... ...to production, ensuring solutions meet reliability, security, and compliance expectations... ...requirements. Integrate AI observability tooling into CI/CD so new models, prompts...Full timeManual laborWork at officeLocal areaFlexible hours$85 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Jack Dorsey . Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience) Type:... ...platforms , Kubernetes , CI/CD systems , observability , and infrastructure automation ....Remote jobContract workSummer work$72k - $138k per year
...and on the job coaching -- We are looking for a passionate AI Research Engineer to join our team. You will work at the intersection of... ...evaluating, and deploying generative AI (GenAI) systems that are both reliable and impactful. This role blends fundamental research, model...Permanent employmentFlexible hours- ...We are looking for a product-focused AI Software Engineer to turn advances in AI into restaurant products that work reliably in the real world. You will build across customer... ...systems. ~ A strong quality bar for testing, observability, security, reliability, and user experience...Long term contractFull timeTemporary work
$100k - $150k per year
...growth, with a focus on practical AI adoption, stronger internal... ...is hiring a Senior Full Stack Engineer to build AI-enabled products,... ...complex production tasks into reliable, reviewable steps. · Systems... ...Improve system reliability, observability, documentation, maintainability...Long term contractFull timeInternshipWork at office3 days per week$80k - $130k per year
DevOps SRE Consultant Position Description The DevOps Site Reliability Engineering consultant is responsible for environment... .... This role is part of an AI first software engineering... ...avec le soutien des plateformes, l’observabilité et l’automatisation afin d’aider...$172k - $229k per year
...looking for a Staff Software Engineer to set and drive our company-wide... ..., shipping, and operating reliable software. This is a high-impact... ...systems, validate changes, observe production behavior, manage risk... ...to equipping contractors with AI-driven tools to conquer chaos,...Long term contractPermanent employmentFor contractorsWork at officeLocal areaWork from homeFlexible hours- ...intersection of product and platform engineers, financial partners, and AI systems—connecting them to ensure... ...Stripe’s open banking ecosystem remain reliable at scale. What you'll do... ...improve anomaly detection systems and observability tooling to ensure data flowing...
$80 - $120 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Summers , and Jack Dorsey . Position: Incident management / reliability / SRE Evaluator Type: Contract Compensation: $80–$...Remote jobContract workSummer workWork at office
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - Dynatrace & AI Observability. Be the first to apply!
- senior site reliability engineer Toronto, ON
- site reliability engineer remote Toronto, ON
- site reliability engineer intern Toronto, ON
- site reliability engineer Toronto, ON
- site carpenter Toronto, ON
- website developer Toronto, ON
- site safety Toronto, ON
- site maintenance Toronto, ON
- senior site reliability engineer
- site reliability engineer remote

