Site Reliability Engineer (SRE) - Dynatrace & AI Observability
Astra North Infoteck Inc.
Site Reliability Engineer
Location: Toronto, ON
Work Model: Hybrid (2 days per week in-person at the Toronto office preferred)
Required Skills
• Site Reliability Engineering (SRE)
• DevOps
• Dynatrace
Role Summary
• Design, implement, and optimize Site Reliability Engineering (SRE) and DevOps practices to ensure high system availability, performance, and reliability across distributed environments.
• Leverage Dynatrace, Davis AI, automation, and cloud technologies to enable proactive monitoring, intelligent automation, and operational excellence.
Role Description
Dynatrace & AI-Driven Observability
• Lead the implementation and optimization of the Dynatrace platform across applications and infrastructure.
• Utilize Dynatrace Davis AI for automated root cause analysis, anomaly detection, event correlation, predictive performance insights, and alert noise reduction.
• Configure and manage OneAgent deployments, Smartscape topology mapping, service flow, and distributed tracing.
• Define and monitor SLIs, SLOs, and user experience metrics.
• Build custom dashboards, alerts, and observability pipelines.
• Integrate Dynatrace with CI/CD pipelines for release validation and performance gating.
• Integrate Dynatrace with incident management tools such as PagerDuty and ServiceNow.
• Enable self-healing automation using Dynatrace event triggers and AI-driven insights.
Automation & Configuration Management
• Design and implement automation solutions using Ansible.
• Automate configuration management, application deployments, and environment provisioning.
• Develop reusable Ansible playbooks and roles for scalable operations.
• Automate operational tasks, patching, compliance processes, and remediation workflows.
• Integrate Ansible with CI/CD pipelines and monitoring systems.
Cloud & DevOps
• Design and manage cloud-native solutions on AWS with exposure to Azure.
• Develop infrastructure using Terraform, CloudFormation, or AWS CDK.
• Build and manage CI/CD pipelines using GitHub Actions, Jenkins, or GitLab CI.
• Develop and deploy serverless solutions using AWS Lambda, API Gateway, and Step Functions.
• Automate DevOps and operational workflows using Python (boto3) and Bash scripting.
• Deploy and maintain production environments through automated pipelines.
• Optimize cloud infrastructure for cost, performance, and scalability.
Monitoring & Reliability Engineering
• Monitor and manage AWS CloudWatch and Azure Monitor/Log Analytics.
• Design unified observability across multi-cloud environments.
• Implement logging and distributed tracing strategies.
• Work with Docker, Kubernetes, ECS, and AKS environments.
• Design fault-tolerant, highly available, and disaster recovery solutions.
• Support incident response, on-call activities, and root cause analysis (RCA).
Required Qualifications
• Proven experience with Dynatrace APM, Real User Monitoring (RUM), and infrastructure monitoring.
• Strong hands-on experience with Dynatrace Davis AI capabilities.
• Experience with Ansible for automation and configuration management.
• Deep knowledge of AWS services and cloud-native architectures.
• Experience with Infrastructure as Code using Terraform, CloudFormation, or AWS CDK.
• Proficiency in Python (boto3) and Bash scripting.
• Experience supporting production-scale environments.
• Business Analyst experience.
• Scrum Master experience.
Nice to Have
• Dynatrace Associate or Professional certification.
• Experience with Dynatrace APIs and automation.
• Experience building self-healing systems using AI-driven triggers.
• Familiarity with Prometheus, Grafana, and the ELK Stack.
• Azure cloud experience and certifications.
• Experience with GitOps and Platform Engineering.
- ...Site Reliability Engineer Location: Toronto ON Work Model: Hybrid (2 days per week in-person at the Toronto office preferred) Required Skills Site Reliability Engineering (SRE) DevOps Dynatrace Role Summary Design implement and optimize...SuggestedFull timeWork at office2 days per week
- ...Platform Engineer – DevOps, Site Reliability Engineering (SRE) & Dynatrace Required Skills • Strong experience as a Platform Engineer with expertise in DevOps and Site Reliability Engineering (SRE). • Experience designing, implementing, automating, and supporting...SuggestedPermanent employment
- ...Job Role: Senior Platform Engineer DevOps SRE & Dynatrace Observability Location : Toronto - Hybrid (3 Days Work from Office) Job Summary We are seeking a highly skilled and experienced Platform Engineer with strong DevOps and Site Reliability Engineering (SRE...SuggestedFull timeWork at office
$141k - $191k per year
...at Thomson Reuters and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence... .... About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for:...SuggestedWork at officeLocal areaFlexible hours2 days per week3 days per week- ...is part of the Production Support and Reliability Engineering team, responsible for ensuring the stability... ...Support, Production Support, SRE, or Platform Operations roles. Strong... ...reliability. Knowledge of monitoring and observability tools for proactive system management....SuggestedPermanent employmentFull timeLocal area
$140k - $180k per year
...visit: The role Our SRE team runs the shared infrastructure that all other engineering teams at Tripstack depend on.... ...Prometheus, Thanos, and Grafana observability stack. We are in the middle... ...the deployment patterns and reliability standards that apply across both...RemplacementFull timeWork at officeImmediate startFlexible hours$100k - $125k per year
...Position Summary We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer, you will play a crucial role in enhancing the reliability, performance, and scalability of our systems...Full timeWork at officeFlexible hours$104.24k - $143.3k per year
...methodologies at an innovative and growing company. Our Lead Site Reliability Engineers will provide a stable infrastructure platform throughout our... ...or a migration of an existing client to our Azure cloud. The SRE will be responsible for aligning client requirements to...Full timeWork at officeImmediate startRemote workShift workWeekend work2 days per week$110k - $120k per year
...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for... ...UAT, production), ensuring changes are safe, repeatable, and observable. Design and maintain automated CI/CD pipelines and enforce...Temporary workInternshipWork at officeRemote work- ...Senior Site Reliability Engineer - Edge Location : Ottawa/Toronto, On-Site Reports to: Head of Security The Role We build rugged edge... ...Additional equity granted based on impact We use AI tools to support parts of the hiring process, including screening...Full time
$110k - $151.8k per year
...Secure Every Identity, from AI to Human Identity is the key to... ...s talk. The Auth0 Platform Observability team owns the observability tooling... ...looking for an Observability Engineer to help ensure that our... ...you have experience within the Site Reliability Engineering (SRE)...Full timeLocal areaWorldwide$80k - $130k per year
DevOps SRE Position Description The DevOps Site Reliability Engineering consultant is responsible for environment support, reliability, and operational stability across development, test, and pre production environments. This role combines hands on environment configuration...- ...Observability Engineer – Kubernetes, Prometheus, Grafana & Cloud Monitoring Role Overview • Experienced Observability Engineer to join... ...self-healing infrastructure using modern observability tools and AI/ML capabilities Key Responsibilities • Design,...Long term contractPermanent employment
$78.62 - $89.34 per hour
Our client, is seeking a talented and proactive Site Reliability Specialist / Senior Cloud Platform Engineer to join their core Cloud Engineering division. In this engineering-focused role, you will move far beyond basic operational support to act as a principal architect...Long term contractPermanent employmentContract workWork at office- ...Palona’s AI agents operate in real restaurant environments: noisy... ...for an applied AI Modeling Engineer to improve the intelligence, accuracy... ...hypothesis to experiment to reliable deployment. What you’ll... ...work is reproducible, tested, observable, and usable by other engineers...Long term contractFull timeTemporary workInternshipImmediate start
- ...is leading the industry on cutting-edge AI technology, revolutionizing performance... ...seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy... ...at manufacturing and test partner sites, including in Taiwan. Tenstorrent...Permanent employmentInternshipSecond job
- ...Palona’s AI agents operate continuously in production, handle real-time guest... ...part of the product: latency, reliability, deployment safety, observability, security, and cost directly shape the... ...We are looking for an Infrastructure Engineer who combines cloud and reliability depth...Long term contractFull timeTemporary workInternship
$135k - $170k per year
...General Information: Job Title: Applied AI Engineer Location: Toronto, ON (Onsite/Hybrid... ...systems for cost, latency, and reliability Collaborate across teams where needed... ...experience (vLLM, TGI, llama.cpp) Observability tooling (Langfuse, LangSmith) Prior...Long term contractFull timeInternshipWork at officeImmediate startRemote workFlexible hours$140.6k - $190.6k per year
...The Lead AI Forward Engineer designs and guides the delivery of AI-powered solutions that reduce... ...to production, ensuring solutions meet reliability, security, and compliance expectations... ...requirements. Integrate AI observability tooling into CI/CD so new models, prompts...Full timeManual laborWork at officeLocal areaFlexible hours- ...Manager: Senior Director, Data Engineering Contract Terms: Permanent, 37... ...dedicated engineers , providing reliable and consistent data to the stakeholder... ...platform operation effort as a Site Reliability Developer! Among many other duties, the SRE team within Ticketmaster's TM1...Permanent employmentFull timeContract workLocal areaRemote workFlexible hours
- ...We are looking for a product-focused AI Software Engineer to turn advances in AI into restaurant products that work reliably in the real world. You will build across customer... ...systems. ~ A strong quality bar for testing, observability, security, reliability, and user experience...Long term contractFull timeTemporary work
$85 per hour
...creative and technical talent with leading AI research labs. Headquartered in San... ...Jack Dorsey . Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience) Type:... ...platforms , Kubernetes , CI/CD systems , observability , and infrastructure automation ....Remote jobContract workSummer work$250k per year
...Role: Observability Engineer – Trading Client: Elite FinTech Compensation: $120,000 - $250,000 CAD + Bonus Location: Toronto... ...Working with multiple technical teams to ensure visibility and reliability across systems. Key Responsibilities Monitoring Tools...Permanent employmentImmediate start$72k - $138k per year
...and on the job coaching -- We are looking for a passionate AI Research Engineer to join our team. You will work at the intersection of... ...evaluating, and deploying generative AI (GenAI) systems that are both reliable and impactful. This role blends fundamental research, model...Permanent employmentFlexible hours- ...Performance Test Engineer – LoadRunner, Performance Testing, Dynatrace & Splunk Required Skills: • Minimum 6+ years of experience in Performance Testing. • Proficiency in LoadRunner with hands-on scripting experience. • Experience leading end-to-end Performance...
$80k - $180k per year
...Zafin is an AI platform company helping regulated institutions modernize... ...? The AI Evaluation Engineer ensures AI agent solutions are accurate, reliable, safe, and production-ready within... ...CI/CD, automated evaluation, AI observability, and engineering delivery practices...Full time$80k - $190k per year
...Zafin is an AI platform company helping regulated institutions... ...Opportunity? The AI Mobile Engineer- IOS is a senior technical implementation... ...designs into robust, reliable, maintainable implementations... ...monitoring, logging, and observability for agent systems in production...Full time- ...production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, equity trading,...Permanent employmentRelocationFlexible hours
$135k - $210k per year
...Overview: Guidepoint seeks an experienced AI Engineer as an integral member of the Toronto-... ...like LLM Judges or MLflow, AI observability, and system monitoring. Evaluate and... ...# Live Coding & System Design Test (On-site, 2 hours) # Technical Leadership Interview...Full timeWorldwide- ...building a category-defining AI platform for restaurants. This... ...looking for an AI Full-Stack Engineer who treats growth as an engineering... ..., SEO and generative-engine optimization, paid acquisition... ...acquisition products. Establish reliable end-to-end attribution from source...Long term contractFull timeTemporary work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - Dynatrace & AI Observability. Be the first to apply!
- site reliability engineer intern Toronto, ON
- site reliability engineer Toronto, ON
- site reliability engineer remote Toronto, ON
- senior site reliability engineer Toronto, ON
- site maintenance Toronto, ON
- site safety Toronto, ON
- website developer Toronto, ON
- site reliability engineer intern
- site reliability engineer
- site reliability engineer sre
