Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer (SRE) - Dynatrace & AI Observability

Temporary

Astra North Infoteck Inc.

Site Reliability Engineer 

Location: Toronto, ON

Work Model: Hybrid (2 days per week in-person at the Toronto office preferred)

Required Skills

• Site Reliability Engineering (SRE)

• DevOps

• Dynatrace

Role Summary

• Design, implement, and optimize Site Reliability Engineering (SRE) and DevOps practices to ensure high system availability, performance, and reliability across distributed environments.

• Leverage Dynatrace, Davis AI, automation, and cloud technologies to enable proactive monitoring, intelligent automation, and operational excellence.

Role Description

Dynatrace & AI-Driven Observability

• Lead the implementation and optimization of the Dynatrace platform across applications and infrastructure.

• Utilize Dynatrace Davis AI for automated root cause analysis, anomaly detection, event correlation, predictive performance insights, and alert noise reduction.

• Configure and manage OneAgent deployments, Smartscape topology mapping, service flow, and distributed tracing.

• Define and monitor SLIs, SLOs, and user experience metrics.

• Build custom dashboards, alerts, and observability pipelines.

• Integrate Dynatrace with CI/CD pipelines for release validation and performance gating.

• Integrate Dynatrace with incident management tools such as PagerDuty and ServiceNow.

• Enable self-healing automation using Dynatrace event triggers and AI-driven insights.

Automation & Configuration Management

• Design and implement automation solutions using Ansible.

• Automate configuration management, application deployments, and environment provisioning.

• Develop reusable Ansible playbooks and roles for scalable operations.

• Automate operational tasks, patching, compliance processes, and remediation workflows.

• Integrate Ansible with CI/CD pipelines and monitoring systems.

Cloud & DevOps

• Design and manage cloud-native solutions on AWS with exposure to Azure.

• Develop infrastructure using Terraform, CloudFormation, or AWS CDK.

• Build and manage CI/CD pipelines using GitHub Actions, Jenkins, or GitLab CI.

• Develop and deploy serverless solutions using AWS Lambda, API Gateway, and Step Functions.

• Automate DevOps and operational workflows using Python (boto3) and Bash scripting.

• Deploy and maintain production environments through automated pipelines.

• Optimize cloud infrastructure for cost, performance, and scalability.

Monitoring & Reliability Engineering

• Monitor and manage AWS CloudWatch and Azure Monitor/Log Analytics.

• Design unified observability across multi-cloud environments.

• Implement logging and distributed tracing strategies.

• Work with Docker, Kubernetes, ECS, and AKS environments.

• Design fault-tolerant, highly available, and disaster recovery solutions.

• Support incident response, on-call activities, and root cause analysis (RCA).

Required Qualifications

• Proven experience with Dynatrace APM, Real User Monitoring (RUM), and infrastructure monitoring.

• Strong hands-on experience with Dynatrace Davis AI capabilities.

• Experience with Ansible for automation and configuration management.

• Deep knowledge of AWS services and cloud-native architectures.

• Experience with Infrastructure as Code using Terraform, CloudFormation, or AWS CDK.

• Proficiency in Python (boto3) and Bash scripting.

• Experience supporting production-scale environments.

• Business Analyst experience.

• Scrum Master experience.

Nice to Have

• Dynatrace Associate or Professional certification.

• Experience with Dynatrace APIs and automation.

• Experience building self-healing systems using AI-driven triggers.

• Familiarity with Prometheus, Grafana, and the ELK Stack.

• Azure cloud experience and certifications.

• Experience with GitOps and Platform Engineering.

Vacancy posted a month ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer (SRE) - Dynatrace & AI Observability in Toronto, ON vacancy
  •  ...Site Reliability Engineer Location: Toronto ON Work Model: Hybrid (2 days per week in-person at the Toronto office preferred) Required Skills Site Reliability Engineering (SRE) DevOps Dynatrace Role Summary Design implement and optimize... 
    Suggested
    Full time
    Work at office
    2 days per week

    Astra North Infoteck Inc.

    Toronto, ON
    a month ago
  •  ...Platform Engineer – DevOps, Site Reliability Engineering (SRE) & Dynatrace Required Skills • Strong experience as a Platform Engineer with expertise in DevOps and Site Reliability Engineering (SRE). • Experience designing, implementing, automating, and supporting... 
    Suggested
    Permanent employment

    Astra North Infoteck Inc.

    Toronto, ON
    27 days ago
  •  ...Job Role: Senior Platform Engineer DevOps SRE & Dynatrace Observability Location : Toronto - Hybrid (3 Days Work from Office) Job Summary We are seeking a highly skilled and experienced Platform Engineer with strong DevOps and Site Reliability Engineering (SRE... 
    Suggested
    Full time
    Work at office

    Astra North Infoteck Inc.

    Toronto, ON
    14 days ago
  • $141k - $191k per year

     ...at Thomson Reuters and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence...  .... About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for:... 
    Suggested
    Work at office
    Local area
    Flexible hours
    2 days per week
    3 days per week

    Thomson Reuters

    Toronto, ON
    more than 2 months ago
  •  ...is part of the Production Support and Reliability Engineering team, responsible for ensuring the stability...  ...Support, Production Support, SRE, or Platform Operations roles. Strong...  ...reliability. Knowledge of monitoring and observability tools for proactive system management.... 
    Suggested
    Permanent employment
    Full time
    Local area

    Capgemini

    Toronto, ON
    5 hours ago
  • $140k - $180k per year

     ...visit: The role Our SRE team runs the shared infrastructure that all other engineering teams at Tripstack depend on....  ...Prometheus, Thanos, and Grafana observability stack. We are in the middle...  ...the deployment patterns and reliability standards that apply across both... 
    Remplacement
    Full time
    Work at office
    Immediate start
    Flexible hours

    Etraveli Group

    Toronto, ON
    18 hours ago
  • $100k - $125k per year

     ...Position Summary We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,  you will play a crucial role in enhancing the reliability, performance, and scalability of our systems... 
    Full time
    Work at office
    Flexible hours

    Tipalti

    Toronto, ON
    23 hours ago
  • $104.24k - $143.3k per year

     ...methodologies at an innovative and growing company. Our Lead Site Reliability Engineers will provide a stable infrastructure platform throughout our...  ...or a migration of an existing client to our Azure cloud. The SRE will be responsible for aligning client requirements to... 
    Full time
    Work at office
    Immediate start
    Remote work
    Shift work
    Weekend work
    2 days per week

    SimCorp

    Toronto, ON
    8 days ago
  • $110k - $120k per year

     ...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for...  ...UAT, production), ensuring changes are safe, repeatable, and observable. Design and maintain automated CI/CD pipelines and enforce... 
    Temporary work
    Internship
    Work at office
    Remote work

    Momentum Financial Services Group

    Toronto, ON
    a month ago
  •  ...Senior Site Reliability Engineer - Edge Location : Ottawa/Toronto, On-Site Reports to: Head of Security The Role We build rugged edge...  ...Additional equity granted based on impact We use AI tools to support parts of the hiring process, including screening... 
    Full time

    dominion%20dynamics

    Toronto, ON
    22 hours ago
  • $110k - $151.8k per year

     ...Secure Every Identity, from AI to Human Identity is the key to...  ...s talk. The Auth0 Platform Observability team owns the observability tooling...  ...looking for an Observability Engineer to help ensure that our...  ...you have experience within the Site Reliability Engineering (SRE)... 
    Full time
    Local area
    Worldwide

    Okta

    Toronto, ON
    8 days ago
  • $80k - $130k per year

    DevOps SRE Position Description The DevOps Site Reliability Engineering consultant is responsible for environment support, reliability, and operational stability across development, test, and pre production environments. This role combines hands on environment configuration... 
    Toronto, ON
    5 days ago
  •  ...Observability Engineer – Kubernetes, Prometheus, Grafana & Cloud Monitoring Role Overview • Experienced Observability Engineer to join...  ...self-healing infrastructure using modern observability tools and AI/ML capabilities Key Responsibilities • Design,... 
    Long term contract
    Permanent employment

    Astra North Infoteck Inc.

    Toronto, ON
    29 days ago
  • $78.62 - $89.34 per hour

    Our client, is seeking a talented and proactive Site Reliability Specialist / Senior Cloud Platform Engineer to join their core Cloud Engineering division. In this engineering-focused role, you will move far beyond basic operational support to act as a principal architect... 
    Long term contract
    Permanent employment
    Contract work
    Work at office

    Randstad

    Toronto, ON
    27 days ago
  •  ...Palona’s AI agents operate in real restaurant environments: noisy...  ...for an applied AI Modeling Engineer to improve the intelligence, accuracy...  ...hypothesis to experiment to reliable deployment. What you’ll...  ...work is reproducible, tested, observable, and usable by other engineers... 
    Long term contract
    Full time
    Temporary work
    Internship
    Immediate start

    Palona AI

    Toronto, ON
    1 day ago
  •  ...is leading the industry on cutting-edge AI technology, revolutionizing performance...  ...seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy...  ...at manufacturing and test partner sites, including in Taiwan.   Tenstorrent... 
    Permanent employment
    Internship
    Second job

    Tenstorrent

    Toronto, ON
    7 days ago
  •  ...Palona’s AI agents operate continuously in production, handle real-time guest...  ...part of the product: latency, reliability, deployment safety, observability, security, and cost directly shape the...  ...We are looking for an Infrastructure Engineer who combines cloud and reliability depth... 
    Long term contract
    Full time
    Temporary work
    Internship

    Palona AI

    Toronto, ON
    1 day ago
  • $135k - $170k per year

     ...General Information: Job Title: Applied AI Engineer Location: Toronto, ON (Onsite/Hybrid...  ...systems for cost, latency, and reliability Collaborate across teams where needed...  ...experience (vLLM, TGI, llama.cpp) Observability tooling (Langfuse, LangSmith) Prior... 
    Long term contract
    Full time
    Internship
    Work at office
    Immediate start
    Remote work
    Flexible hours

    Fulfillment IQ

    Toronto, ON
    8 days ago
  • $140.6k - $190.6k per year

     ...The Lead AI Forward Engineer designs and guides the delivery of AI-powered solutions that reduce...  ...to production, ensuring solutions meet reliability, security, and compliance expectations...  ...requirements. Integrate AI observability tooling into CI/CD so new models, prompts... 
    Full time
    Manual labor
    Work at office
    Local area
    Flexible hours

    Thomson Reuters

    Toronto, ON
    1 day ago
  •  ...Manager: Senior Director, Data Engineering Contract Terms: Permanent, 37...  ...dedicated engineers , providing reliable and consistent data to the stakeholder...  ...platform operation effort as a Site Reliability Developer! Among many other duties, the SRE team within Ticketmaster's TM1... 
    Permanent employment
    Full time
    Contract work
    Local area
    Remote work
    Flexible hours

    Live Nation

    Toronto, ON
    3 days ago
  •  ...We are looking for a product-focused AI Software Engineer to turn advances in AI into restaurant products that work reliably in the real world. You will build across customer...  ...systems. ~ A strong quality bar for testing, observability, security, reliability, and user experience... 
    Long term contract
    Full time
    Temporary work

    Palona AI

    Toronto, ON
    1 day ago
  • $85 per hour

     ...creative and technical talent with leading AI research labs. Headquartered in San...  ...Jack Dorsey . Position: DevOps / SRE / Cloud Engineer (Coding Agent Experience) Type:...  ...platforms , Kubernetes , CI/CD systems , observability , and infrastructure automation .... 
    Remote job
    Contract work
    Summer work

    Mercor

    Toronto, ON
    1 day ago
  • $250k per year

     ...Role: Observability Engineer – Trading Client:  Elite FinTech Compensation: $120,000 - $250,000 CAD + Bonus Location:  Toronto...  ...Working with multiple technical teams to ensure visibility and reliability across systems. Key Responsibilities Monitoring Tools... 
    Permanent employment
    Immediate start

    Hunter Bond

    Toronto, ON
    19 days ago
  • $72k - $138k per year

     ...and on the job coaching -- We are looking for a passionate AI Research Engineer to join our team. You will work at the intersection of...  ...evaluating, and deploying generative AI (GenAI) systems that are both reliable and impactful. This role blends fundamental research, model... 
    Permanent employment
    Flexible hours

    Deloitte

    Toronto, ON
    14 hours ago
  •  ...Performance Test Engineer – LoadRunner, Performance Testing, Dynatrace & Splunk Required Skills: • Minimum 6+ years of experience in Performance Testing. • Proficiency in LoadRunner with hands-on scripting experience. • Experience leading end-to-end Performance... 

    Astra North Infoteck Inc.

    Toronto, ON
    28 days ago
  • $80k - $180k per year

     ...Zafin is an AI platform company helping regulated institutions modernize...  ...?  The AI Evaluation Engineer ensures AI agent solutions are accurate, reliable, safe, and production-ready within...  ...CI/CD, automated evaluation, AI observability, and engineering delivery practices... 
    Full time

    Zafin

    Toronto, ON
    19 days ago
  • $80k - $190k per year

     ...Zafin is an AI platform company helping regulated institutions...  ...Opportunity?  The  AI Mobile Engineer- IOS is a senior technical implementation...  ...designs into robust, reliable, maintainable implementations...  ...monitoring, logging, and observability for agent systems in production... 
    Full time

    Zafin

    Toronto, ON
    7 days ago
  •  ...production support activity at a high level including ITIL (information technology infrastructure library), monitoring, DevOps, SRE (site reliability engineering), and disaster recovery. How to discuss common financial topics, including financial markets, equity trading,... 
    Permanent employment
    Relocation
    Flexible hours

    mthree Recruiting Portal

    Toronto, ON
    5 days ago
  • $135k - $210k per year

     ...Overview:   Guidepoint seeks an experienced AI Engineer as an integral member of the Toronto-...  ...like LLM Judges or MLflow, AI observability, and system monitoring.  Evaluate and...  ...# Live Coding & System Design Test (On-site, 2 hours) # Technical Leadership Interview... 
    Full time
    Worldwide

    Guidepoint

    Toronto, ON
    25 days ago
  •  ...building a category-defining AI platform for restaurants. This...  ...looking for an AI Full-Stack Engineer who treats growth as an engineering...  ..., SEO and generative-engine optimization, paid acquisition...  ...acquisition products. Establish reliable end-to-end attribution from source... 
    Long term contract
    Full time
    Temporary work

    Palona AI

    Toronto, ON
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer (SRE) - Dynatrace & AI Observability. Be the first to apply!