Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer, AI/ML Infrastructure

Boson AI

Overview

We2;re looking for a Senior Site Reliability Engineer to help us run one of the most exciting GPU clusters aroundour Toronto datacenter packed with NVIDIA H100 and A100 GPUs, over 20PB of Ceph storage, terabit networking, and hundreds of servers.

Youll be hands-on with the full lifecycle of HPC infrastructure: planning, building, testing, deploying, and keeping everything running smoothly. That means troubleshooting issues as they arise, monitoring performance, developing automation to make our lives easier, and working closely with engineering and science teams to ensure they have what they need. Youll also help us plan for future capacity and evaluate new technologies as we continue to scale.

Responsibilities

  • Manage and optimize HPC cluster operations
  • Deploy and maintain infrastructure-as-code solutions
  • Support ML/research teams with cluster usage optimization
  • Operate, troubleshoot and optimize Ceph storage clusters
  • Develop automation and tooling

Minimum Qualifications

  • 5+ years of experience in SRE or HPC operations
  • Proficiency in Linux systems administration (Ubuntu/Debian)
  • Experience with Kubernetes and container orchestration
  • Experience with Ceph >1PB deployments and maintenance
  • Knowledge of security best practices in multi-tenant environments
  • Understanding of L2/L3 networking fundamentals
  • Skilled in Python and Bash scripting

Preferred Qualifications

  • Experience with infrastructure-as-code tools (Ansible/Terraform)
  • Experience with GitOps (Helm, ArgoCD)
  • Strong grasp of RDMA, InfiniBand, and GPUDirect technologies
  • Familiarity with deep learning frameworks such as PyTorch and TensorFlow
  • Familiarity in at least one cloud platform: AWS, Azure or GCP

If youre a natural problem-solver with a passion for continuous learning, wed love to hear from you.

We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.

#J-18808-Ljbffr
Vacancy posted 3 days ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer, AI/ML Infrastructure in Toronto, ON vacancy
  • $110k - $120k per year

     ...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for...  ...Information Security and Compliance teams to validate that infrastructure and deployment practices meet data protection and privacy... 
    Suggested
    Temporary work
    Internship
    Work at office
    Remote work

    Momentum Financial Services Group

    Toronto, ON
    20 days ago
  •  ...WFO) 12 months We are seeking an experienced Site Reliability Engineer to join our Enterprise Kubernetes Platform team at a...  ...toil through intelligent tooling, self-healing systems, and AI-assisted operational workflows. You will work alongside platform... 
    Suggested
    Contract work

    Astra North Infoteck Inc.

    Toronto, ON
    6 days ago
  •  ...Site Reliability Engineer Location: Toronto ON Work Model: Hybrid (2 days per week in-person at the Toronto office preferred) Required...  ...across distributed environments. Leverage Dynatrace Davis AI automation and cloud technologies to enable proactive... 
    Suggested
    Full time
    Work at office
    2 days per week

    Astra North Infoteck Inc.

    Toronto, ON
    13 days ago
  • $100k - $125k per year

     ...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,...  ...systems in real time.  Why join Tipalti? Tipalti is the AI-powered platform for finance automation, elevating how finance... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Tipalti

    Toronto, ON
    19 hours ago
  • $78.62 - $89.34 per hour

    Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering...  ...relational and non-relational database design with deep Infrastructure as Code (IaC) scripting skills, enabling development teams... 
    Suggested
    Long term contract
    Permanent employment
    Full time
    Contract work
    Work at office

    Randstad

    Toronto, ON
    4 days ago
  •  ...proprietary generative media models and AI native creative workflows, tackling unsolved...  ..., and craft as much as research and engineering – shipping experiences that creatives actually...  ...The Role As a Software Engineer, ML Infrastructure at Ideogram, you'll build the systems... 
    Full time
    Work at office

    ideogram

    Toronto, ON
    18 hours ago
  • $110k - $150k per year

     ...give you the space to grow.   About the Role We are seeking ML/AI Engineers to contribute to major projects.  The ML / AI Engineer...  ...end ML lifecycle, ensuring models and AI services are scalable, reliable, secure, and deliver measurable business value. The role will... 
    Permanent employment
    Full time
    Remote work
    Flexible hours

    Levio

    Toronto, ON
    19 hours ago
  •  ...Senior Site Reliability Engineer - Edge Location : Ottawa/Toronto, On-Site Reports to: Head of Security The Role You own the edge compute...  ...Additional equity granted based on impact We use AI tools to support parts of the hiring process, including screening... 
    Full time

    dominion%20dynamics

    Toronto, ON
    18 hours ago
  •  ...-clearing broker-dealer and brokerage infrastructure for stocks, ETFs, options, crypto, fixed...  ...is a diverse group of experienced engineers, traders, and brokerage professionals...  ...encourage you to apply. Your Role As a Site Reliability Engineer (SRE) at Alpaca, you will... 
    Home office

    Alpaca

    Toronto, ON
    3 days ago
  • $164.6k - $235.1k per year

     ...About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization...  ...automation platforms. Partner with infra lead to align Tubi’s infrastructure & SRE roadmap. Partner with tech leaders to align the SRE... 
    Long term contract
    Remplacement
    Full time
    Contract work
    Temporary work
    Local area
    Flexible hours

    Tubi - Canada

    Toronto, ON
    14 days ago
  • $188.2k - $268.9k per year

     ...We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design...  ...to support ML workloads at scale, ensuring reliability, observability, and operational excellence.... 
    Long term contract
    Remplacement
    Full time
    Temporary work
    Work at office
    Local area
    Flexible hours

    Tubi - Canada

    Toronto, ON
    16 days ago
  • $102.7k - $137k per year

     ...Position Title: Sr. Site Reliability Engineer, IE&O Position Type: Regular - Full-Time Requisition ID: 42659 McCainers are...  ...native systems, embed observability into applications and infrastructure, automate operational workflows, and help scale SRE practices... 
    Full time

    McCain Foods

    Toronto, ON
    19 hours ago
  • Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure. This remote role positions you to improve our platform's performance and scalability. In this senior role, you will lead the design and maintenance of... 
    Remote work

    Viafoura

    Toronto, ON
    19 hours ago
  •  ...you work fast, flexibly, and collaboratively — without compromising standards — we want to hear from you. We’re looking for an AI/ML Engineer who will develop, optimize, and scale machine learning models that power our next generation of user experiences. Working closely... 
    Full time
    Worldwide

    USMobile

    Toronto, ON
    27 days ago
  • $78.62 - $89.34 per hour

    Our client, is seeking a talented and proactive Site Reliability Specialist / Senior Cloud Platform Engineer to join their core Cloud Engineering division....  ...structural lifecycle support of enterprise Azure infrastructure and Databricks platforms. This position is ideal... 
    Long term contract
    Permanent employment
    Contract work
    Work at office

    Randstad

    Toronto, ON
    4 days ago
  •  ...Site Reliability Engineer – APM, Dynatrace, Observability Location • Toronto, ON • Hybrid – 2 days in office per week Required...  ...• Amazon ECS • Azure Functions • Understanding of AI-based system fundamentals, including how AI systems are built... 
    Contract work
    Work at office
    2 days per week

    Astra North Infoteck Inc.

    Toronto, ON
    14 days ago
  •  ...Platform Engineer – DevOps, Site Reliability Engineering (SRE) & Dynatrace Required Skills • Strong experience as a Platform Engineer with...  ..., automating, and supporting enterprise-scale platform infrastructure. • Strong knowledge of high availability, reliability... 
    Permanent employment

    Astra North Infoteck Inc.

    Toronto, ON
    4 days ago
  • $141k - $191k per year

     ...software development, and technology infrastructure? If yes, come join our team at Thomson...  ...Manager, you will lead a team of 10+ engineers, oversee their development and ensure...  ...the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible... 
    Work at office
    Local area
    Flexible hours
    2 days per week
    3 days per week

    Thomson Reuters

    Toronto, ON
    more than 2 months ago
  • $130k - $145k per year

     ...are seeking a highly skilled and passionate Senior DevOps Engineer, AI infrastructure to join our team! You will play a key role in designing, developing...  ..., memory management, and tool orchestration for safe and reliable execution. Process and analyze large datasets of... 
    Full time
    Internship
    Work at office
    Flexible hours

    Financeit

    Toronto, ON
    12 days ago
  • $60 - $75 per hour

     ...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors...  ...Qualifications Must-Have ~3+ years of experience in software engineering or data science & analytics . Application Process (... 
    Remote job
    Contract work
    Summer work

    Mercor

    Toronto, ON
    20 days ago
  •  ...belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an engineer, a builder, and a...  ..., cost-effective, and secure by default.  Scaling cloud infrastructure to support our Kubernetes-based ecosystem.  Maintaining... 
    Full time
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to friday
    Flexible hours

    iManage

    Toronto, ON
    more than 2 months ago
  •  ...Engineering Manager, AI Conversation Platform Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises...  ...Driving an ambitious vision for AI/ML that benefits our users Setting the... 

    Stripe

    Toronto, ON
    13 days ago
  • $101k - $169k per year

     ...take control and receive the best solutions to their complex challenges.   Lead and oversee complex, high-impact engagements in the AI Model Risk Space, ensuring alignment with industry best practices, regulatory expectations, and emerging AI governance standards. Guide... 
    Permanent employment
    Contract work
    Flexible hours

    Deloitte

    Toronto, ON
    1 day ago
  • $120k per year

     ...Python and Kubernetes Software Engineer - Data, AI/ML & Analytics Join to apply for the Python and Kubernetes Software Engineer - Data, AI...  ...building open source solutions for public cloud and private infrastructure. As a software engineer on the team, you'll collaborate... 
    Full time
    Freelance
    Local area
    Work from home
    Worldwide

    Canonical

    Toronto, ON
    3 days ago
  •  ...We are looking for a Database Reliability Engineer to join our team. This is not a traditional...  ...database knowledge. You think in code, manage infrastructure via Terraform, and treat database...  ...Services?  Take a look at our careers site and you’ll find everything you’d expect... 
    Permanent employment
    Full time
    Internship
    Remote work
    Worldwide

    MUFG Investor Services

    Toronto, ON
    12 days ago
  • $135k - $210k per year

     ...Overview:   Guidepoint seeks an experienced AI Engineer as an integral member of the Toronto-based AI team....  ...Technology Hub serves as the base of our Data/AI/ML team, dedicated to building a modern data infrastructure for advanced analytics and the development of responsible... 
    Full time
    Worldwide

    Guidepoint

    Toronto, ON
    3 days ago
  • $100k - $115k per year

     ...delivering strategic technology and AI-driven initiatives. The ideal...  ...owners, architects, and engineering teams to align priorities....  ...Engineering, Security, Compliance, and Infrastructure teams.   Executive &...  ..., analytics platforms, and AI/ML-enabled solutions. Experience... 
    Permanent employment
    Full time
    Work at office
    Local area
    3 days per week

    INFOYA

    Toronto, ON
    22 days ago
  • $72k - $138k per year

     ...-- We are looking for a passionate AI Research Engineer to join our team. You will work at the intersection...  ...AI (GenAI) systems that are both reliable and impactful. This role blends...  ...Intelligence (Al) practice is comprised of Al/ML experts with hands-on experience in developing... 
    Permanent employment
    Flexible hours

    Deloitte

    Toronto, ON
    10 hours ago
  • $94k - $113k per year

     ...company data to drive business solutions. Feature engineering, selection and optimization models with ML techniques. Improve data collection procedures...  ...~ Familiarity with LLMs, prompt engineering, and AI application development using APIs (e.g. OpenAI, Anthropic... 
    Full time

    Project X Ltd.

    Toronto, ON
    3 days ago
  •  ...AI Automation Engineer – AI/ML, Test Automation, Python/Java & SAP Location: • Toronto, ON - Hybrid (4 Days WFO) Duration: • 12 months Role Description: • We are seeking experienced automation engineers to support our Quality Engineering and... 
    Permanent employment
    Shift work

    Astra North Infoteck Inc.

    Toronto, ON
    11 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer, AI/ML Infrastructure. Be the first to apply!