Site Reliability Engineer, AI/ML Infrastructure
Boson AI
Overview
We2;re looking for a Senior Site Reliability Engineer to help us run one of the most exciting GPU clusters aroundour Toronto datacenter packed with NVIDIA H100 and A100 GPUs, over 20PB of Ceph storage, terabit networking, and hundreds of servers.
Youll be hands-on with the full lifecycle of HPC infrastructure: planning, building, testing, deploying, and keeping everything running smoothly. That means troubleshooting issues as they arise, monitoring performance, developing automation to make our lives easier, and working closely with engineering and science teams to ensure they have what they need. Youll also help us plan for future capacity and evaluate new technologies as we continue to scale.
Responsibilities
- Manage and optimize HPC cluster operations
- Deploy and maintain infrastructure-as-code solutions
- Support ML/research teams with cluster usage optimization
- Operate, troubleshoot and optimize Ceph storage clusters
- Develop automation and tooling
Minimum Qualifications
- 5+ years of experience in SRE or HPC operations
- Proficiency in Linux systems administration (Ubuntu/Debian)
- Experience with Kubernetes and container orchestration
- Experience with Ceph >1PB deployments and maintenance
- Knowledge of security best practices in multi-tenant environments
- Understanding of L2/L3 networking fundamentals
- Skilled in Python and Bash scripting
Preferred Qualifications
- Experience with infrastructure-as-code tools (Ansible/Terraform)
- Experience with GitOps (Helm, ArgoCD)
- Strong grasp of RDMA, InfiniBand, and GPUDirect technologies
- Familiarity with deep learning frameworks such as PyTorch and TensorFlow
- Familiarity in at least one cloud platform: AWS, Azure or GCP
If youre a natural problem-solver with a passion for continuous learning, wed love to hear from you.
We may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist our recruitment team but do not replace human judgment. Final hiring decisions are ultimately made by humans. If you would like more information about how your data is processed, please contact us.
#J-18808-Ljbffr$110k - $120k per year
...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for... ...Information Security and Compliance teams to validate that infrastructure and deployment practices meet data protection and privacy...SuggestedTemporary workInternshipWork at officeRemote work- ...WFO) 12 months We are seeking an experienced Site Reliability Engineer to join our Enterprise Kubernetes Platform team at a... ...toil through intelligent tooling, self-healing systems, and AI-assisted operational workflows. You will work alongside platform...SuggestedContract work
- ...Site Reliability Engineer Location: Toronto ON Work Model: Hybrid (2 days per week in-person at the Toronto office preferred) Required... ...across distributed environments. Leverage Dynatrace Davis AI automation and cloud technologies to enable proactive...SuggestedFull timeWork at office2 days per week
$100k - $125k per year
...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,... ...systems in real time. Why join Tipalti? Tipalti is the AI-powered platform for finance automation, elevating how finance...SuggestedFull timeWork at officeFlexible hours$78.62 - $89.34 per hour
Our client, is seeking a talented and proactive Site Reliability Engineer (SRE) / Senior Database Platform Engineer to join their core Data Engineering... ...relational and non-relational database design with deep Infrastructure as Code (IaC) scripting skills, enabling development teams...SuggestedLong term contractPermanent employmentFull timeContract workWork at office- ...proprietary generative media models and AI native creative workflows, tackling unsolved... ..., and craft as much as research and engineering – shipping experiences that creatives actually... ...The Role As a Software Engineer, ML Infrastructure at Ideogram, you'll build the systems...Full timeWork at office
$110k - $150k per year
...give you the space to grow. About the Role We are seeking ML/AI Engineers to contribute to major projects. The ML / AI Engineer... ...end ML lifecycle, ensuring models and AI services are scalable, reliable, secure, and deliver measurable business value. The role will...Permanent employmentFull timeRemote workFlexible hours- ...Senior Site Reliability Engineer - Edge Location : Ottawa/Toronto, On-Site Reports to: Head of Security The Role You own the edge compute... ...Additional equity granted based on impact We use AI tools to support parts of the hiring process, including screening...Full time
- ...-clearing broker-dealer and brokerage infrastructure for stocks, ETFs, options, crypto, fixed... ...is a diverse group of experienced engineers, traders, and brokerage professionals... ...encourage you to apply. Your Role As a Site Reliability Engineer (SRE) at Alpaca, you will...Home office
$164.6k - $235.1k per year
...About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization... ...automation platforms. Partner with infra lead to align Tubi’s infrastructure & SRE roadmap. Partner with tech leaders to align the SRE...Long term contractRemplacementFull timeContract workTemporary workLocal areaFlexible hours$188.2k - $268.9k per year
...We are seeking a Director of Machine Learning Engineering and Infrastructure to lead a hybrid team bridging advanced ML engineering with world-class infrastructure design... ...to support ML workloads at scale, ensuring reliability, observability, and operational excellence....Long term contractRemplacementFull timeTemporary workWork at officeLocal areaFlexible hours$102.7k - $137k per year
...Position Title: Sr. Site Reliability Engineer, IE&O Position Type: Regular - Full-Time Requisition ID: 42659 McCainers are... ...native systems, embed observability into applications and infrastructure, automate operational workflows, and help scale SRE practices...Full time- Advance your career as a Senior Site Reliability Engineer at Viafoura, specializing in Kubernetes and AWS infrastructure. This remote role positions you to improve our platform's performance and scalability. In this senior role, you will lead the design and maintenance of...Remote work
- ...you work fast, flexibly, and collaboratively — without compromising standards — we want to hear from you. We’re looking for an AI/ML Engineer who will develop, optimize, and scale machine learning models that power our next generation of user experiences. Working closely...Full timeWorldwide
$78.62 - $89.34 per hour
Our client, is seeking a talented and proactive Site Reliability Specialist / Senior Cloud Platform Engineer to join their core Cloud Engineering division.... ...structural lifecycle support of enterprise Azure infrastructure and Databricks platforms. This position is ideal...Long term contractPermanent employmentContract workWork at office- ...Site Reliability Engineer – APM, Dynatrace, Observability Location • Toronto, ON • Hybrid – 2 days in office per week Required... ...• Amazon ECS • Azure Functions • Understanding of AI-based system fundamentals, including how AI systems are built...Contract workWork at office2 days per week
- ...Platform Engineer – DevOps, Site Reliability Engineering (SRE) & Dynatrace Required Skills • Strong experience as a Platform Engineer with... ..., automating, and supporting enterprise-scale platform infrastructure. • Strong knowledge of high availability, reliability...Permanent employment
$141k - $191k per year
...software development, and technology infrastructure? If yes, come join our team at Thomson... ...Manager, you will lead a team of 10+ engineers, oversee their development and ensure... ...the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible...Work at officeLocal areaFlexible hours2 days per week3 days per week$130k - $145k per year
...are seeking a highly skilled and passionate Senior DevOps Engineer, AI infrastructure to join our team! You will play a key role in designing, developing... ..., memory management, and tool orchestration for safe and reliable execution. Process and analyze large datasets of...Full timeInternshipWork at officeFlexible hours$60 - $75 per hour
...Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors... ...Qualifications Must-Have ~3+ years of experience in software engineering or data science & analytics . Application Process (...Remote jobContract workSummer work- ...belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a... ..., cost-effective, and secure by default. Scaling cloud infrastructure to support our Kubernetes-based ecosystem. Maintaining...Full timeWork at officeLocal areaRemote workWorldwideMonday to fridayFlexible hours
- ...Engineering Manager, AI Conversation Platform Who we are About Stripe Stripe is a financial infrastructure platform for businesses. Millions of companies—from the world’s largest enterprises... ...Driving an ambitious vision for AI/ML that benefits our users Setting the...
$101k - $169k per year
...take control and receive the best solutions to their complex challenges. Lead and oversee complex, high-impact engagements in the AI Model Risk Space, ensuring alignment with industry best practices, regulatory expectations, and emerging AI governance standards. Guide...Permanent employmentContract workFlexible hours$120k per year
...Python and Kubernetes Software Engineer - Data, AI/ML & Analytics Join to apply for the Python and Kubernetes Software Engineer - Data, AI... ...building open source solutions for public cloud and private infrastructure. As a software engineer on the team, you'll collaborate...Full timeFreelanceLocal areaWork from homeWorldwide- ...We are looking for a Database Reliability Engineer to join our team. This is not a traditional... ...database knowledge. You think in code, manage infrastructure via Terraform, and treat database... ...Services? Take a look at our careers site and you’ll find everything you’d expect...Permanent employmentFull timeInternshipRemote workWorldwide
$135k - $210k per year
...Overview: Guidepoint seeks an experienced AI Engineer as an integral member of the Toronto-based AI team.... ...Technology Hub serves as the base of our Data/AI/ML team, dedicated to building a modern data infrastructure for advanced analytics and the development of responsible...Full timeWorldwide$100k - $115k per year
...delivering strategic technology and AI-driven initiatives. The ideal... ...owners, architects, and engineering teams to align priorities.... ...Engineering, Security, Compliance, and Infrastructure teams. Executive &... ..., analytics platforms, and AI/ML-enabled solutions. Experience...Permanent employmentFull timeWork at officeLocal area3 days per week$72k - $138k per year
...-- We are looking for a passionate AI Research Engineer to join our team. You will work at the intersection... ...AI (GenAI) systems that are both reliable and impactful. This role blends... ...Intelligence (Al) practice is comprised of Al/ML experts with hands-on experience in developing...Permanent employmentFlexible hours$94k - $113k per year
...company data to drive business solutions. Feature engineering, selection and optimization models with ML techniques. Improve data collection procedures... ...~ Familiarity with LLMs, prompt engineering, and AI application development using APIs (e.g. OpenAI, Anthropic...Full time- ...AI Automation Engineer – AI/ML, Test Automation, Python/Java & SAP Location: • Toronto, ON - Hybrid (4 Days WFO) Duration: • 12 months Role Description: • We are seeking experienced automation engineers to support our Quality Engineering and...Permanent employmentShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer, AI/ML Infrastructure. Be the first to apply!
- senior site reliability engineer Toronto, ON
- site reliability engineer intern Toronto, ON
- site reliability engineer Toronto, ON
- site reliability engineer remote Toronto, ON
- site safety Toronto, ON
- website developer Toronto, ON
- site maintenance Toronto, ON
- site carpenter Toronto, ON
- machine learning researcher Toronto, ON
- machine learning part time Toronto, ON
