Site Reliability Engineer, Metal
$100k per yearTenstorrent
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities.
Tenstorrent is building large-scale AI systems across internal clusters and customer deployments. This role sits at the intersection of site reliability, infrastructure operations, and customer engineering, ensuring our systems are reliable, observable, and production-ready.
This role is hybrid, based out of Toronto, ON; Austin, TX; or Santa Clara, CA.
We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting.
Who You Are
- Experienced in site reliability, infrastructure, or systems engineering in distributed environments.
- Strong Linux systems knowledge with the ability to troubleshoot complex multi-layer issues.
- Proficient with observability tools such as Prometheus, Grafana, and alerting systems.
- Comfortable with scripting and automation using Python, Go, or similar languages.
- Solid understanding of networking fundamentals and how systems behave at scale.
What We Need
- Ensure reliability and operational health of Tenstorrent systems across internal and customer environments.
- Troubleshoot complex issues across compute, networking, and software layers.
- Partner with engineering teams and customers to resolve production incidents.
- Design and improve monitoring, observability, and alerting systems.
- Build automation to reduce operational toil and improve system reliability.
What You Will Learn
- How large-scale AI infrastructure is operated across internal clusters and customer deployments.
- How distributed systems behave under real-world production conditions.
- How observability and automation drive reliability at scale.
- How hardware, networking, and software systems interact in AI environments.
- How customer-facing AI infrastructure is deployed, supported, and optimized.
Compensation for all engineers at Tenstorrent ranges from $100k - $500k including base and variable compensation targets. Experience, skills, education, background and location all impact the actual offer made.
Tenstorrent offers a highly competitive compensation package and benefits, and we are an equal opportunity employer.
This offer of employment is contingent upon the applicant being eligible to access U.S. export-controlled technology. Due to U.S. export laws, including those codified in the U.S. Export Administration Regulations (EAR), the Company is required to ensure compliance with these laws when transferring technology to nationals of certain countries (such as EAR Country Groups D:1, E1, and E2). These requirements apply to persons located in the U.S. and all countries outside the U.S. As the position offered will have direct and/or indirect access to information, systems, or technologies subject to these laws, the offer may be contingent upon your citizenship/permanent residency status or ability to obtain prior license approval from the U.S. Commerce Department or applicable federal agency. If employment is not possible due to U.S. export laws, any offer of employment will be rescinded.
- ...the world and our goal is preserving uncensored Internet access and online privacy for all. Right now we are looking for a Site Reliability Engineer to help us tame DNS. This position can be hybrid at our Toronto offices or remote in Canada only but you MUST reside in the...SuggestedFull timeDirect hireWork at officeRemote work
$115k - $130k per year
..., and continuously raising the bar. Kaseya is hiring a Site Reliability Engineer to keep our production systems healthy as we scale. You'll own... ...enforce SLOs, SLIs, and error budgets that keep our systems reliable Lead incident response, troubleshooting, and blameless...SuggestedLong term contractFull timeWorldwide$110k - $120k per year
...professional development support, discounts through Perkopolis, and recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring the availability, performance, and resilience of the...SuggestedFull timeTemporary workInternshipWork at officeRemote work- ...challenges with continuous learning opportunities, then Tescys could be a good fit for you! About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a team at the heart of platform reliability for mission-critical...SuggestedRemote jobLong term contractPermanent employmentFull time
$100k - $125k per year
...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,... ...culture Tech at Tipalti Our tech teams are the engine behind our business. Tipalti’s tech ecosystem is extremely...SuggestedFull timeWork at officeFlexible hours- ...competitive advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in software... ...experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly evaluate products...Full time
$154k - $200k per year
...client base with operations throughout North America, Central America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic technical leadership across infrastructure and software development...Long term contractFull time- ...About the Role Fivetran is looking for a high-performance engineer to be a part of a team of Site Reliability Engineers. You will be working closely with engineering teams, product managers, as well as support and sales engineers to build the future of the Fivetran Data...Full timeWork at officeRemote work
- ...of both work styles in a workplace that is intentional about belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a systems thinker. You’ll create middleware and platform guardrails...Full timeWork at officeLocal areaRemote workWorldwideMonday to fridayFlexible hours
- ...Docebians around the world and help us reinvent the way people learn, because learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident response while also shaping the underlying infrastructure that...Full timeFor contractorsWork at officeWorldwide3 days per week
- ...learning never stops. The Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team of engineers dedicated... ...Product, Engineering, and Support leadership to integrate reliable SRE practices into early planning and product delivery...Long term contractFull timeFor contractorsWork at officeWorldwide3 days per week
$140k - $182k per year
...client base with operations throughout North America, Central America, Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and software development. You will support and help inform the evolution of...Full time- ...for keeping a large scale, multi-tenant SaaS platform running reliably for customers around the world. You'll take a hands-on lead role... ...just symptoms, and to raise the reliability bar across the wider engineering org. What You'll Be Doing Take point on major incidents...Full timeImmediate startFlexible hours
- ...and how we use AI in our recruiting process here . The Site Reliability Engineering organization at Pinterest is accountable for ensuring... ...Demonstrated ability to write effective prompts to get high-quality, reliable outputs from LLMs ~ Demonstrated ability to use AI to...Full timeWork at office
- ...Senior Site Reliability Engineer - Edge Location : Ottawa/Toronto, On-Site Reports to: Head of Security The Role You own the edge compute module — the standard hardware stack, OS image, and runtime that integrates with Dominion Dynamics's mesh radios, sensors...Full time
- ...visibility, and optimize spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s... ...Excellence & Automation Design, develop, and automate reliable cloud infrastructure and platform services. Apply Infrastructure...Full timeManual laborLocal areaFlexible hours
$130k - $180k per year
...Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure... ...implement, and maintain CI/CD pipelines, enabling rapid and reliable software releases. Automate and optimize our infrastructure...Full timeRemote workVisa sponsorshipWork visaFlexible hours$243k - $297k per year
...more resilient businesses. As Relay continues to scale, the reliability, performance, and resilience of our platform are no longer... ...role responsible not only for guiding a strong team of Site Reliability Engineers, but for shaping how reliability strategy influences engineering...Long term contractFull timeInternshipWork at officeTrial periodFlexible hours$110k - $125k per year
...systems bringing #BetterGlobalHealth to patients everyday! Apply today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across...Full timeRemote workFlexible hours- ...San Francisco and founded in 2014, Tubi is part of Tubi Media Group, a division of Fox Corporation. About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that applies a developer's mindset...RemplacementFull timeContract workTemporary workFlexible hours
- ...customers. Cohere is a team of researchers, engineers, designers, and more, who are passionate... ...building high-performance, scalable and reliable machine learning systems? Do you want to... ...NLP applications? We are looking for a Site Reliability Engineer to join the Model...Full timeWork at officeRemote workFlexible hours
$136k - $187k per year
...users worldwide. Our commitment to reliability is a key foundation of our product and... ...availability expectations is a core engineering focus. As a Senior Site Reliability Engineer, you'll join... ...solutions that make our system more reliable by design. What you’ll do: Design...Local areaRemote workWorldwide$144k - $200k per year
**The Team** Platform Engineering is the department within SRE that is responsible for a range... ...role in developing and maintaining the reliable and globally connected multi-cloud network... ...Overview** We are seeking a talented Site Reliability Engineer (SRE) with a strong...Full timeWork at officeRemote workWorldwideFlexible hours$104.24k - $143.3k per year
...Microsoft Azure cloud technologies in a team that adopts DevOps methodologies at an innovative and growing company. Our Lead Site Reliability Engineers will provide a stable infrastructure platform throughout our clients initial onboarding journey with SimCorp -whether that...Full timeWork at officeImmediate startRemote workShift workWeekend work2 days per week- ...that are particularly strong in a few areas, and have some interest and capabilities in others. About the Role: As a Site Reliability Engineer , you’ll join the global Platform SRE team responsible for building, operating, and scaling Kong’s multi-region SaaS...Full time
$150k - $250k per year
...About The Role We're looking for a Senior Site Reliability Engineer to help us run one of the most exciting GPU clusters around—our Toronto datacenter packed with NVIDIA H100 and A100 GPUs, over 20PB of Ceph storage, terabit networking, and hundreds of servers. You'll...Full time- ...Platform Engineer – DevOps, Site Reliability Engineering (SRE) & Dynatrace Required Skills • Strong experience as a Platform Engineer with expertise in DevOps and Site Reliability Engineering (SRE). • Experience designing, implementing, automating, and supporting...Permanent employment
$141k - $191k per year
...and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence. About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for: Team Leadership...Work at officeLocal areaFlexible hours2 days per week3 days per week- ...Site Reliability Engineer Location: Toronto ON Work Model: Hybrid (2 days per week in-person at the Toronto office preferred) Required Skills Site Reliability Engineering (SRE) DevOps Dynatrace Role Summary Design implement and optimize...Full timeWork at office2 days per week
- ...services to define SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of the storage layer that underpins... ...over manual processes. We are a small team of software engineers with a strong bias towards software solutions to avoid toil...Long term contractFull timeWork at officeLocal areaImmediate startRemote workWorldwideShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer, Metal. Be the first to apply!
- site reliability engineer intern Toronto, ON
- site reliability engineer remote Toronto, ON
- site reliability engineer Toronto, ON
- senior site reliability engineer Toronto, ON
- website developer Toronto, ON
- site safety Toronto, ON
- site maintenance Toronto, ON
- site reliability engineer intern
- site reliability engineer remote
- site reliability engineer sre
