Performance & Reliability Engineer
Cerebras Systems
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. Our novel wafer-scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programming simplicity of a single device. This approach allows Cerebras to deliver industry-leading training and inference speeds and empowers machine learning users to effortlessly run large-scale ML applications, without the hassle of managing hundreds of GPUs or TPUs.
Cerebras' current customers include top model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras , to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
Thanks to the groundbreaking wafer-scale architecture, Cerebras Inference offers the fastest Generative AI inference solution in the world, over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
About The Role
Join Cerebras as a Performance & Reliability Engineer within our innovative Co-Design and Next Generation Team. Our groundbreaking CS-3 system has set new benchmarks in high-performance ML training and inference solutions. It leverages a dinner-plate sized chip with 44GB of on-chip memory to surpass traditional hardware capabilities. This role focuses on characterizing and optimizing the performance and reliability of state-of-the-art AI models running on Cerebras' breakthrough hardware.
Responsibilities
- Characterize and enhance the performance and reliability of advanced ML hardware/software systems, with emphasis on reducing power and thermal fluctuations.
- Analyze ML workloads, software kernels, and hardware architecture for power and performance impacts, and synthesize high-level insights across these layers.
- Develop creative software solutions to improve reliability and performance, collaborating cross-functionally to deploy these solutions in production.
- Influence the design of Cerebras' next-generation AI architecture and software stack through rigorous workload analysis and computational efficiency optimization.
- Partner with ML engineers, researchers, and reliability specialists to understand model behavior and drive system-level improvements from a software perspective.
- Collaborate with teams in architecture, silicon, and research to advance our computational platforms and influence future system designs.
Skills & Qualifications
- BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field.
- 3+ years of relevant experience in performance engineering, reliability, computer architecture, and/or software design.
- Proficiency in Python or other scripting languages.
- Experience with C/C++ and assembly programming.
- Demonstrated expertise with system-level performance and reliability optimization.
- Strong verbal and written communication skills.
- Nice to have: Hands-on experience with ML models, ML frameworks, and collective communication.
- Nice to have: Understanding of thermal management principles and power delivery for advanced semiconductors.
Why Join Cerebras
People who are serious about software make their own hardware. At Cerebras we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open source their cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.
Read our blog: Five Reasons to Join Cerebras in 2026.
Apply today and become part of the forefront of groundbreaking advancements in AI!
Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.
This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.
$110k - $125k per year
...patients everyday! Apply today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure...PerformanceFull timeRemote workFlexible hours- ...unlocking real-time iteration and increasing intelligence via additional agentic computation. About The Role Join Cerebras as a Performance Engineer within our innovative Runtime Team. Our groundbreaking CS-3 system, hosted by a distributed set of modern and powerful x86...PerformanceFull timeLocal area
$115k - $125k per year
...Putting People First, Outstanding Corporate Citizenship, High Performance Culture, and Rigorous Financial Discipline. Mining... ...York Stock Exchange (symbol: KGC). Job Summary The Reliability Engineer, as part of the Asset Management team, plays a vital role in...PerformanceLong term contractTemporary workFor contractorsCasual workLocal areaImmediate startRemote workOverseas$110k - $120k per year
...Discretionary Annual Bonus – Rewarding both individual and company performance Comprehensive Benefits – Health and dental coverage with... ...programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring...PerformanceFull timeTemporary workInternshipWork at officeRemote work- ...cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost... .... Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy... ...of AI hardware by helping build highly reliable systems that power tomorrow's largest...PerformancePermanent employmentFull timeInternshipSecond job
$100k - $125k per year
...experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability... ...role in enhancing the reliability, performance, and scalability of our systems and... ...Tipalti Our tech teams are the engine behind our business. Tipalti’s tech...PerformanceFull timeWork at officeFlexible hours- ...Most Innovative Companies, and Forbes World’s Best Bank. Visit our institutional page About the role Senior System Engineer - Systems Performance Team The Systems Performance team is part of the Computing Squad (Foundation / Runtime Platforms). You will be part of a...PerformanceLong term contractFull timeRemote workWork from homeRelocation packageFlexible hours
$100k per year
...the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining... ...seniorities. Tenstorrent is looking for an experienced Reliability Engineer to drive and execute reliability testing for our products...PerformancePermanent employmentFull timeInternship- ...À titre de conseiller principal, gestion de la performance financière au sein de l’équipe Gestion de la performance financière - Entreprises de la Banque Nationale, cela signifie agir à titre de partenaire stratégique auprès des dirigeants d’entreprise et contribuer directement...PerformanceApprenticeshipFlexible hours
$154k - $200k per year
...America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with... ...establishing standards that minimize manual work, increase performance and reduce incident frequency and severity of incidents Own...PerformanceLong term contractFull time$146k - $201.3k per year
...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity Reporting to the Director of Quality & Performance, this role as Engineering Manager of Performance & Resilience will drive performance and...PerformanceFull timeLocal areaWorldwide- ...Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team... ...: Lead, coach, and develop a high-performing team of SREs, taking ownership of hiring... ...onboarding, mentoring, and continuous performance management. Inspire the Culture:...PerformanceLong term contractFull timeFor contractorsWork at officeWorldwide3 days per week
- ..., collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a systems... ...based on industry data. Rewarding me with an annual performance-based bonus. Offering comprehensive Health/Vision/Dental...PerformanceFull timeWork at officeLocal areaRemote workWorldwideMonday to fridayFlexible hours
- ...learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident... ...Run structured reliability work such as capacity forecasting, performance tuning, and controlled failure testing. Improve how the org...PerformanceFull timeFor contractorsWork at officeWorldwide3 days per week
$153.82k - $277k per year
...thrive, we can’t wait to meet you. Site Reliability Engineers (SREs) at Braze are responsible for... ...and helping translate them into reliable, highly scalable technology stacks. You... ...Configure, tune, and operate Braze's high-performance NGINX routing, proxying, and ingress...PerformancePermanent employmentFull timeInternshipWork at officeLocal areaRemote workFlexible hoursRotating shift$140k - $180k per year
...the shared infrastructure that all other engineering teams at Tripstack depend on. This... ...sits, and the deployment patterns and reliability standards that apply across both. This... ...acting as incident commander, managing performance, and developing engineers into senior roles...PerformanceRemplacementFull timeWork at officeImmediate startFlexible hours- ...advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in software... ...experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly evaluate products and services...PerformanceFull time
- ...As Chief Advisor, Financial Performance Management within the Financial Performance Management - Commercial Banking team at National Bank means acting as a strategic partner to business leaders and contributing directly to the organization’s financial objectives. In...PerformanceFlexible hours
$140k - $182k per year
...Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and software... ...automation of our infrastructure to minimize manual work, increase performance, and decrease the frequency and severity of incidents...PerformanceFull time$140k - $155k per year
...profiles! This is a hands-on senior engineering role focused on improving production... ...enable engineering teams to ship secure, reliable, and scalable software with confidence.... ...reducing operational toil, improving system performance, increasing platform reliability, and...PerformanceRemote jobPermanent employmentFull timeFlexible hours$165k - $220k per year
...terms. About This Team And Role The Mozilla Firefox Performance team is a community of engineers who care deeply about delivering the fastest browser... ...browser as well as helping other teams write fast and reliable code to make Firefox excellent for users. Do you value...PerformanceImmediate startHome office- ...spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in... ...are highly available, scalable, and performant. This role blends deep technical expertise... ...Design, develop, and automate reliable cloud infrastructure and platform services...PerformanceFull timeManual laborLocal areaFlexible hours
$170k - $185k per year
...Interested in joining one of Canada’s top-performing asset managers? We are seeking a Head of Platform Engineering, Reliability & Control to lead our horizontal engineering... ...and strengthening end-to-end data pipeline performance. ~ Standardize developer experience by...PerformanceFull timeWork at office3 days per week- ...Schedule: 40 hours/week – 100% remote work Job Description We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability, performance, and resilience of high-availability SaaS platforms. Working in an AWS and Kubernetes...PerformancePermanent employmentWork at officeLocal areaRemote work
$243k - $297k per year
...more resilient businesses. As Relay continues to scale, the reliability, performance, and resilience of our platform are no longer just technical... ...not only for guiding a strong team of Site Reliability Engineers, but for shaping how reliability strategy influences engineering...PerformanceLong term contractFull timeInternshipWork at officeTrial periodFlexible hours$130k - $180k per year
...we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud... ...CI/CD pipelines, enabling rapid and reliable software releases. Automate and optimize... ...of the end-to-end availability and performance of our cloud infrastructure;...PerformanceFull timeRemote workVisa sponsorshipWork visaFlexible hours- ...and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure... ...-scale AI training and serving: high-performance networks, GPU clusters, storage,... ...the operational tooling that keeps them reliable. This is a hands-on role for someone...PerformanceFull timeRemote work
$172k - $229k per year
...BuildOps is looking for a Staff Software Engineer to set and drive our company-wide... ...strategy for building, shipping, and operating reliable software. This is a high-impact,... ...validation, developer platforms, testability, performance engineering, or incident learning....PerformanceLong term contractPermanent employmentFull timeFor contractorsWork at officeLocal areaWork from homeFlexible hours- ...best for our customers. Cohere is a team of researchers, engineers, designers, and more, who are passionate about their... ...future! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and...PerformanceFull timeWork at officeRemote workFlexible hours
- ...What youll be doing As a member of CIBCs Application Reliability Engineering Platform team the Consultant Site Reliability Engineering... ...most optimal for you to thrive in your role. To successfully perform the work youll be on-site full-time. Youll have the flexibility...PerformanceFull timeContract work3 days per week1 day per week
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Performance & Reliability Engineer. Be the first to apply!

