Performance & Reliability Engineer
Cerebras Systems
Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. Our novel wafer-scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programming simplicity of a single device. This approach allows Cerebras to deliver industry-leading training and inference speeds and empowers machine learning users to effortlessly run large-scale ML applications, without the hassle of managing hundreds of GPUs or TPUs.
Cerebras' current customers include top model labs, global enterprises, and cutting-edge AI-native startups. OpenAI recently announced a multi-year partnership with Cerebras , to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference.
Thanks to the groundbreaking wafer-scale architecture, Cerebras Inference offers the fastest Generative AI inference solution in the world, over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.
About The Role
Join Cerebras as a Performance & Reliability Engineer within our innovative Co-Design and Next Generation Team. Our groundbreaking CS-3 system has set new benchmarks in high-performance ML training and inference solutions. It leverages a dinner-plate sized chip with 44GB of on-chip memory to surpass traditional hardware capabilities. This role focuses on characterizing and optimizing the performance and reliability of state-of-the-art AI models running on Cerebras' breakthrough hardware.
Responsibilities
- Characterize and enhance the performance and reliability of advanced ML hardware/software systems, with emphasis on reducing power and thermal fluctuations.
- Analyze ML workloads, software kernels, and hardware architecture for power and performance impacts, and synthesize high-level insights across these layers.
- Develop creative software solutions to improve reliability and performance, collaborating cross-functionally to deploy these solutions in production.
- Influence the design of Cerebras' next-generation AI architecture and software stack through rigorous workload analysis and computational efficiency optimization.
- Partner with ML engineers, researchers, and reliability specialists to understand model behavior and drive system-level improvements from a software perspective.
- Collaborate with teams in architecture, silicon, and research to advance our computational platforms and influence future system designs.
Skills & Qualifications
- BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field.
- 3+ years of relevant experience in performance engineering, reliability, computer architecture, and/or software design.
- Proficiency in Python or other scripting languages.
- Experience with C/C++ and assembly programming.
- Demonstrated expertise with system-level performance and reliability optimization.
- Strong verbal and written communication skills.
- Nice to have: Hands-on experience with ML models, ML frameworks, and collective communication.
- Nice to have: Understanding of thermal management principles and power delivery for advanced semiconductors.
Why Join Cerebras
People who are serious about software make their own hardware. At Cerebras we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:
- Build a breakthrough AI platform beyond the constraints of the GPU.
- Publish and open source their cutting-edge AI research.
- Work on one of the fastest AI supercomputers in the world.
- Enjoy job stability with startup vitality.
- Our simple, non-corporate work culture that respects individual beliefs.
Read our blog: Five Reasons to Join Cerebras in 2026.
Apply today and become part of the forefront of groundbreaking advancements in AI!
Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer. We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.
This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.
- ...world, over 10 times faster than GPU-based hyperscale cloud inference services. About The Role Join Cerebras as a Performance & Reliability Engineer within our innovative Co-Design and Next Generation Team. Our groundbreaking CS-3 system has set new benchmarks in high-...PerformanceFull time
$110k - $125k per year
...patients everyday! Apply today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure...PerformanceFull timeRemote workFlexible hours- ...unlocking real-time iteration and increasing intelligence via additional agentic computation. About The Role Join Cerebras as a Performance Engineer within our innovative Runtime Team. Our groundbreaking CS-3 system, hosted by a distributed set of modern and powerful x86...PerformanceFull timeLocal area
- ...The primary objective of the Database Reliability Engineer r is to provide expertise across... ...they operate efficiently, securely, and reliable within private and public cloud environments... ...up and restoring data, monitoring performance, and optimizing database operations...PerformanceLong term contractFull timeInternshipRotating shift
$100k per year
...the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining... ...seniorities. Tenstorrent is looking for an experienced Reliability Engineer to drive and execute reliability testing for our products...PerformancePermanent employmentFull timeInternship$115k - $130k per year
...and results. We are a high-growth, high-performance organization that values individuals... ...the bar. Kaseya is hiring a Site Reliability Engineer to keep our production systems healthy... ...and error budgets that keep our systems reliable Lead incident response,...PerformanceLong term contractFull timeWorldwide$110k - $120k per year
...Discretionary Annual Bonus – Rewarding both individual and company performance Comprehensive Benefits – Health and dental coverage with... ...programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring...PerformanceFull timeTemporary workInternshipWork at officeRemote work- ...fit for you! About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a... ...will help maintain, optimize, and ensure the reliability and performance of the systems that power our cloud infrastructure across AWS...PerformanceRemote jobLong term contractPermanent employmentFull time
- ...the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining... ...of all seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next generation...PerformancePermanent employmentInternshipSecond job
- ...cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost... .... Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy... ...of AI hardware by helping build highly reliable systems that power tomorrow's largest...PerformancePermanent employmentFull timeInternshipSecond job
- ...the world and our goal is preserving uncensored Internet access and online privacy for all. Right now we are looking for a Site Reliability Engineer to help us tame DNS. This position can be hybrid at our Toronto offices or remote in Canada only but you MUST reside in the...Full timeDirect hireWork at officeRemote work
$100k - $125k per year
...experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability... ...role in enhancing the reliability, performance, and scalability of our systems and... ...Tipalti Our tech teams are the engine behind our business. Tipalti’s tech...PerformanceFull timeWork at officeFlexible hours$100k per year
...cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost... ...This role sits at the intersection of site reliability, infrastructure operations, and customer engineering, ensuring our systems are reliable, observable, and production-ready....PerformancePermanent employmentFull time- ...advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in software... ...experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly evaluate products and services...PerformanceFull time
- ...Most Innovative Companies, and Forbes World’s Best Bank. Visit our institutional page About the role Senior System Engineer - Systems Performance Team The Systems Performance team is part of the Computing Squad (Foundation / Runtime Platforms). You will be part of a...PerformanceLong term contractFull timeRemote workWork from homeRelocation packageFlexible hours
- ...scale, multi-tenant SaaS platform running reliably for customers around the world. You'll... ...raise the reliability bar across the wider engineering org. What You'll Be Doing Take... ...work such as capacity forecasting, performance tuning, and controlled failure testing....PerformanceFull timeImmediate startFlexible hours
- ...recruiting process here . The Site Reliability Engineering organization at Pinterest is... ...of innovation Manage capacity and performance to help scale our infrastructure both... ...effective prompts to get high-quality, reliable outputs from LLMs ~ Demonstrated ability...PerformanceFull timeWork at office
- ...learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident... ...Run structured reliability work such as capacity forecasting, performance tuning, and controlled failure testing. Improve how the org...PerformanceFull timeFor contractorsWork at officeWorldwide3 days per week
$154k - $200k per year
...America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with... ...establishing standards that minimize manual work, increase performance and reduce incident frequency and severity of incidents Own...PerformanceLong term contractFull time- ...About the Role Fivetran is looking for a high-performance engineer to be a part of a team of Site Reliability Engineers. You will be working closely with engineering teams, product managers, as well as support and sales engineers to build the future of the Fivetran Data...PerformanceFull timeWork at officeRemote work
- ..., collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a systems... ...based on industry data. Rewarding me with an annual performance-based bonus. Offering comprehensive Health/Vision/Dental...PerformanceFull timeWork at officeLocal areaRemote workWorldwideMonday to fridayFlexible hours
$140k - $182k per year
...Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and software... ...automation of our infrastructure to minimize manual work, increase performance, and decrease the frequency and severity of incidents...PerformanceFull time- ...Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team... ...: Lead, coach, and develop a high-performing team of SREs, taking ownership of hiring... ...onboarding, mentoring, and continuous performance management. Inspire the Culture:...PerformanceLong term contractFull timeFor contractorsWork at officeWorldwide3 days per week
$100k - $120k per year
...and expand self-service analytics, we need a platform engineer to own the reliability, performance, and operational excellence of the systems that make it... ...Prometheus, Grafana, or Datadog. ~ Expertise in MPP/OLAP engines, specifically Apache Doris or alternatives like...PerformanceLong term contractFull timeWork at office$136k - $187k per year
...worldwide. Our commitment to reliability is a key foundation of our... ...availability expectations is a core engineering focus. As a Senior Site... ...that make our system more reliable by design. What you’ll do... ...improving the availability, performance, and observability of our services...PerformanceLocal areaRemote workWorldwide- ...a division of Fox Corporation. About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are... ...have a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through...PerformanceRemplacementFull timeContract workTemporary workFlexible hours
$146k - $201.3k per year
...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity Reporting to the Director of Quality & Performance, this role as Engineering Manager of Performance & Resilience will drive performance and...PerformanceFull timeLocal areaWorldwide- ...spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in... ...are highly available, scalable, and performant. This role blends deep technical expertise... ...Design, develop, and automate reliable cloud infrastructure and platform services...PerformanceFull timeManual laborLocal areaFlexible hours
$130k - $180k per year
...we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud... ...CI/CD pipelines, enabling rapid and reliable software releases. Automate and optimize... ...of the end-to-end availability and performance of our cloud infrastructure;...PerformanceFull timeRemote workVisa sponsorshipWork visaFlexible hours$243k - $297k per year
...more resilient businesses. As Relay continues to scale, the reliability, performance, and resilience of our platform are no longer just technical... ...not only for guiding a strong team of Site Reliability Engineers, but for shaping how reliability strategy influences engineering...PerformanceLong term contractFull timeInternshipWork at officeTrial periodFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Performance & Reliability Engineer. Be the first to apply!
