Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Performance Reliability Engineer

Full-time

Cerebras Systems

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. Our novel wafer-scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programming simplicity of a single device. This approach allows Cerebras to deliver industry-leading training and inference speeds and empowers machine learning users to effortlessly run large-scale ML applications, without the hassle of managing hundreds of GPUs or TPUs.  

Cerebras' current customers include global corporations across multiple industries, national labs, and top-tier healthcare systems. In January, we announced a multi-year, multi-million-dollar partnership with Mayo Clinic, underscoring our commitment to transforming AI applications across various fields. In August, we launched Cerebras Inference, the fastest Generative AI inference solution in the world, over 10 times faster than GPU-based hyperscale cloud inference services.

About The Role

Join Cerebras as a Performance & Reliability Engineer within our innovative Co-Design and Next Generation Team. Our groundbreaking CS-3 system has set new benchmarks in high-performance ML training and inference solutions. It leverages a dinner-plate sized chip with 44GB of on-chip memory to surpass traditional hardware capabilities. This role focuses on characterizing and optimizing the performance and reliability of state-of-the-art AI models running on Cerebras' breakthrough hardware.

 

Responsibilities


  • Characterize and enhance the performance and reliability of advanced ML hardware/software systems, with emphasis on reducing power and thermal fluctuations.

  • Analyze ML workloads, software kernels, and hardware architecture for power and performance impacts, and synthesize high-level insights across these layers.

  • Develop creative software solutions to improve reliability and performance, collaborating cross-functionally to deploy these solutions in production.

  • Influence the design of Cerebras' next-generation AI architecture and software stack through rigorous workload analysis and computational efficiency optimization.

  • Partner with ML engineers, researchers, and reliability specialists to understand model behavior and drive system-level improvements from a software perspective.

  • Collaborate with teams in architecture, silicon, and research to advance our computational platforms and influence future system designs.

Skills & Qualifications


  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field.

  • 3+ years of relevant experience in performance engineering, reliability, computer architecture, and/or software design.

  • Proficiency in Python or other scripting languages.

  • Experience with C/C++ and assembly programming.

  • Demonstrated expertise with system-level performance and reliability optimization.

  • Strong verbal and written communication skills.

  • Nice to have: Hands-on experience with ML models, ML frameworks, and collective communication.

  • Nice to have: Understanding of thermal management principles and power delivery for advanced semiconductors.

Why Join Cerebras


People who are serious about software make their own hardware. At Cerebras we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:


  1. Build a breakthrough AI platform beyond the constraints of the GPU.

  2. Publish and open source their cutting-edge AI research.

  3. Work on one of the fastest AI supercomputers in the world.

  4. Enjoy job stability with startup vitality.

  5. Our simple, non-corporate work culture that respects individual beliefs.

Read our blog:  Five Reasons to Join Cerebras in 2025.

Apply today and become part of the forefront of groundbreaking advancements in AI!


Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer.  We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Vacancy posted 5 hours ago
Similar jobs that could be interesting for youBased on the Performance Reliability Engineer in Toronto, ON vacancy
  • $110k - $125k per year

     ...patients everyday! Apply today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure... 
    Performance
    Full time
    Remote work
    Flexible hours

    Smile Digital Health

    Toronto, ON
    5 hours ago
  •  ...unlocking real-time iteration and increasing intelligence via additional agentic computation. About The Role Join Cerebras as a Performance Engineer within our innovative Runtime Team. Our groundbreaking   CS-3 system, hosted by a distributed set of modern and powerful x86... 
    Performance
    Full time
    Local area

    Cerebras Systems

    Toronto, ON
    5 hours ago
  • $115k - $130k per year

     ...and results. We are a high-growth, high-performance organization that values individuals...  ...the bar.  Kaseya is hiring a Site Reliability Engineer to keep our production systems healthy...  ...and error budgets that keep our systems reliable Lead incident response,... 
    Performance
    Long term contract
    Full time
    Worldwide

    Kaseya Careers

    Toronto, ON
    5 hours ago
  • $110k - $120k per year

     ...Discretionary Annual Bonus – Rewarding both individual and company performance Comprehensive Benefits – Health and dental coverage with...  ...programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring... 
    Performance
    Full time
    Temporary work
    Internship
    Work at office
    Remote work

    Momentum Financial Services Group

    Toronto, ON
    5 hours ago
  •  ...cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost...  .... Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy...  ...of AI hardware by helping build highly reliable systems that power tomorrow's largest... 
    Performance
    Permanent employment
    Full time
    Internship
    Second job

    Tenstorrent

    Toronto, ON
    5 hours ago
  •  ...fit for you! About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a...  ...will help maintain, optimize, and ensure the reliability and performance of the systems that power our cloud infrastructure across AWS... 
    Performance
    Remote job
    Long term contract
    Permanent employment
    Full time

    Tecsys Inc.

    Toronto, ON
    5 hours ago
  •  ...the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining...  ...of all seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next generation... 
    Performance
    Permanent employment
    Internship
    Second job

    Tenstorrent

    Toronto, ON
    5 days ago
  • $100k per year

     ...the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining...  ...seniorities. Tenstorrent is looking for an experienced Reliability Engineer to drive and execute reliability testing for our products... 
    Performance
    Permanent employment
    Full time
    Internship

    Tenstorrent

    Toronto, ON
    5 hours ago
  •  ...The primary objective of the Database Reliability Engineer r is to provide expertise across...  ...they operate efficiently, securely, and reliable within private and public cloud environments...  ...up and restoring data, monitoring performance, and optimizing database operations... 
    Performance
    Long term contract
    Full time
    Internship
    Rotating shift

    Orion Health

    Toronto, ON
    5 hours ago
  •  ...Most Innovative Companies, and Forbes World’s Best Bank. Visit our institutional page  About the role Senior System Engineer - Systems Performance Team The Systems Performance team is part of the Computing Squad (Foundation / Runtime Platforms). You will be part of a... 
    Performance
    Long term contract
    Full time
    Remote work
    Work from home
    Relocation package
    Flexible hours

    Nubank

    Toronto, ON
    5 hours ago
  •  ...scale, multi-tenant SaaS platform running reliably for customers around the world. You'll...  ...raise the reliability bar across the wider engineering org. What You'll Be Doing Take...  ...work such as capacity forecasting, performance tuning, and controlled failure testing.... 
    Performance
    Full time
    Immediate start
    Flexible hours

    HighlightTA

    Toronto, ON
    5 hours ago
  •  ...recruiting process here . The Site Reliability Engineering organization at Pinterest is...  ...of innovation Manage capacity and performance to help scale our infrastructure both...  ...effective prompts to get high-quality, reliable outputs from LLMs ~ Demonstrated ability... 
    Performance
    Full time
    Work at office

    Pinterest

    Toronto, ON
    5 hours ago
  •  ...learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident...  ...Run structured reliability work such as capacity forecasting, performance tuning, and controlled failure testing. Improve how the org... 
    Performance
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    5 hours ago
  • $154k - $200k per year

     ...America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with...  ...establishing standards that minimize manual work, increase performance and reduce incident frequency and severity of incidents Own... 
    Performance
    Long term contract
    Full time

    Movable Ink

    Toronto, ON
    5 hours ago
  •  ...About the Role Fivetran is looking for a high-performance engineer to be a part of a team of Site Reliability Engineers. You will be working closely with engineering teams, product managers, as well as support and sales engineers to build the future of the Fivetran Data... 
    Performance
    Full time
    Work at office
    Remote work

    Fivetran

    Toronto, ON
    5 hours ago
  •  ...Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team...  ...: Lead, coach, and develop a high-performing team of SREs, taking ownership of hiring...  ...onboarding, mentoring, and continuous performance management. Inspire the Culture:... 
    Performance
    Long term contract
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    5 hours ago
  • $100k - $120k per year

     ...and expand self-service analytics, we need a platform engineer to own the reliability, performance, and operational excellence of the systems that make it...  ...Prometheus, Grafana, or Datadog.   ~ Expertise in MPP/OLAP engines, specifically Apache Doris or alternatives like... 
    Performance
    Long term contract
    Full time
    Work at office

    Rumble Inc.

    Toronto, ON
    5 hours ago
  •  ..., collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an engineer, a builder, and a systems...  ...based on industry data.   Rewarding me with an annual performance-based bonus.   Offering comprehensive Health/Vision/Dental... 
    Performance
    Full time
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to friday
    Flexible hours

    Imanage

    Toronto, ON
    5 hours ago
  •  ...the world and our goal is preserving uncensored Internet access and online privacy for all. Right now we are looking for a Site Reliability Engineer to help us tame DNS. This position can be hybrid at our Toronto offices or remote in Canada only but you MUST reside in the... 
    Full time
    Direct hire
    Work at office
    Remote work

    Funded.club

    Toronto, ON
    5 hours ago
  • $140k - $182k per year

     ...Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and software...  ...automation of our infrastructure to minimize manual work, increase performance, and decrease the frequency and severity of incidents... 
    Performance
    Full time

    Movable Ink

    Toronto, ON
    5 hours ago
  • $100k - $125k per year

     ...experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability...  ...role in enhancing the reliability, performance, and scalability of our systems and...  ...Tipalti  Our tech teams are the engine behind our business. Tipalti’s tech... 
    Performance
    Full time
    Work at office
    Flexible hours

    Tipalti

    Toronto, ON
    5 hours ago
  • $100k per year

     ...cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost...  ...This role sits at the intersection of site reliability, infrastructure operations, and customer engineering, ensuring our systems are reliable, observable, and production-ready.... 
    Performance
    Permanent employment
    Full time

    Tenstorrent

    Toronto, ON
    5 hours ago
  •  ...advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in software...  ...experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly evaluate products and services... 
    Performance
    Full time

    Serigor Inc

    Toronto, ON
    5 hours ago
  • $146k - $201.3k per year

     ...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity Reporting to the Director of Quality & Performance, this role as Engineering Manager of Performance & Resilience will drive performance and... 
    Performance
    Full time
    Local area
    Worldwide

    Okta

    Toronto, ON
    5 hours ago
  • $136k - $187k per year

     ...worldwide. Our commitment to reliability is a key foundation of our...  ...availability expectations is a core engineering focus. As a Senior Site...  ...that make our system more reliable by design. What you’ll do...  ...improving the availability, performance, and observability of our services... 
    Performance
    Local area
    Remote work
    Worldwide

    Okta

    Toronto, ON
    5 hours ago
  •  ...a division of Fox Corporation. About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are...  ...have a stellar experience. We own the availability, latency, performance, and capacity of our platform, and we achieve our goals through... 
    Performance
    Remplacement
    Full time
    Contract work
    Temporary work
    Flexible hours

    Tubi

    Toronto, ON
    5 hours ago
  •  ...spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in...  ...are highly available, scalable, and performant. This role blends deep technical expertise...  ...Design, develop, and automate reliable cloud infrastructure and platform services... 
    Performance
    Full time
    Manual labor
    Local area
    Flexible hours

    Emburse

    Toronto, ON
    5 hours ago
  • $170k - $185k per year

     ...Interested in joining one of Canada’s top-performing asset managers? We are seeking a Head of Platform Engineering, Reliability & Control to lead our horizontal engineering...  ...and strengthening end-to-end data pipeline performance.   ~ Standardize developer experience by... 
    Performance
    Full time
    Work at office
    3 days per week

    Connor, Clark

    Toronto, ON
    5 hours ago
  • $243k - $297k per year

     ...more resilient businesses. As Relay continues to scale, the reliability, performance, and resilience of our platform are no longer just technical...  ...not only for guiding a strong team of Site Reliability Engineers, but for shaping how reliability strategy influences engineering... 
    Performance
    Long term contract
    Full time
    Internship
    Work at office
    Trial period
    Flexible hours

    Relay

    Toronto, ON
    5 hours ago
  • $130k - $180k per year

     ...we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud...  ...CI/CD pipelines, enabling rapid and reliable software releases. Automate and optimize...  ...of the end-to-end availability and performance of our cloud infrastructure;... 
    Performance
    Full time
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Acquird.io

    Toronto, ON
    5 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Performance Reliability Engineer. Be the first to apply!