Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Performance & Reliability Engineer

Full-time

Cerebras Systems

Cerebras Systems builds the world's largest AI chip, 56 times larger than GPUs. Our novel wafer-scale architecture provides the AI compute power of dozens of GPUs on a single chip, with the programming simplicity of a single device. This approach allows Cerebras to deliver industry-leading training and inference speeds and empowers machine learning users to effortlessly run large-scale ML applications, without the hassle of managing hundreds of GPUs or TPUs.  

Cerebras' current customers include top model labs, global enterprises, and cutting-edge AI-native startups.  OpenAI recently announced a multi-year partnership with Cerebras , to deploy 750 megawatts of scale, transforming key workloads with ultra high-speed inference. 

Thanks to the groundbreaking wafer-scale architecture, Cerebras Inference offers the fastest Generative AI inference solution in the world, over 10 times faster than GPU-based hyperscale cloud inference services. This order of magnitude increase in speed is transforming the user experience of AI applications, unlocking real-time iteration and increasing intelligence via additional agentic computation.

About The Role

Join Cerebras as a Performance & Reliability Engineer within our innovative Co-Design and Next Generation Team. Our groundbreaking CS-3 system has set new benchmarks in high-performance ML training and inference solutions. It leverages a dinner-plate sized chip with 44GB of on-chip memory to surpass traditional hardware capabilities. This role focuses on characterizing and optimizing the performance and reliability of state-of-the-art AI models running on Cerebras' breakthrough hardware.

 

Responsibilities


  • Characterize and enhance the performance and reliability of advanced ML hardware/software systems, with emphasis on reducing power and thermal fluctuations.

  • Analyze ML workloads, software kernels, and hardware architecture for power and performance impacts, and synthesize high-level insights across these layers.

  • Develop creative software solutions to improve reliability and performance, collaborating cross-functionally to deploy these solutions in production.

  • Influence the design of Cerebras' next-generation AI architecture and software stack through rigorous workload analysis and computational efficiency optimization.

  • Partner with ML engineers, researchers, and reliability specialists to understand model behavior and drive system-level improvements from a software perspective.

  • Collaborate with teams in architecture, silicon, and research to advance our computational platforms and influence future system designs.

Skills & Qualifications


  • BS, MS, or PhD in Computer Science, Electrical Engineering, or a related field.

  • 3+ years of relevant experience in performance engineering, reliability, computer architecture, and/or software design.

  • Proficiency in Python or other scripting languages.

  • Experience with C/C++ and assembly programming.

  • Demonstrated expertise with system-level performance and reliability optimization.

  • Strong verbal and written communication skills.

  • Nice to have: Hands-on experience with ML models, ML frameworks, and collective communication.

  • Nice to have: Understanding of thermal management principles and power delivery for advanced semiconductors.

Why Join Cerebras


People who are serious about software make their own hardware. At Cerebras we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras:


  1. Build a breakthrough AI platform beyond the constraints of the GPU.

  2. Publish and open source their cutting-edge AI research.

  3. Work on one of the fastest AI supercomputers in the world.

  4. Enjoy job stability with startup vitality.

  5. Our simple, non-corporate work culture that respects individual beliefs.

Read our blog:  Five Reasons to Join Cerebras in 2026.

Apply today and become part of the forefront of groundbreaking advancements in AI!


Cerebras Systems is committed to creating an equal and diverse environment and is proud to be an equal opportunity employer.  We celebrate different backgrounds, perspectives, and skills. We believe inclusive teams build better products and companies. We try every day to build a work environment that empowers people to do their best work through continuous learning, growth and support of those around them.

This website or its third-party tools process personal data. For more details, click here to review our CCPA disclosure notice.

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Performance & Reliability Engineer in Toronto, ON vacancy
  • $110k - $125k per year

     ...patients everyday! Apply today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability, and performance of production-grade services deployed across multiple cloud vendors and infrastructure... 
    Performance
    Full time
    Remote work
    Flexible hours

    Smile Digital Health

    Toronto, ON
    1 day ago
  •  ...unlocking real-time iteration and increasing intelligence via additional agentic computation. About The Role Join Cerebras as a Performance Engineer within our innovative Runtime Team. Our groundbreaking   CS-3 system, hosted by a distributed set of modern and powerful x86... 
    Performance
    Full time
    Local area

    Cerebras Systems

    Toronto, ON
    1 day ago
  • $115k - $125k per year

     ...Putting People First, Outstanding Corporate Citizenship, High Performance Culture, and Rigorous Financial Discipline. Mining...  ...York Stock Exchange (symbol: KGC). Job Summary The Reliability Engineer, as part of the Asset Management team, plays a vital role in... 
    Performance
    Long term contract
    Temporary work
    For contractors
    Casual work
    Local area
    Immediate start
    Remote work
    Overseas

    Kinross Gold Corporation

    Toronto, ON
    16 hours ago
  • $110k - $120k per year

     ...Discretionary Annual Bonus – Rewarding both individual and company performance Comprehensive Benefits – Health and dental coverage with...  ...programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring... 
    Performance
    Full time
    Temporary work
    Internship
    Work at office
    Remote work

    Momentum Financial Services Group

    Toronto, ON
    1 day ago
  •  ...cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost...  .... Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy...  ...of AI hardware by helping build highly reliable systems that power tomorrow's largest... 
    Performance
    Permanent employment
    Full time
    Internship
    Second job

    Tenstorrent

    Toronto, ON
    1 day ago
  • $100k - $125k per year

     ...experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability...  ...role in enhancing the reliability, performance, and scalability of our systems and...  ...Tipalti  Our tech teams are the engine behind our business. Tipalti’s tech... 
    Performance
    Full time
    Work at office
    Flexible hours

    Tipalti

    Toronto, ON
    19 hours ago
  •  ...Most Innovative Companies, and Forbes World’s Best Bank. Visit our institutional page  About the role Senior System Engineer - Systems Performance Team The Systems Performance team is part of the Computing Squad (Foundation / Runtime Platforms). You will be part of a... 
    Performance
    Long term contract
    Full time
    Remote work
    Work from home
    Relocation package
    Flexible hours

    Nubank

    Toronto, ON
    1 day ago
  • $100k per year

     ...the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining...  ...seniorities. Tenstorrent is looking for an experienced Reliability Engineer to drive and execute reliability testing for our products... 
    Performance
    Permanent employment
    Full time
    Internship

    Tenstorrent

    Toronto, ON
    1 day ago
  •  ...À titre de conseiller principal, gestion de la performance financière au sein de l’équipe Gestion de la performance financière - Entreprises de la Banque Nationale, cela signifie agir à titre de partenaire stratégique auprès des dirigeants d’entreprise et contribuer directement... 
    Performance
    Apprenticeship
    Flexible hours

    Banque Nationale

    Toronto, ON
    3 days ago
  • $154k - $200k per year

     ...America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with...  ...establishing standards that minimize manual work, increase performance and reduce incident frequency and severity of incidents Own... 
    Performance
    Long term contract
    Full time

    Movable Ink

    Toronto, ON
    1 day ago
  • $146k - $201.3k per year

     ...opportunity to do career-defining work. We're all in on this mission. If you are too, let's talk. The Engineering Opportunity Reporting to the Director of Quality & Performance, this role as Engineering Manager of Performance & Resilience will drive performance and... 
    Performance
    Full time
    Local area
    Worldwide

    Okta

    Toronto, ON
    1 day ago
  •  ...Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team...  ...: Lead, coach, and develop a high-performing team of SREs, taking ownership of hiring...  ...onboarding, mentoring, and continuous performance management. Inspire the Culture:... 
    Performance
    Long term contract
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    1 day ago
  •  ..., collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an engineer, a builder, and a systems...  ...based on industry data.   Rewarding me with an annual performance-based bonus.   Offering comprehensive Health/Vision/Dental... 
    Performance
    Full time
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to friday
    Flexible hours

    Imanage

    Toronto, ON
    1 day ago
  •  ...learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident...  ...Run structured reliability work such as capacity forecasting, performance tuning, and controlled failure testing. Improve how the org... 
    Performance
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    1 day ago
  • $153.82k - $277k per year

     ...thrive, we can’t wait to meet you. Site Reliability Engineers (SREs) at Braze are responsible for...  ...and helping translate them into reliable, highly scalable technology stacks. You...  ...Configure, tune, and operate Braze's high-performance NGINX routing, proxying, and ingress... 
    Performance
    Permanent employment
    Full time
    Internship
    Work at office
    Local area
    Remote work
    Flexible hours
    Rotating shift

    Braze

    Toronto, ON
    19 hours ago
  • $140k - $180k per year

     ...the shared infrastructure that all other engineering teams at Tripstack depend on. This...  ...sits, and the deployment patterns and reliability standards that apply across both. This...  ...acting as incident commander, managing performance, and developing engineers into senior roles... 
    Performance
    Remplacement
    Full time
    Work at office
    Immediate start
    Flexible hours

    Tripstack

    Toronto, ON
    1 day ago
  •  ...advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in software...  ...experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly evaluate products and services... 
    Performance
    Full time

    Serigor Inc

    Toronto, ON
    1 day ago
  •  ...As Chief Advisor, Financial Performance Management within the Financial Performance Management - Commercial Banking team at National Bank means acting as a strategic partner to business leaders and contributing directly to the organization’s financial objectives. In... 
    Performance
    Flexible hours

    National Bank

    Toronto, ON
    3 days ago
  • $140k - $182k per year

     ...Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and software...  ...automation of our infrastructure to minimize manual work, increase performance, and decrease the frequency and severity of incidents... 
    Performance
    Full time

    Movable Ink

    Toronto, ON
    1 day ago
  • $140k - $155k per year

     ...profiles! This is a hands-on senior engineering role focused on improving production...  ...enable engineering teams to ship secure, reliable, and scalable software with confidence....  ...reducing operational toil, improving system performance, increasing platform reliability, and... 
    Performance
    Remote job
    Permanent employment
    Full time
    Flexible hours

    Caseware

    Toronto, ON
    1 day ago
  • $165k - $220k per year

     ...terms. About This Team And Role The Mozilla Firefox Performance team is a community of engineers who care deeply about delivering the fastest browser...  ...browser as well as helping other teams write fast and reliable code to make Firefox excellent for users. Do you value... 
    Performance
    Immediate start
    Home office

    Mozilla

    Toronto, ON
    7 days ago
  •  ...spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in...  ...are highly available, scalable, and performant. This role blends deep technical expertise...  ...Design, develop, and automate reliable cloud infrastructure and platform services... 
    Performance
    Full time
    Manual labor
    Local area
    Flexible hours

    Emburse

    Toronto, ON
    1 day ago
  • $170k - $185k per year

     ...Interested in joining one of Canada’s top-performing asset managers? We are seeking a Head of Platform Engineering, Reliability & Control to lead our horizontal engineering...  ...and strengthening end-to-end data pipeline performance.   ~ Standardize developer experience by... 
    Performance
    Full time
    Work at office
    3 days per week

    Connor, Clark

    Toronto, ON
    1 day ago
  •  ...Schedule: 40 hours/week – 100% remote work Job Description We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability, performance, and resilience of high-availability SaaS platforms. Working in an AWS and Kubernetes... 
    Performance
    Permanent employment
    Work at office
    Local area
    Remote work

    TOTEM Recruteur de talent

    Toronto, ON
    8 days ago
  • $243k - $297k per year

     ...more resilient businesses. As Relay continues to scale, the reliability, performance, and resilience of our platform are no longer just technical...  ...not only for guiding a strong team of Site Reliability Engineers, but for shaping how reliability strategy influences engineering... 
    Performance
    Long term contract
    Full time
    Internship
    Work at office
    Trial period
    Flexible hours

    Relay

    Toronto, ON
    1 day ago
  • $130k - $180k per year

     ...we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud...  ...CI/CD pipelines, enabling rapid and reliable software releases. Automate and optimize...  ...of the end-to-end availability and performance of our cloud infrastructure;... 
    Performance
    Full time
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Acquird.io

    Toronto, ON
    1 day ago
  •  ...and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure...  ...-scale AI training and serving: high-performance networks, GPU clusters, storage,...  ...the operational tooling that keeps them reliable. This is a hands-on role for someone... 
    Performance
    Full time
    Remote work

    bosonai

    Toronto, ON
    3 days ago
  • $172k - $229k per year

     ...BuildOps is looking for a Staff Software Engineer to set and drive our company-wide...  ...strategy for building, shipping, and operating reliable software. This is a high-impact,...  ...validation, developer platforms, testability, performance engineering, or incident learning.... 
    Performance
    Long term contract
    Permanent employment
    Full time
    For contractors
    Work at office
    Local area
    Work from home
    Flexible hours

    Buildops

    Toronto, ON
    1 day ago
  •  ...best for our customers. Cohere is a team of researchers, engineers, designers, and more, who are passionate about their...  ...future! Why this role? Are you energized by building high-performance, scalable and reliable machine learning systems? Do you want to help define and... 
    Performance
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    Toronto, ON
    1 day ago
  •  ...What youll be doing As a member of CIBCs Application Reliability Engineering Platform team the Consultant Site Reliability Engineering...  ...most optimal for you to thrive in your role. To successfully perform the work youll be on-site full-time. Youll have the flexibility... 
    Performance
    Full time
    Contract work
    3 days per week
    1 day per week

    Canadian Imperial Bank of Commerce

    Toronto, ON
    12 days ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Performance & Reliability Engineer. Be the first to apply!