Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

AI & HPC Infrastructure Engineer

Full-time

First Principles, Llc

About FirstPrinciples

FirstPrinciples is a research organization building AI infrastructure for discovery in fundamental science. Currently, our work focuses on building systems like Theo, the AI Physicist, which is a domain-specialized system for research in fundamental physics.

We’re a fast-growing, remote-first team of builders, researchers, engineers, and thinkers working across Canada, the US, the UK, and expanding globally. What brings us together is a shared curiosity about how the universe works, and a belief that we can build systems that help us explore it more effectively.

We spend our time working on questions that don’t have clear answers, like how to design AI that can reason through scientific problems, and how the scientific process as a whole might evolve. This is work that sits somewhere between creativity and rigorous thinking, and often requires comfort with ambiguity and iteration. If you’re someone who enjoys tackling big, abstract problems and building the infrastructure that makes ambitious research possible, you’ll likely find the work here interesting.

Why This Role Exists

We’re building the next generation of infrastructure for AI-driven scientific discovery, and we need someone who can help own the systems that make our research and inference workloads reliable, scalable, and fast.

This role is about building and operating the compute foundation behind our AI Physicist: Kubernetes clusters, Linux systems, GPU infrastructure, cloud environments, HPC-style compute, deployment workflows, monitoring, and automation. As our workloads grow, we need infrastructure that can support both experimentation and production-like inference across cloud, bare metal, and hybrid environments.

You’ll play a central role in shaping how we run compute at FirstPrinciples. That includes provisioning and managing clusters, improving reliability and observability, reducing operational toil, supporting researchers and engineers, and helping us make practical decisions about when to use managed cloud services, self-managed Kubernetes, Slurm-style systems, or owned hardware.

We’re looking for someone hands-on, systems-oriented, and comfortable working in a fast-moving research environment. You should have strong Kubernetes and Linux fundamentals, good operational instincts, and enough experience with cloud and HPC/GPU infrastructure to help us build toward a robust bare metal and multi-cloud inference platform.

What You’ll Do

  • Design, deploy, and operate Kubernetes infrastructure for AI inference, research, and engineering workloads

  • Set up and manage GPU and HPC-style compute environments, including scheduling, utilization, job management, and node-level troubleshooting

  • Work with systems such as Kubernetes, Slurm or similar schedulers, container runtimes, GPU drivers & libraries (ie; CUDA), storage systems, and observability tools

  • Build and manage Linux-based compute environments, including provisioning, networking, storage, monitoring, access control, and lifecycle management

  • Help architect bare metal, cloud, and hybrid infrastructure across AWS, GCP, Azure, or equivalent platforms

  • Own the reliability and operational health of infrastructure systems, including monitoring, alerting, incident response, capacity planning, and performance tuning

  • Improve deployment workflows, automation, configuration management, secrets management, and infrastructure-as-code practices

  • Partner with ML engineers, researchers, and software engineers to understand workload requirements and translate them into practical infrastructure designs

  • Evaluate tradeoffs between managed cloud services, self-managed Kubernetes, HPC schedulers, bare metal deployments, and multi-cloud architectures

  • Build tooling, documentation, runbooks, and operational practices that help the team move quickly without making infrastructure fragile or opaque

  • Balance speed and robustness, knowing when to prototype quickly and when to harden systems for long-term use

Who You Are

  • Strong infrastructure builder with experience operating production, research, cloud, or high-performance compute systems

  • Deeply comfortable with Linux administration, including debugging networking, storage, system services, permissions, performance issues, and node-level failures

  • Experienced with Kubernetes in real environments, including cluster operations, deployments, networking, observability, scaling, and troubleshooting

  • Comfortable working with cloud infrastructure on AWS, GCP, Azure, or equivalent platforms

  • Familiar with infrastructure automation and configuration tools such as Terraform, Ansible, Helm, ArgoCD, GitOps workflows, or similar systems

  • Experienced with GPU-heavy, compute-heavy, or HPC-style workloads, especially in environments involving AI, ML, research computing, or scientific workloads

  • Able to work across bare metal and cloud environments, and interested in the practical tradeoffs between the two

  • Comfortable reasoning about resource scheduling, cluster utilization, autoscaling, storage, networking, and observability for distributed workloads

  • Practical and ownership-oriented; you can take ambiguous infrastructure needs and turn them into working systems

  • Comfortable collaborating across disciplines, especially with researchers and engineers who may not think in infrastructure terms

  • Able to operate independently as a senior or strong intermediate contributor, while knowing when to bring others into important technical decisions

  • Motivated by building foundational systems that make ambitious technical and scientific work possible

Bonus

  • Hands-on experience with production-grade LLM inference and serving engines, such as vLLM, SGLang, or TensorRT

  • Experience working at an AI company, ML infrastructure team, research lab, university compute environment, HPC center, or scientific computing organization

  • Experience supporting model inference, model serving, distributed training, high-throughput batch workloads, or internal ML platforms

  • Hands-on experience with Slurm or similar HPC schedulers, including job scheduling, resource allocation, queue management, and cluster configuration

  • Experience operating GPU infrastructure, including NVIDIA drivers, CUDA, container runtimes, scheduling, utilization, and hardware failure modes

  • Experience with RDMA, InfiniBand, high-performance networking, distributed filesystems (ie. Lustre, BeeGFS), object storage, or storage systems for compute-heavy workloads

  • Experience with Kubernetes operators, custom controllers, CRDs, or platform tooling for AI/ML workloads

  • Experience with Prometheus, Grafana, Loki, OpenTelemetry, Datadog, or similar monitoring, logging, and observability tools

  • Experience with container registries, image optimization, CI/CD systems, deployment pipelines, and secure software delivery

  • Experience leading engineering operations or infrastructure efforts while remaining hands-on technically

  • Familiarity with security, access control, secrets management, and reliability practices in production or research environments

What You’ll Get

  • The opportunity to work on foundational problems at the intersection of AI and physics

  • A high-trust, low-bureaucracy environment with real ownership

  • Remote-first work with flexibility in how you structure your day

  • Exposure to cutting-edge ideas across AI, scientific discovery, infrastructure, and emerging technologies

  • A culture that values curiosity, depth of thinking, and first-principles reasoning

  • The chance to shape the compute and inference infrastructure behind advanced AI systems for scientific discovery

Vacancy posted 8 hours ago
Similar jobs that could be interesting for youBased on the AI & HPC Infrastructure Engineer in Canada vacancy
  •  ...enterprises who are building AI systems to power magical experiences...  ...is a team of researchers, engineers, designers, and more, who are...  ...Why this team? The internal infrastructure team is responsible for building...  ...Build and scale ML-optimized HPC infrastructure : Deploy and... 
    Suggested
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    Canada
    8 hours ago
  • usd80k - usd120k per year

     ...PayPal, CoreWeave , Synthesia and Mistral.ai choose Lago for their billing. We’re...  ...What You’ll Do Own and scale Lago’s infrastructure to ensure reliability, performance, and...  ...strategies. Collaborate with engineers to improve deployment processes and DevOps... 
    Suggested
    Permanent employment
    Full time
    Contract work
    Remote work
    Flexible hours

    Lago

    Canada
    8 hours ago
  • $163k - $194k per year

     ...any specified location above.  We are AI Native We are building an AI native...  ...your candidacy. About The Team The Infrastructure Team is a small, high-leverage team with...  ...daily. We run the platforms that every engineering team at Life360 depends on, including... 
    Suggested
    Full time
    Summer work
    Remote work
    Flexible hours

    Life360

    Canada
    8 hours ago
  •  ...center. Our platform combines the best of AI and human intelligence to help contact...  ...About the role: As a member of the infrastructure team you are responsible for designing,...  ...our core infrastructure that allows the engineering team to execute quickly, productively, and... 
    Suggested
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Home office

    Cresta

    Canada
    8 hours ago
  • $218.42k - $302.84k per year

     ...Job Description We’re seeking a talented software engineer, specializing in security and infrastructure, to help grow our product security team. We’re looking...  ...please contact your recruiter. Our hiring process uses AI-assisted tools to help screen and assess candidates.... 
    Suggested
    Full time
    Internship
    Work at office
    Remote work
    Home office
    Flexible hours

    Tailscale

    Canada
    8 hours ago
  • $106.25k - $125k per year

     ...providers; or helping artists use the latest AI tools and make thoughtful decisions with...  ...an invaluable role in our success. The engineering team at Warner Music Group makes all of...  ...Your role: We are seeking a Senior Infrastructure Engineer to lead our transition from... 
    Full time
    Local area
    Worldwide
    Shift work

    Warnermusic

    Canada
    8 hours ago
  • $145k - $185k per year

     ...are looking for a well-versed, passionate Engineer who wants to play a key role in site...  ...and cloud operations of our global cloud infrastructure. We’re seeking individuals with creative...  ...Washington We may use artificial intelligence (AI) tools to support parts of the hiring... 
    Full time
    Work at office
    Local area
    Immediate start
    Remote work
    Home office
    Flexible hours

    Vgs

    Canada
    8 hours ago
  •  ...ready to evolve. We're looking for true builders: engineers who orchestrate systems, ship production-grade AI, and modernize massive legacy codebases at speed....  ...adversarial testing → deploy → iterate    Architect AI Infrastructure  ~ Implement complex RAG pipelines and optimize... 
    Full time
    Internship
    Shift work

    Valsoft

    Canada
    8 hours ago
  •  ...our products are reflected in how our team works: we embrace AI as a core productivity multiplier, with all team members...  ...services of GitLab. An overview of this role As Engineering Manager, Infrastructure Platforms at GitLab, you’ll focus on building and supporting... 
    Full time

    Gitlab

    Canada
    8 hours ago
  • $104k - $123.5k per year

     .... With 465 billion automated optimizations per second, the AI-powered StackAdapt Marketing Platform seamlessly connects brand...  ...and marketing channels. As a Senior Quality Engineer in Quality Infrastructure at StackAdapt, you'll lead the design and implementation of... 
    Full time
    Local area
    Remote work
    Work from home
    Home office

    Stackadapt

    Canada
    8 hours ago
  •  ...that characterize most projects. Our Generative AI platform, Imogen, harnesses this method using advanced data engineering, compiler and LLM based techniques to create...  ...right thing. Do what works. Be kind. As an Infrastructure Software Engineer in Delivery at Mechanical Orchard... 
    Remote job
    Full time
    Local area

    Mechanical Orchard

    Canada
    8 hours ago
  • $146.25k - $195k per year

     ...residing in these provinces. This is Engineering at Lattice Lattice's Engineering team...  ...people and organizations to thrive. As AI becomes fundamental to every product...  ...Lead the development of AI evaluation infrastructure, quality metrics, experimentation capabilities... 
    Long term contract
    Full time
    Work at office
    Remote work
    Work from home

    Lattice

    Canada
    8 hours ago
  •  ...enterprise merchants, enabling businesses to ship direct-from-factory from manufacturing hubs like China to destinations worldwide. As an AI Engineer, you will own the design, development, and deployment of AI-powered systems that make our operations faster, our team smarter, and... 
    Full time
    Remote work
    Worldwide

    Portless

    Canada
    8 hours ago
  •  ...Tiger Analytics is a global leader in AI and analytics, helping Fortune 1000 companies...  ....We are looking for a highly skilled  AI Engineer with 7+ years of experience in software...  ...engineering, with a heavy focus on Python, AWS infrastructure, and Generative AI. The ideal candidate... 
    Long term contract
    Full time
    Local area

    Tiger Analytics

    Canada
    8 hours ago
  •  ...commercial real estate owners bring their portfolios to net-zero with an AI-driven sustainability, asset, and investment management platform....  ...in our category. The role We're looking for an AI Engineer to build AI-powered features, help teammates do their best work,... 
    Full time
    Flexible hours

    Cambio Ai Inc.

    Canada
    8 hours ago
  •  ...Spellbook is the most comprehensive AI copilot for transactional lawyers. It works directly inside Microsoft Word to help legal...  ...ABOUT THE ROLE Spellbook is seeking a generalist Platform and Infrastructure Engineer to join our team. In this role, you'll work across our entire... 
    Full time
    Contract work
    Flexible hours

    Spellbook

    Canada
    8 hours ago
  •  ...analytics, data warehousing, observability, and AI workloads. The company’s sustained,...  ...We are looking for exceptional backend engineers who can work across a variety of...  ...functionally to build and maintain the robust infrastructure powering ClickPipes. This role is ideal for... 
    Full time
    Local area
    Remote work
    Home office
    Flexible hours

    Clickhouse

    Canada
    8 hours ago
  •  ...Role Summary: Censys is seeking a Senior Software Engineer to join our SOC/TH team focused on AI and LLMs . The SOC-TH team builds intelligence-...  ...responders triage, identify, analyze, and monitor malicious infrastructure at internet scale. We are defining Censys as the... 
    Full time
    Remote work
    Worldwide

    Censys

    Canada
    8 hours ago
  •  ...us. We have successfully hosted over 100 engineering internships, and always have our interns...  ...this role We are looking for an Infrastructure Engineering Intern with 4-16 months of experience...  ...cutting-edge artificial intelligence (AI) technology to make our hiring process... 
    Remote job
    Full time
    Internship
    Work at office
    Local area
    Home office
    Night shift

    Supercom

    Canada
    8 hours ago
  •  ...growth.  The Mission  We’re building AI startups inside an established portfolio...  ...vertical software companies—and we need engineers who can make ideas real, fast.  Our mandate...  ..., industry expertise, existing infrastructure, and revenue to validate against—without... 
    Long term contract
    Full time
    Internship
    Work at office
    Immediate start
    Remote work

    Valsoft

    Canada
    8 hours ago
  • $150k - $180k per year

     ...· Lifelong learners with a drive to excel  · Resilient people who rise to the occasion  About the role:   We are seeking an AI/ML Engineer to join BPM’s Enterprise Technology Solutions team. This role is for a builder — someone who doesn’t just experiment with AI but makes... 
    Remote job
    Full time
    Local area
    Immediate start
    Flexible hours
    Shift work

    Bpm Llp

    Canada
    8 hours ago
  •  ...in real-time analytics, data warehousing, observability, and AI workloads. The company’s sustained, accelerating momentum...  ...Come be a part of our journey! About the Team The Cloud Infrastructure Engineering team builds and manages the foundational blocks of ClickHouse... 
    Full time
    Local area
    Remote work
    Home office
    Flexible hours

    Clickhouse

    Canada
    8 hours ago
  •  ...Staffinity is currently seeking an AI Platform Engineer - DevOps for a premier corporate client located in Halifax. This is a full-time...  ...and automated platform services. Rather than managing manual infrastructure operations, this role is explicitly focused on developing value... 
    Full time
    Work at office
    3 days per week

    Staffinity Inc.

    Canada
    8 hours ago
  •  ...delivery system for successful local restaurants and chains. About The Role:   We're looking for a rare kind of engineer: someone who thinks AI-first, ships end-to-end, and moves at a pace that makes a small team feel like a large one. You won't be handed tickets -... 
    Full time
    For contractors
    Local area
    Remote work
    Weekend work

    Ad Sauce

    Canada
    8 hours ago
  •  ...MaintainX is the world's leading AI-powered maintenance and asset management platform, serving 14,000+ customers including Duracell, Shell...  ...at the point of work. Document Intelligence is the horizontal engine that turns that raw material — any file a customer hands us — into... 
    Full time
    Immediate start

    Maintainx

    Canada
    8 hours ago
  • $123.75k - $165k per year

     ...This is Engineering at Lattice Lattice’s Engineering team is continuously working to better both our product and our craft. We use...  ...technical architecture but also an amazing product experience. The Infrastructure Platform team focuses on the cloud services, infrastructure,... 
    Full time
    Work at office
    Remote work
    Work from home

    Lattice

    Canada
    8 hours ago
  •  ...time of application Summary We are seeking an experienced Infrastructure Engineer to join our Infrastructure & DevOps team. This role is...  ...deploying, and configuring Azure infrastructure across analytics, AI, compute, storage, networking, and platform services. The successful... 
    Contract work
    Internship
    Remote work

    PLATO

    Canada
    9 days ago
  • $271k - $286k per year

     ...measurement. As Senior Staff Software Engineer on the Data Infrastructure team, you will set the technical direction...  ...built on Apache Iceberg, a multi-engine compute platform spanning stream processing...  ...and operational maturity. Pioneer AI-native Data Infrastructure Engineering... 
    Long term contract
    Permanent employment
    Full time
    Work at office
    Remote work
    Work from home
    Flexible hours

    Instacart

    Canada
    8 hours ago
  • $153k - $213k per year

     ...later without any hidden fees or compounding interest. The Engineering team builds systems that power Affirm’s mission. We take pride...  ...The Trust Infra team elevates the security posture of our infrastructure and services by embedding security in everything from provisioning... 
    Remote job
    Full time
    Work at office
    Flexible hours

    Affirm

    Canada
    8 hours ago
  • $75 per hour

     ...connects specialists with project-based AI opportunities for leading tech companies...  ...environments — a virtual company with codebase, infrastructure, and context (tickets, docs,...  ...NOT: Not data labeling; Not prompt engineering; Not cybersecurity or red-teaming — there... 
    Hourly pay
    Permanent employment
    Full time
    Temporary work
    Freelance

    Mindrift

    Canada
    8 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to AI & HPC Infrastructure Engineer. Be the first to apply!

Related searches