Lead Site Reliability Engineer
$154k - $200k per yearMovable Ink
Movable Ink scales content personalization for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable Ink to maximize revenue, simplify workflow and boost marketing agility. Headquartered in New York City with close to 600 employees, Movable Ink serves its global client base with operations throughout North America, Central America, Europe, Australia, and Japan.
As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic technical leadership across infrastructure and software development. You will own the design and evolution of major systems within our multi-cloud, multi-region, active-active content serving platform that serves upwards of 25 Billion requests daily. Through a combination of architectural vision, cross-team collaboration and mentorship, you will help drive the reliability initiatives and define the technical strategy that scales our platform to 50 Billion requests per day and beyond.
Responsibilities:
- Define and drive the automation strategy for infrastructure tooling, establishing standards that minimize manual work, increase performance and reduce incident frequency and severity of incidents
- Own the design, reliability and evolution of core platform applications, mentoring team members on best practices and ensuring systems meet long-term business objectives
- Architect and lead the logging platform strategy, driving its design and balancing availability, retention and cost optimization
- Establish capacity planning and performance management frameworks, proactively identifying scaling opportunities and guiding teams through complex troubleshooting scenarios
- Lead cross-functional reliability initiatives with SRE and service engineering teams, influencing architectural decisions and championing practices that ensure resilient service delivery
- Demonstrate a high level of autonomy in anticipating, identifying, and addressing systemic weaknesses and opportunities for platform improvement without direct supervision.
Qualifications:
- Proven track record in Site Reliability or Software Engineering, designing, building, and owning scalable, resilient services with a focus on long-term reliability strategy
- Deep expertise in architecting and operating complex distributed systems such as Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB/Cassandra, with the ability to guide teams through distributed system challenges
- Designing and owning automation strategies to manage services at scale, with expertise in establishing performance analysis frameworks and mentoring others on diagnostics and resolution
- Deep, hands-on experience (6+ years) in Site Reliability or Software Engineering, specifically leading and shaping multi-cloud architecture and strategy (AWS and GCP).
- Experience architecting and leading large-scale observability platforms, including defining observability standards and SLO frameworks. We use Prometheus and Thanos with Grafana Alloy, Loki and Tempo
- Experience leading on-call excellence, including driving improvements to monitoring and alerting strategies, automating runbooks and mentoring team members on incident response best practices. Every member of the SRE team does a week long on-call rotation
- Expert-level proficiency with infrastructure as code, including defining IaC standards and patterns across teams. We use Terraform and Chef
- Advanced Kubernetes expertise, including cluster architecture design, multi-tenancy strategies, and guiding teams on container orchestration best practices. We use EKS and GKE
- Proficiency in multiple programming languages with the ability to design and review code that meets reliability standards. We use NodeJS, Golang, Ruby, Python and shell scripting
- Advanced Linux systems expertise, with the ability to diagnose complex system-level issues and mentor others on performance tuning and troubleshooting
The base pay range for this position is $154,000-$200,000 CAD/year, which can include additional bonus depending on the position ultimately offered, in addition to a full range of medical, financial, and/or other benefits. The base pay offered may vary depending on job-related knowledge, skills, and experience.
Studies have shown that women, communities of color, and historically underrepresented people are less likely to apply to jobs unless they meet every single qualification. We are committed to building a diverse and inclusive culture where all Inkers can thrive. If you’re excited about the role but don’t meet all of the abovementioned qualifications, we encourage you to apply. Our differences bring a breadth of knowledge and perspectives that makes us collectively stronger.
We welcome and employ people regardless of race, color, gender identity or expression, religion, genetic information, parental or pregnancy status, national origin, sexual orientation, age, citizenship, marital status, ethnicity, family or marital status, physical and mental ability, political affiliation, disability, Veteran status, or other protected characteristics. We are proud to be an equal opportunity employer.
$104.24k - $143.3k per year
...Join some of the most innovative thinkers in FinTech as we lead the evolution of financial technology. If you are an innovative... ...at an innovative and growing company. Our Lead Site Reliability Engineers will provide a stable infrastructure platform throughout our...SuggestedFull timeWork at officeImmediate startRemote workShift workWeekend work2 days per week- ...Windscribe is a leading cyber security and privacy company launched in April 2016 and now with more than 70 million users. We... ...and online privacy for all. Right now we are looking for a Site Reliability Engineer to help us tame DNS. This position can be hybrid at our Toronto...SuggestedFull timeDirect hireWork at officeRemote work
$115k - $130k per year
...About Kaseya Kaseya is the leading provider of AI-powered IT management and cybersecurity software, serving Managed Service... ..., and continuously raising the bar. Kaseya is hiring a Site Reliability Engineer to keep our production systems healthy as we scale. You'll own...SuggestedLong term contractFull timeWorldwide- ...growing innovator offering supply chain solutions to industry leading healthcare systems, hospitals, and pharmacy businesses to... ...good fit for you! About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a...SuggestedRemote jobLong term contractPermanent employmentFull time
$110k - $120k per year
...behind Money Mart—Canada’s largest non-bank branch network—and a leader in financial solutions for underserved communities. From... ...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring...SuggestedFull timeTemporary workInternshipWork at officeRemote work$100k - $125k per year
...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,... ...management, tax compliance, and treasury. Tipalti partners with leading financial institutions such as Citi, Wells Fargo, J.P....Full timeWork at officeFlexible hours$100k per year
...Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations... .... This role sits at the intersection of site reliability, infrastructure operations, and customer engineering, ensuring our systems are reliable, observable, and...Permanent employmentFull time- ...world and help us reinvent the way people learn, because learning never stops. Role Overview As a Senior Site Reliability Engineer, you'll take a hands-on lead role in high severity incident response while also shaping the underlying infrastructure that supports the...Full timeFor contractorsWork at officeWorldwide3 days per week
$140k - $182k per year
...America, Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and... ...on Cloud platforms (AWS/GCP) ~ Experience architecting and leading large-scale observability platforms, including defining observability...Full time- ...help us reinvent the way people learn, because learning never stops. The Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team of engineers dedicated to safeguarding the operational health and resilience of the Docebo...Long term contractFull timeFor contractorsWork at officeWorldwide3 days per week
- ...belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a... ..., and documentation over process. You’ll engage in and often lead architectural discussions, reduce toil, and deliver scalable,...Full timeWork at officeLocal areaRemote workWorldwideMonday to fridayFlexible hours
- ...competitive advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in software... ...experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly evaluate products...Full time
- ...large scale, multi-tenant SaaS platform running reliably for customers around the world. You'll take a hands-on lead role in high severity incident response while also... ...to raise the reliability bar across the wider engineering org. What You'll Be Doing Take point on...Full timeImmediate startFlexible hours
- ...About the Role Fivetran is looking for a high-performance engineer to be a part of a team of Site Reliability Engineers. You will be working closely with engineering teams, product managers, as well as support and sales engineers to build the future of the Fivetran Data...Full timeWork at officeRemote work
- ...and how we use AI in our recruiting process here . The Site Reliability Engineering organization at Pinterest is accountable for ensuring... ...Demonstrated ability to write effective prompts to get high-quality, reliable outputs from LLMs ~ Demonstrated ability to use AI to...Full timeWork at office
- ...visibility, and optimize spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s... ...standards for scalability, observability, and fault tolerance. Lead cross-functional troubleshooting of complex issues spanning...Full timeManual laborLocal areaFlexible hours
- ...Group, a division of Fox Corporation. About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a... ...seeking an experienced and visionary Senior SRE Manager to lead and grow our newly built Site Reliability Engineering team....RemplacementFull timeContract workTemporary workFlexible hours
$130k - $180k per year
...Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure (... ...systems remain stable and responsive even during off-hours. Lead the development, implementation, and achievement of service-...Full timeRemote workVisa sponsorshipWork visaFlexible hours$243k - $297k per year
...businesses. As Relay continues to scale, the reliability, performance, and resilience of our... ...not only for guiding a strong team of Site Reliability Engineers, but for shaping how reliability... ...meet you! What You'll Be Doing Lead and evolve Relay’s Site Reliability Engineering...Long term contractFull timeInternshipWork at officeTrial periodFlexible hours$110k - $125k per year
...healthcare stakeholders to collect and exchange data with our leading FHIR-based data liberation platform. At its heart, the... ...today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability...Full timeRemote workFlexible hours- ...Senior Site Reliability Engineer - Edge Location : Ottawa/Toronto, On-Site Reports to: Head of Security The Role You own the edge compute module — the standard hardware stack, OS image, and runtime that integrates with Dominion Dynamics's mesh radios, sensors...Full time
- ...customers. Cohere is a team of researchers, engineers, designers, and more, who are passionate... ...building high-performance, scalable and reliable machine learning systems? Do you want to... ...NLP applications? We are looking for a Site Reliability Engineer to join the Model...Full timeWork at officeRemote workFlexible hours
$144k - $200k per year
**The Team** Platform Engineering is the department within SRE that is responsible for a range... ...role in developing and maintaining the reliable and globally connected multi-cloud network... ...Overview** We are seeking a talented Site Reliability Engineer (SRE) with a strong...Full timeWork at officeRemote workWorldwideFlexible hours$136k - $187k per year
...users worldwide. Our commitment to reliability is a key foundation of our product and... ...availability expectations is a core engineering focus. As a Senior Site Reliability Engineer, you'll join... ...solutions that make our system more reliable by design. What you’ll do: Design...Local areaRemote workWorldwide- ...interest and capabilities in others. About the Role: As a Site Reliability Engineer , you’ll join the global Platform SRE team responsible for... ...of operational playbooks and postmortem practices. Lead and contribute to scaling initiatives that improve elasticity...Full time
$150k - $250k per year
...About The Role We're looking for a Senior Site Reliability Engineer to help us run one of the most exciting GPU clusters around—our Toronto datacenter packed with NVIDIA H100 and A100 GPUs, over 20PB of Ceph storage, terabit networking, and hundreds of servers. You'll...Full time- ...Platform Engineer – DevOps, Site Reliability Engineering (SRE) & Dynatrace Required Skills • Strong experience as a Platform Engineer with expertise... ...when required. Observability & Monitoring • Lead the implementation and administration of enterprise...Permanent employment
$141k - $191k per year
...Reuters and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence. About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for: Team...Work at officeLocal areaFlexible hours2 days per week3 days per week- ...Site Reliability Engineer Location: Toronto ON Work Model: Hybrid (2 days per week in-person at the Toronto office preferred) Required... ...Role Description Dynatrace & AI-Driven Observability Lead the implementation and optimization of the Dynatrace platform...Full timeWork at office2 days per week
- ...SLOs, shape capacity plans, and ensure the reliability, durability, and operational safety of... .... We are a small team of software engineers with a strong bias towards software solutions... ...candidates may also have experience with: Leading major architectural shifts, such as...Long term contractFull timeWork at officeLocal areaImmediate startRemote workWorldwideShift work
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead process engineer Toronto, ON
- lead software engineer Toronto, ON
- lead structural engineer Toronto, ON
- site reliability engineer intern Toronto, ON
- site reliability engineer remote Toronto, ON
- site reliability engineer Toronto, ON
- senior site reliability engineer Toronto, ON
- website developer Toronto, ON
- site safety Toronto, ON
- site maintenance Toronto, ON
