Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

$154k - $200k per year
Full-time

Movable Ink








Movable Ink scales content personalization for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable Ink to maximize revenue, simplify workflow and boost marketing agility. Headquartered in New York City with close to 600 employees, Movable Ink serves its global client base with operations throughout North America, Central America, Europe, Australia, and Japan.

As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic technical leadership across infrastructure and software development. You will own the design and evolution of major systems within our multi-cloud, multi-region, active-active content serving platform that serves upwards of 25 Billion requests daily. Through a combination of architectural vision, cross-team collaboration and mentorship, you will help drive the reliability initiatives and define the technical strategy that scales our platform to 50 Billion requests per day and beyond.

Responsibilities:


  • Define and drive the automation strategy for infrastructure tooling, establishing standards that minimize manual work, increase performance and reduce incident frequency and severity of incidents

  • Own the design, reliability and evolution of core platform applications, mentoring team members on best practices and ensuring systems meet long-term business objectives

  • Architect and lead the logging platform strategy, driving its design and balancing availability, retention and cost optimization

  • Establish capacity planning and performance management frameworks, proactively identifying scaling opportunities and guiding teams through complex troubleshooting scenarios

  • Lead cross-functional reliability initiatives with SRE and service engineering teams, influencing architectural decisions and championing practices that ensure resilient service delivery

  • Demonstrate a high level of autonomy in anticipating, identifying, and addressing systemic weaknesses and opportunities for platform improvement without direct supervision.

Qualifications:


  • Proven track record in Site Reliability or Software Engineering, designing, building, and owning scalable, resilient services with a focus on long-term reliability strategy

  • Deep expertise in architecting and operating complex distributed systems such as Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB/Cassandra, with the ability to guide teams through distributed system challenges

  • Designing and owning automation strategies to manage services at scale, with expertise in establishing performance analysis frameworks and mentoring others on diagnostics and resolution

  • Deep, hands-on experience (6+ years) in Site Reliability or Software Engineering, specifically leading and shaping multi-cloud architecture and strategy (AWS and GCP).

  • Experience architecting and leading large-scale observability platforms, including defining observability standards and SLO frameworks. We use Prometheus and Thanos with Grafana Alloy, Loki and Tempo

  • Experience leading on-call excellence, including driving improvements to monitoring and alerting strategies, automating runbooks and mentoring team members on incident response best practices. Every member of the SRE team does a week long on-call rotation

  • Expert-level proficiency with infrastructure as code, including defining IaC standards and patterns across teams. We use Terraform and Chef

  • Advanced Kubernetes expertise, including cluster architecture design, multi-tenancy strategies, and guiding teams on container orchestration best practices. We use EKS and GKE

  • Proficiency in multiple programming languages with the ability to design and review code that meets reliability standards. We use NodeJS, Golang, Ruby, Python and shell scripting

  • Advanced Linux systems expertise, with the ability to diagnose complex system-level issues and mentor others on performance tuning and troubleshooting

The base pay range for this position is $154,000-$200,000 CAD/year, which can include additional bonus depending on the position ultimately offered, in addition to a full range of medical, financial, and/or other benefits. The base pay offered may vary depending on job-related knowledge, skills, and experience.

Studies have shown that women, communities of color, and historically underrepresented people are less likely to apply to jobs unless they meet every single qualification. We are committed to building a diverse and inclusive culture where all Inkers can thrive. If you’re excited about the role but don’t meet all of the abovementioned qualifications, we encourage you to apply. Our differences bring a breadth of knowledge and perspectives that makes us collectively stronger.

We welcome and employ people regardless of race, color, gender identity or expression, religion, genetic information, parental or pregnancy status, national origin, sexual orientation, age, citizenship, marital status, ethnicity, family or marital status, physical and mental ability, political affiliation, disability, Veteran status, or other protected characteristics. We are proud to be an equal opportunity employer.

Vacancy posted 16 hours ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Toronto, ON vacancy
  • $140k - $180k per year

     ...into Etraveli Group’s world leading tech platform - giving partners...  ...that all other engineering teams at Tripstack depend on....  ...the deployment patterns and reliability standards that apply across both...  ...depth - interconnects, BGP, site-to-site VPN, cross-region peering... 
    Suggested
    Remplacement
    Full time
    Work at office
    Immediate start
    Flexible hours

    Tripstack

    Toronto, ON
    16 hours ago
  • $100k - $125k per year

     ...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,...  ...management, tax compliance, and treasury. Tipalti partners with leading financial institutions such as Citi, Wells Fargo, J.P.... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Tipalti

    Toronto, ON
    16 hours ago
  • $110k - $120k per year

     ...behind Money Mart—Canada’s largest non-bank branch network—and a leader in financial solutions for underserved communities. From...  ...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for ensuring... 
    Suggested
    Full time
    Temporary work
    Internship
    Work at office
    Remote work

    Momentum Financial Services Group

    Toronto, ON
    16 hours ago
  •  ...help us reinvent the way people learn, because learning never stops. The Adventure Ahead As the Manager of Site Reliability Engineering (SRE), you will lead a talented team of engineers dedicated to safeguarding the operational health and resilience of the Docebo... 
    Suggested
    Long term contract
    Full time
    For contractors
    Work at office
    Worldwide
    3 days per week

    Docebo

    Toronto, ON
    16 hours ago
  • $140k - $182k per year

     ...America, Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and...  ...on Cloud platforms (AWS/GCP) ~ Experience architecting and leading large-scale observability platforms, including defining observability... 
    Suggested
    Full time

    Movable Ink

    Toronto, ON
    16 hours ago
  • $140k - $155k per year

     ...This is a hands-on senior engineering role focused on improving production...  ...teams to ship secure, reliable, and scalable software with confidence...  ...across the organization. Lead response efforts for high-severity...  ...on cloud-native technologies, site reliability engineering... 
    Remote job
    Permanent employment
    Full time
    Flexible hours

    Caseware

    Toronto, ON
    16 hours ago
  • $153.82k - $277k per year

     ...a place where you can thrive, we can’t wait to meet you. Site Reliability Engineers (SREs) at Braze are responsible for keeping all internal-facing...  ...NGINX and Kubernetes is important. WHAT YOU'LL DO Lead NGINX & Kubernetes Ingress Infrastructure Architect and Operate... 
    Permanent employment
    Full time
    Internship
    Work at office
    Local area
    Remote work
    Flexible hours
    Rotating shift

    Braze

    Toronto, ON
    16 hours ago
  •  ...belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means…  You are an engineer, a builder, and a...  ..., and documentation over process. You’ll engage in and often lead architectural discussions, reduce toil, and deliver scalable,... 
    Full time
    Work at office
    Local area
    Remote work
    Worldwide
    Monday to friday
    Flexible hours

    Imanage

    Toronto, ON
    16 hours ago
  • $80 - $110 per hour

     ...Singapore and also operating in Denmark, Spain and Vietnam. The Site Reliability Engineer  will improve the availability, performance, scalability and...  ...device operations and routine production changes.   ~ Lead technically during incidents, drive evidence-based learning... 
    Remote job
    Full time
    Contract work

    Axon-networks

    Toronto, ON
    16 hours ago
  •  ...competitive advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in software...  ...experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly evaluate products... 
    Full time

    Serigor Inc

    Toronto, ON
    16 hours ago
  •  ...visibility, and optimize spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s...  ...standards for scalability, observability, and fault tolerance. Lead cross-functional troubleshooting of complex issues spanning... 
    Full time
    Manual labor
    Local area
    Flexible hours

    Emburse

    Toronto, ON
    16 hours ago
  •  ...AI more natural, capable, and useful. We are looking for a Site Reliability Engineer to help build and operate the infrastructure behind that work...  ...storage, scheduling, and the operational tooling that keeps them reliable. This is a hands-on role for someone who enjoys taking... 
    Full time
    Remote work

    bosonai

    Toronto, ON
    25 days ago
  • $92k - $118k per year

     ...Build reliable, resilient cloud platforms that keep critical financial services running. The Role We are looking for an experienced Site Reliability and DevOps Engineer to join the SRE team within our growing Corporate Action Processing group. You will bring 3+ years... 
    Internship
    Immediate start

    Capco

    Toronto, ON
    6 days ago
  •  ...customers. Cohere is a team of researchers, engineers, designers, and more, who are passionate...  ...building high-performance, scalable and reliable machine learning systems? Do you want to...  ...NLP applications? We are looking for a Site Reliability Engineer to join the Model... 
    Full time
    Work at office
    Remote work
    Flexible hours

    Cohere

    Toronto, ON
    16 hours ago
  • $110k - $125k per year

     ...healthcare stakeholders to collect and exchange data with our leading FHIR-based data liberation platform.   At its heart, the...  ...today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability... 
    Full time
    Remote work
    Flexible hours

    Smile Digital Health

    Toronto, ON
    16 hours ago
  • $130k - $180k per year

     ...Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure (...  ...systems remain stable and responsive even during off-hours. Lead the development, implementation, and achievement of service-... 
    Full time
    Remote work
    Visa sponsorship
    Work visa
    Flexible hours

    Acquird.io

    Toronto, ON
    16 hours ago
  • $144k - $200k per year

    **The Team** Platform Engineering is the department within SRE that is responsible for a range...  ...role in developing and maintaining the reliable and globally connected multi-cloud network...  ...Overview** We are seeking a talented Site Reliability Engineer (SRE) with a strong... 
    Full time
    Work at office
    Remote work
    Worldwide
    Flexible hours

    Mongodb

    Toronto, ON
    16 hours ago
  •  ...Employment Status: Permanent Schedule: 40 hours/week – 100% remote work Job Description We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability, performance, and resilience of high-availability SaaS platforms. Working... 
    Permanent employment
    Work at office
    Local area
    Remote work

    TOTEM Recruteur de talent

    Toronto, ON
    6 days ago
  • $116k - $235.1k per year

     ...About the Role: Site Reliability Engineering (SRE) at Tubi is not a traditional operations team. We are a software engineering organization that...  ...Participate in a 24/7 on-call rotation, acting as a key technical leader and incident commander during critical service disruptions.... 
    Long term contract
    Remplacement
    Full time
    Temporary work
    Work at office
    Local area
    Immediate start
    Flexible hours
    2 days per week

    Tubi - Canada

    Toronto, ON
    6 days ago
  •  ...youll be doing As a member of CIBCs Application Reliability Engineering Platform team the Consultant Site Reliability Engineering will play a key role in...  ...before client impact. Participate in and as required lead incident response and post-mortem analysis to identify... 
    Full time
    Contract work
    3 days per week
    1 day per week

    Canadian Imperial Bank of Commerce

    Toronto, ON
    a month ago
  • $141k - $191k per year

     ...Reuters and develop your career. As an SRE Manager, you will lead a team of 10+ engineers, oversee their development and ensure operational excellence. About the Role: In this opportunity as Site Reliability Engineering Manager , you will be responsible for: Team... 
    Work at office
    Local area
    Flexible hours
    2 days per week
    3 days per week

    Thomson Reuters

    Toronto, ON
    more than 2 months ago
  •  ...to reimagine what’s possible. Join us and help the world’s leading organizations unlock the value of technology and build a more...  ...PostgreSQL) This role is part of the Production Support and Reliability Engineering team, responsible for ensuring the stability, availability,... 
    Permanent employment
    Full time
    Local area

    Capgemini

    Toronto, ON
    a month ago
  • $120k - $170k per year

     ...Going   Magnet Forensics is a global leader in the development of digital investigative...  ...skilled and motivated Senior DevOps Engineer to join our dynamic team and play a key role...  ...incidents, and drive improvements that increase reliability and operational efficiency; Write and... 
    Full time
    Work at office
    Local area
    Flexible hours

    Magnet Forensics

    Toronto, ON
    16 hours ago
  •  ...programs. We are seeking enthusiastic, reliable, and motivated individuals to join our...  ...Program team for the 2026–2027 school year. Site Coordinators oversee the daily operations...  ...Development Certification. Experience leading sports, arts, recreation, educational, health... 
    Full time
    Contract work
    Seasonal work
    Local area
    Monday to friday

    Bgc St. Alban's Club

    Toronto, ON
    16 hours ago
  • $82.2 - $89.34 per hour

    We are seeking a highly skilled Site Reliability Specialist IV for an enterprise-level contract opportunity based in Toronto. In this role, you will take on a premier cloud engineering, platform automation, and operational reliability capacity, specializing in designing, building... 
    Long term contract
    Permanent employment
    Contract work

    Randstad

    Toronto, ON
    23 days ago
  •  ...Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance...  .... Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy...  ...products. A strong problem-solver who can lead root-cause investigations on failures... 
    Permanent employment
    Full time
    Internship
    Second job

    Tenstorrent

    Toronto, ON
    16 hours ago
  •  ...Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance...  .... Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy...  ...testing at manufacturing and test partner sites, including in Taiwan.   Tenstorrent... 
    Permanent employment
    Internship
    Second job

    Tenstorrent

    Toronto, ON
    6 days ago
  • $100k per year

     ...Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost...  ...seniorities. Tenstorrent is looking for an experienced Reliability Engineer to drive and execute reliability testing for our products from... 
    Permanent employment
    Full time
    Internship

    Tenstorrent

    Toronto, ON
    16 hours ago
  •  ...single device. This approach allows Cerebras to deliver industry-leading training and inference speeds and empowers machine learning...  .... About The Role Join Cerebras as a Performance & Reliability Engineer within our innovative Co-Design and Next Generation Team. Our... 
    Full time

    Cerebras Systems

    Toronto, ON
    16 hours ago
  •  ...Site Reliability Engineer Location: Toronto ON Work Model: Hybrid (2 days per week in-person at the Toronto office preferred) Required...  ...Role Description Dynatrace & AI-Driven Observability Lead the implementation and optimization of the Dynatrace platform... 
    Full time
    Work at office
    2 days per week

    Astra North Infoteck Inc.

    Toronto, ON
    more than 2 months ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!