Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Lead Site Reliability Engineer

Full-time

Movable Ink









Movable Ink scales content personalization for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable Ink to maximize revenue, simplify workflow and boost marketing agility. Headquartered in New York City with close to 600 employees, Movable Ink serves its global client base with operations throughout North America, Central America, Europe, Australia, and Japan.

As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic technical leadership across infrastructure and software development. You will own the design and evolution of major systems within our multi-cloud, multi-region, active-active content serving platform that serves upwards of 25 Billion requests daily. Through a combination of architectural vision, cross-team collaboration and mentorship, you will help drive the reliability initiatives and define the technical strategy that scales our platform to 50 Billion requests per day and beyond.

Responsibilities:


  • Define and drive the automation strategy for infrastructure tooling, establishing standards that minimize manual work, increase performance and reduce incident frequency and severity of incidents

  • Own the design, reliability and evolution of core platform applications, mentoring team members on best practices and ensuring systems meet long-term business objectives

  • Architect and lead the logging platform strategy, driving its design and balancing availability, retention and cost optimization

  • Establish capacity planning and performance management frameworks, proactively identifying scaling opportunities and guiding teams through complex troubleshooting scenarios

  • Lead cross-functional reliability initiatives with SRE and service engineering teams, influencing architectural decisions and championing practices that ensure resilient service delivery

  • Demonstrate a high level of autonomy in anticipating, identifying, and addressing systemic weaknesses and opportunities for platform improvement without direct supervision.

Qualifications:


  • Proven track record in Site Reliability or Software Engineering, designing, building, and owning scalable, resilient services with a focus on long-term reliability strategy

  • Deep expertise in architecting and operating complex distributed systems such as Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB/Cassandra, with the ability to guide teams through distributed system challenges

  • Designing and owning automation strategies to manage services at scale, with expertise in establishing performance analysis frameworks and mentoring others on diagnostics and resolution

  • Deep, hands-on experience (6+ years) in Site Reliability or Software Engineering, specifically leading and shaping multi-cloud architecture and strategy (AWS and GCP).

  • Experience architecting and leading large-scale observability platforms, including defining observability standards and SLO frameworks. We use Prometheus and Thanos with Grafana Alloy, Loki and Tempo

  • Experience leading on-call excellence, including driving improvements to monitoring and alerting strategies, automating runbooks and mentoring team members on incident response best practices. Every member of the SRE team does a week long on-call rotation

  • Expert-level proficiency with infrastructure as code, including defining IaC standards and patterns across teams. We use Terraform and Chef

  • Advanced Kubernetes expertise, including cluster architecture design, multi-tenancy strategies, and guiding teams on container orchestration best practices. We use EKS and GKE

  • Proficiency in multiple programming languages with the ability to design and review code that meets reliability standards. We use NodeJS, Golang, Ruby, Python and shell scripting

  • Advanced Linux systems expertise, with the ability to diagnose complex system-level issues and mentor others on performance tuning and troubleshooting

Studies have shown that women, communities of color, and historically underrepresented people are less likely to apply to jobs unless they meet every single qualification. We are committed to building a diverse and inclusive culture where all Inkers can thrive. If you’re excited about the role but don’t meet all of the abovementioned qualifications, we encourage you to apply. Our differences bring a breadth of knowledge and perspectives that makes us collectively stronger.

We welcome and employ people regardless of race, color, gender identity or expression, religion, genetic information, parental or pregnancy status, national origin, sexual orientation, age, citizenship, marital status, ethnicity, family or marital status, physical and mental ability, political affiliation, disability, Veteran status, or other protected characteristics. We are proud to be an equal opportunity employer.

Vacancy posted 7 hours ago
Similar jobs that could be interesting for youBased on the Lead Site Reliability Engineer in Remote vacancy
  • $101.2k - $136.9k per year

     ...member facing services. We are looking for a highly skilled Site Reliability Engineer (SRE) who will help build, operate and continuously improve the...  ...and improve runbooks and self-service capabilities.   ~ Lead or contribute to incident response and follow-through with... 
    Suggested
    Permanent employment
    Full time
    Internship
    Work at office
    Immediate start
    Home office
    Flexible hours
    2 days per week

    Vancity

    Remote
    7 hours ago
  •  ...on to find out more. The role: We’re looking for a Site Reliability Engineer to manage, maintain, improve and provide support on our platform...  ..., you will have the opportunity to work with an industry leading distributed team. Having access to expertise from across the... 
    Suggested
    Full time
    Remote work

    Tyk Technologies Limited

    Remote
    7 hours ago
  • $20 per day

     ...as one of Canada’s fastest-growing companies and backed by leading U.S. investors, Hiive is profitable, well-capitalized, and building...  ...careers page to see how you can grow with us! As a Site Reliability Engineer at Hiive, you will be responsible for ensuring the... 
    Suggested
    Full time
    Summer holiday
    Relocation

    Hiive

    Remote
    7 hours ago
  •  ...MaintainX is the world's leading Asset and Work Intelligence platform for industrial...  ...modern, IoT-enabled, cloud-based tool for reliability, safety, and operations of physical...  ...$2.5 billion. We’re looking for a Site Reliability Engineer (SRE) to help advance MaintainX’s... 
    Suggested
    Full time

    Maintainx

    Remote
    7 hours ago
  •  ...growing innovator offering supply chain solutions to industry leading healthcare systems, hospitals, and pharmacy businesses to...  ...good fit for you! About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a... 
    Suggested
    Long term contract
    Permanent employment
    Full time
    Work at office
    Remote work

    Tecsys Inc.

    Remote
    7 hours ago
  •  ...trillion in AUM and 22 global investment banks. For more information, please visit .     The Role     CMG is looking for a Site Reliability Engineer (SRE) with a strong focus on monitoring, observability, and alerting to ensure the reliability, performance, and... 
    Remote job
    Full time
    Local area

    Capital Markets Gateway

    Remote
    7 hours ago
  • $120k - $200k per year

     ...and many more.   ABOUT THE ROLE At LayerZero, our Site Reliability Engineering (SRE) team is at the intersection of software and systems engineering...  ...internal systems to those external users interact with—are reliable, meet the uptime expectations of our users, and... 
    Full time

    Layer Zero Labs Llc

    Remote
    7 hours ago
  • $110k - $160k per year

     ...hear from you!  Role Overview  We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering....  ...; Participate in on-call incident response rotation; lead or support incident command during active production incidents... 
    Full time
    Internship
    Work at office
    Local area
    Flexible hours

    Magnet Forensics

    Remote
    7 hours ago
  • $123k - $160k per year

     ...your place here. We are seeking a highly experienced Senior Site Reliability Engineer to own the reliability, performance, and operational excellence of our large-scale, distributed infrastructure. You will lead design and execution of systems that power mission critical... 
    Remote job
    Long term contract
    Full time
    Relocation

    Branch Metrics

    Remote
    7 hours ago
  •  ...Gauss Labs is seeking a highly skilled Site Reliability Engineer to join our team in Vancouver. As an SRE at Gauss Labs, you will play a critical...  ...Incident Response: Participating in on-call rotations and leading incident response efforts to minimize downtime and restore service... 
    Full time

    Gauss Labs

    Remote
    7 hours ago
  •  ...tiptoeing into it. We are rebuilding our engineering culture around a simple belief: AI changes...  ...this role is at the center of keeping it reliable, fast, and scalable. As a Staff SRE, you'...  ...right reasons Incident response — lead or support incident response, drive post-... 
    Remote job
    Full time
    Internship
    Work at office
    Local area
    Flexible hours
    Shift work
    Weekend work

    Babylist, Inc

    Remote
    7 hours ago
  • $120k - $160k per year

     ...ABOUT YOU We are looking for a Site Reliability Engineer (Monetization) who is pragmatic, product-minded, and equally comfortable writing...  ...reliability roadmap for the domain together with product engineering leads Participate in product team planning, refinements, and... 
    Long term contract
    Full time

    Xsolla

    Remote
    7 hours ago
  •  ...might just be in the right place! We’re looking for a Staff Site Reliability Engineer to join our Data team in Canada. As a Staff Data SRE,...  ...improvements. You write production-quality IaC (Terraform), lead solutions from design through delivery, and are the person the... 
    Full time
    Work at office
    Remote work
    Flexible hours
    Shift work

    Lightspeed Commerce

    Remote
    7 hours ago
  •  ...We are seeking a Senior DevOps & Site Reliability Engineer to own the reliability, scalability, performance, and operational excellence of Medeloop...  ...CloudWatch, Sentry) covering metrics, logs, traces, and alerting. Lead incident response: triage, mitigate, and drive blameless post... 
    Hourly pay
    Full time

    Medeloop

    Remote
    7 hours ago
  • $180.4k - $230.4k per year

     ...impact with bold thinking are real—and happening daily at Coalition. About the role We are looking for a Staff Site Reliability Engineer to lead AI enablement across our engineering organization. As AI-assisted development reshapes how software gets built, a new platform... 
    Full time
    Remote work
    Home office
    Flexible hours
    Shift work

    Coalition

    Remote
    7 hours ago
  • $145k - $185k per year

     ...makes simulation at scale possible. We're hiring a Senior Site Reliability Engineer to help build and operate that infrastructure. This role sits...  ...: severity definitions, escalation paths, on-call practices. Lead incident response, debugging, and root-cause analysis. Write... 
    Remote job
    Full time

    Parallel Domain

    Remote
    7 hours ago
  • $197.5k - $225k per year

     ...GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization...  ...observability — define SLOs, alerts, and dashboards. Lead incident response and postmortems, focusing on root cause and... 
    Full time

    Securityscorecard

    Remote
    7 hours ago
  • $140k - $180k per year

     ...About Windscribe Windscribe is a leading cyber security and privacy company launched in April 2016 and now with more than...  ...and online privacy for all. Right now we are looking for a Site Reliability Engineer to help us tame DNS. About the Position Linux system... 
    Full time
    Direct hire
    Work at office
    Remote work

    Funded.club

    Remote
    7 hours ago
  • $69k - $90k per year

     ...UN-endorsed Zero Project.  About the role  As a Junior Site Reliability Engineer (SRE) at Fable, you will help support the reliability, performance...  ..., co-ops, or personal projects Our values   To lead, listen first  You amplify voices that are less often heard... 
    Full time
    Internship

    Fable

    Remote
    7 hours ago
  •  ...name a few.  About the role We’re looking for a Senior Site Reliability Engineer (SRE) to help strengthen and scale our multi-cloud platform...  ...developer tooling that removes friction across engineering Lead rollouts of AI-native tooling for code review, testing, and engineering... 
    Full time
    Internship
    Remote work
    Work from home

    Scalepad

    Remote
    7 hours ago
  •  ...over 250 percent year over year, ClickHouse leads the market in real-time analytics, data warehousing...  ...committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building... 
    Full time
    Local area
    Remote work
    Home office
    Flexible hours

    Clickhouse

    Remote
    7 hours ago
  •  ...visualizing relationships between entities in the system. As a Site Reliability Engineer you will be responsible for the availability, latency,...  ...Monitor, develop and troubleshoot applications to resolve issues, lead incident support, be part of the on-call team Automate... 
    Full time
    Flexible hours

    Behavox

    Remote
    7 hours ago
  • $153k - $187k per year

     ...resilient businesses. We’re looking for an incredible Senior Site Reliability Engineer to join our SRE team. We aim to make reliability, security,...  ...to clarify operational responsibilities Collaborate in leading incident response by driving fast mitigation, clear... 
    Full time
    Internship

    Relay

    Remote
    7 hours ago
  •  ...The Site Reliability Engineering organization at Pinterest is accountable for ensuring overall Pinterest availability as well as enhancing Engineering teams’ capability to design, build and operate robust systems at scale. Pinterest’s applications and infrastructure that... 
    Full time
    Work at office
    Relocation
    Relocation package

    Pinterest

    Remote
    7 hours ago
  •  ...in 2024 – but we're just getting started.      As a Sr. Site Reliability Engineer, you'll be the guardian of our platform's reliability and performance...  .... What You Bring to the Team: Design and implement reliable and scalable AWS architecture to meet the needs of the... 
    Full time
    Work at office
    Local area
    Remote work
    Work from home
    Home office
    Weekend work

    Third-Party Job Posts

    Remote
    7 hours ago
  •  ...you’re relentless about the high quality, reliability, and security our customers demand. You...  ...technical deliverables will reach the entire engineering organization to enable product teams to...  ...confidence. Exemplify cloud-native site reliability best practices with a strong... 
    Long term contract
    Full time
    Remote work
    Flexible hours

    Axon

    Remote
    7 hours ago
  • $260k - $275k per year

     ...enterprises • Solve complex reliability challenges at scale • Influence architecture and engineering culture at a company level •...  ...will focus on creating reusable, reliable, and scalable solutions that abstract...  ..., Platform Engineering, or Site Reliability Engineering role,... 
    Full time

    Saviynt

    Remote
    7 hours ago
  • $150k - $240k per year

     ...obsessed about achieving the high quality and reliability our customers demand. You will work...  ...technical deliverables will reach the entire engineering organization to enable product teams to...  ...-effective. Exemplify cloud-native site reliability best practices. Write code... 
    Long term contract
    Full time
    Remote work

    Axon

    Remote
    7 hours ago
  •  ...une équipe dynamique en tant qu’ Ingénieur·e Fiabilité de Site (Site Reliability Engineer) pour l’un de nos clients. Le Site Reliability Engineering...  ...systems and working with high scale scalable and reliable services. Like to work in a fast-moving environment and... 
    Full time
    Apprenticeship
    Local area
    Worldwide

    Mthree Recruiting Portal

    Remote
    7 hours ago
  •  ...The Co-op Refinery Complex (CRC) is hiring a Lead Electrical Reliability Engineer on a permanent basis to work on-site in Regina, SK. Are you ready to take your engineering career to the next level and make a direct impact on the reliability and performance of assets at... 
    Permanent employment
    Temporary work
    Work from home

    FCL

    Remote
    7 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!