Lead Site Reliability Engineer
Movable Ink
Movable Ink scales content personalization for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable Ink to maximize revenue, simplify workflow and boost marketing agility. Headquartered in New York City with close to 600 employees, Movable Ink serves its global client base with operations throughout North America, Central America, Europe, Australia, and Japan.
As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic technical leadership across infrastructure and software development. You will own the design and evolution of major systems within our multi-cloud, multi-region, active-active content serving platform that serves upwards of 25 Billion requests daily. Through a combination of architectural vision, cross-team collaboration and mentorship, you will help drive the reliability initiatives and define the technical strategy that scales our platform to 50 Billion requests per day and beyond.
Responsibilities:
- Define and drive the automation strategy for infrastructure tooling, establishing standards that minimize manual work, increase performance and reduce incident frequency and severity of incidents
- Own the design, reliability and evolution of core platform applications, mentoring team members on best practices and ensuring systems meet long-term business objectives
- Architect and lead the logging platform strategy, driving its design and balancing availability, retention and cost optimization
- Establish capacity planning and performance management frameworks, proactively identifying scaling opportunities and guiding teams through complex troubleshooting scenarios
- Lead cross-functional reliability initiatives with SRE and service engineering teams, influencing architectural decisions and championing practices that ensure resilient service delivery
- Demonstrate a high level of autonomy in anticipating, identifying, and addressing systemic weaknesses and opportunities for platform improvement without direct supervision.
Qualifications:
- Proven track record in Site Reliability or Software Engineering, designing, building, and owning scalable, resilient services with a focus on long-term reliability strategy
- Deep expertise in architecting and operating complex distributed systems such as Apache Pulsar, Apache Kafka, Grafana Loki, ScyllaDB/Cassandra, with the ability to guide teams through distributed system challenges
- Designing and owning automation strategies to manage services at scale, with expertise in establishing performance analysis frameworks and mentoring others on diagnostics and resolution
- Deep, hands-on experience (6+ years) in Site Reliability or Software Engineering, specifically leading and shaping multi-cloud architecture and strategy (AWS and GCP).
- Experience architecting and leading large-scale observability platforms, including defining observability standards and SLO frameworks. We use Prometheus and Thanos with Grafana Alloy, Loki and Tempo
- Experience leading on-call excellence, including driving improvements to monitoring and alerting strategies, automating runbooks and mentoring team members on incident response best practices. Every member of the SRE team does a week long on-call rotation
- Expert-level proficiency with infrastructure as code, including defining IaC standards and patterns across teams. We use Terraform and Chef
- Advanced Kubernetes expertise, including cluster architecture design, multi-tenancy strategies, and guiding teams on container orchestration best practices. We use EKS and GKE
- Proficiency in multiple programming languages with the ability to design and review code that meets reliability standards. We use NodeJS, Golang, Ruby, Python and shell scripting
- Advanced Linux systems expertise, with the ability to diagnose complex system-level issues and mentor others on performance tuning and troubleshooting
Studies have shown that women, communities of color, and historically underrepresented people are less likely to apply to jobs unless they meet every single qualification. We are committed to building a diverse and inclusive culture where all Inkers can thrive. If you’re excited about the role but don’t meet all of the abovementioned qualifications, we encourage you to apply. Our differences bring a breadth of knowledge and perspectives that makes us collectively stronger.
We welcome and employ people regardless of race, color, gender identity or expression, religion, genetic information, parental or pregnancy status, national origin, sexual orientation, age, citizenship, marital status, ethnicity, family or marital status, physical and mental ability, political affiliation, disability, Veteran status, or other protected characteristics. We are proud to be an equal opportunity employer.
$101.2k - $136.9k per year
...member facing services. We are looking for a highly skilled Site Reliability Engineer (SRE) who will help build, operate and continuously improve the... ...and improve runbooks and self-service capabilities. ~ Lead or contribute to incident response and follow-through with...SuggestedPermanent employmentFull timeInternshipWork at officeImmediate startHome officeFlexible hours2 days per week- ...on to find out more. The role: We’re looking for a Site Reliability Engineer to manage, maintain, improve and provide support on our platform... ..., you will have the opportunity to work with an industry leading distributed team. Having access to expertise from across the...SuggestedFull timeRemote work
$20 per day
...as one of Canada’s fastest-growing companies and backed by leading U.S. investors, Hiive is profitable, well-capitalized, and building... ...careers page to see how you can grow with us! As a Site Reliability Engineer at Hiive, you will be responsible for ensuring the...SuggestedFull timeSummer holidayRelocation- ...MaintainX is the world's leading Asset and Work Intelligence platform for industrial... ...modern, IoT-enabled, cloud-based tool for reliability, safety, and operations of physical... ...$2.5 billion. We’re looking for a Site Reliability Engineer (SRE) to help advance MaintainX’s...SuggestedFull time
- ...growing innovator offering supply chain solutions to industry leading healthcare systems, hospitals, and pharmacy businesses to... ...good fit for you! About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a...SuggestedLong term contractPermanent employmentFull timeWork at officeRemote work
- ...trillion in AUM and 22 global investment banks. For more information, please visit . The Role CMG is looking for a Site Reliability Engineer (SRE) with a strong focus on monitoring, observability, and alerting to ensure the reliability, performance, and...Remote jobFull timeLocal area
$120k - $200k per year
...and many more. ABOUT THE ROLE At LayerZero, our Site Reliability Engineering (SRE) team is at the intersection of software and systems engineering... ...internal systems to those external users interact with—are reliable, meet the uptime expectations of our users, and...Full time$110k - $160k per year
...hear from you! Role Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering.... ...; Participate in on-call incident response rotation; lead or support incident command during active production incidents...Full timeInternshipWork at officeLocal areaFlexible hours$123k - $160k per year
...your place here. We are seeking a highly experienced Senior Site Reliability Engineer to own the reliability, performance, and operational excellence of our large-scale, distributed infrastructure. You will lead design and execution of systems that power mission critical...Remote jobLong term contractFull timeRelocation- ...Gauss Labs is seeking a highly skilled Site Reliability Engineer to join our team in Vancouver. As an SRE at Gauss Labs, you will play a critical... ...Incident Response: Participating in on-call rotations and leading incident response efforts to minimize downtime and restore service...Full time
- ...tiptoeing into it. We are rebuilding our engineering culture around a simple belief: AI changes... ...this role is at the center of keeping it reliable, fast, and scalable. As a Staff SRE, you'... ...right reasons Incident response — lead or support incident response, drive post-...Remote jobFull timeInternshipWork at officeLocal areaFlexible hoursShift workWeekend work
$120k - $160k per year
...ABOUT YOU We are looking for a Site Reliability Engineer (Monetization) who is pragmatic, product-minded, and equally comfortable writing... ...reliability roadmap for the domain together with product engineering leads Participate in product team planning, refinements, and...Long term contractFull time- ...might just be in the right place! We’re looking for a Staff Site Reliability Engineer to join our Data team in Canada. As a Staff Data SRE,... ...improvements. You write production-quality IaC (Terraform), lead solutions from design through delivery, and are the person the...Full timeWork at officeRemote workFlexible hoursShift work
- ...We are seeking a Senior DevOps & Site Reliability Engineer to own the reliability, scalability, performance, and operational excellence of Medeloop... ...CloudWatch, Sentry) covering metrics, logs, traces, and alerting. Lead incident response: triage, mitigate, and drive blameless post...Hourly payFull time
$180.4k - $230.4k per year
...impact with bold thinking are real—and happening daily at Coalition. About the role We are looking for a Staff Site Reliability Engineer to lead AI enablement across our engineering organization. As AI-assisted development reshapes how software gets built, a new platform...Full timeRemote workHome officeFlexible hoursShift work$145k - $185k per year
...makes simulation at scale possible. We're hiring a Senior Site Reliability Engineer to help build and operate that infrastructure. This role sits... ...: severity definitions, escalation paths, on-call practices. Lead incident response, debugging, and root-cause analysis. Write...Remote jobFull time$197.5k - $225k per year
...GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization... ...observability — define SLOs, alerts, and dashboards. Lead incident response and postmortems, focusing on root cause and...Full time$140k - $180k per year
...About Windscribe Windscribe is a leading cyber security and privacy company launched in April 2016 and now with more than... ...and online privacy for all. Right now we are looking for a Site Reliability Engineer to help us tame DNS. About the Position Linux system...Full timeDirect hireWork at officeRemote work$69k - $90k per year
...UN-endorsed Zero Project. About the role As a Junior Site Reliability Engineer (SRE) at Fable, you will help support the reliability, performance... ..., co-ops, or personal projects Our values To lead, listen first You amplify voices that are less often heard...Full timeInternship- ...name a few. About the role We’re looking for a Senior Site Reliability Engineer (SRE) to help strengthen and scale our multi-cloud platform... ...developer tooling that removes friction across engineering Lead rollouts of AI-native tooling for code review, testing, and engineering...Full timeInternshipRemote workWork from home
- ...over 250 percent year over year, ClickHouse leads the market in real-time analytics, data warehousing... ...committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building...Full timeLocal areaRemote workHome officeFlexible hours
- ...visualizing relationships between entities in the system. As a Site Reliability Engineer you will be responsible for the availability, latency,... ...Monitor, develop and troubleshoot applications to resolve issues, lead incident support, be part of the on-call team Automate...Full timeFlexible hours
$153k - $187k per year
...resilient businesses. We’re looking for an incredible Senior Site Reliability Engineer to join our SRE team. We aim to make reliability, security,... ...to clarify operational responsibilities Collaborate in leading incident response by driving fast mitigation, clear...Full timeInternship- ...The Site Reliability Engineering organization at Pinterest is accountable for ensuring overall Pinterest availability as well as enhancing Engineering teams’ capability to design, build and operate robust systems at scale. Pinterest’s applications and infrastructure that...Full timeWork at officeRelocationRelocation package
- ...in 2024 – but we're just getting started. As a Sr. Site Reliability Engineer, you'll be the guardian of our platform's reliability and performance... .... What You Bring to the Team: Design and implement reliable and scalable AWS architecture to meet the needs of the...Full timeWork at officeLocal areaRemote workWork from homeHome officeWeekend work
- ...you’re relentless about the high quality, reliability, and security our customers demand. You... ...technical deliverables will reach the entire engineering organization to enable product teams to... ...confidence. Exemplify cloud-native site reliability best practices with a strong...Long term contractFull timeRemote workFlexible hours
$260k - $275k per year
...enterprises • Solve complex reliability challenges at scale • Influence architecture and engineering culture at a company level •... ...will focus on creating reusable, reliable, and scalable solutions that abstract... ..., Platform Engineering, or Site Reliability Engineering role,...Full time$150k - $240k per year
...obsessed about achieving the high quality and reliability our customers demand. You will work... ...technical deliverables will reach the entire engineering organization to enable product teams to... ...-effective. Exemplify cloud-native site reliability best practices. Write code...Long term contractFull timeRemote work- ...une équipe dynamique en tant qu’ Ingénieur·e Fiabilité de Site (Site Reliability Engineer) pour l’un de nos clients. Le Site Reliability Engineering... ...systems and working with high scale scalable and reliable services. Like to work in a fast-moving environment and...Full timeApprenticeshipLocal areaWorldwide
- ...The Co-op Refinery Complex (CRC) is hiring a Lead Electrical Reliability Engineer on a permanent basis to work on-site in Regina, SK. Are you ready to take your engineering career to the next level and make a direct impact on the reliability and performance of assets at...Permanent employmentTemporary workWork from home
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Lead Site Reliability Engineer. Be the first to apply!
- lead process engineer Remote
- lead software engineer Remote
- lead structural engineer Remote
- site reliability engineer intern Remote
- site reliability engineer remote Remote
- site reliability engineer sre Remote
- site reliability engineer Remote
- senior site reliability engineer Remote
- website developer Remote
- site safety Remote
