Site Reliability Engineer, AI Observability
Appnovation Technologies
About us
Appnovation is a global, full-service digital partner that combines Strategy, Experience & Design, Engineering and Managed Services. We build digital solutions that deliver real impact today and serve as foundations for future growth. Bold ambition. Practical action. Endless possibilities.
We’re looking for a Site Reliability Engineer to keep a shared observability platform for LLM-based applications running for a global life sciences client. The platform is built on Langfuse and self-hosted on Kubernetes on AWS, with ClickHouse as the analytical store, PostgreSQL for metadata and Redis as the ingestion queue, all delivered through Argo CD.
Two things need building rather than maintaining. The platform has no monitoring, alerting or defined service levels today, and you will own putting them in place. Infrastructure is defined as code throughout, in Kubernetes manifests, Helm values and Argo CD applications.
Alongside the platform itself, you will own the runbooks that make an incident survivable by someone other than the author, and the onboarding and support path for the internal teams that depend on the platform.
ROLE RESPONSIBILITIES
- Platform Operations: Diagnose and resolve failures across ClickHouse, PostgreSQL, Redis, the ingestion workers and the Kubernetes layer beneath them, including ingest backpressure from queue depth and worker drain behaviour.
- ClickHouse Operations: Own ClickHouse under the operator model, including Keeper quorum, replication, shard and replica topology, and S3 storage tiering.
- Provable Backup and Restore: Rehearse the restore, time it, document it and test it against its failure modes, rather than assuming a successful backup job means a recoverable system.
- Safe Upgrades: Plan and rehearse upgrades in a lower environment, with a rollback plan that works even if a migration is only partly complete.
- Monitoring and Service Levels: Build monitoring, alerting and service levels from scratch, so problems are found here before a user reports them.
- Infrastructure as Code: Maintain Kubernetes manifests, Helm values and Argo CD applications so every change goes through the delivery pipeline.
- Runbooks and SOPs: Write and maintain runbooks and SOPs that let a colleague resolve an incident without the author present.
- Onboarding and Support: Run the onboarding and support path for internal teams that depend on the platform, and triage what they bring.
- Automation: Turn recurring operational work into automation.
- Upgrade Partnership: Work with the platform engineer who owns what the platform offers. They decide what to adopt and how it is configured; you own the migration and its rollback.
QUALIFICATIONS
- Hands-on experience with Kubernetes on AWS (managed EKS), with routine work done through Helm values and Argo CD applications.
- Hands-on experience running ClickHouse in production, including replication and Keeper quorum, shard and replica topology, and backup and restore, ideally run through an operator.
- Experience building monitoring and alerting from scratch, including service levels that reflect what users actually experience rather than what is easy to measure.
- Experience upgrading self-hosted software safely, including schema migrations, rehearsal in a lower environment and a rollback plan for a partly completed migration.
- PostgreSQL and Redis operations deep enough to debug metadata-store and queue problems, including backpressure and worker drain.
- Strong operational writing: runbooks, SOPs and post-incident reviews that a colleague can follow unaided during an incident.
PREFERRED QUALIFICATIONS
- Observability engineering, including OpenTelemetry Collector pipelines, alerting design and Grafana dashboards.
- OIDC or enterprise SSO integration with a corporate identity provider.
- GitHub Actions for plan and apply pipelines with approval gates.
- Experience running LLM observability tools such as Langfuse, LangSmith or Arize Phoenix.
- Experience in pharma, life sciences or another regulated industry.
WHO YOU ARE
- You don’t trust a backup until you’ve restored from it
- You want to find problems before users do
- You write runbooks for the person on call at 3am, not for yourself
- You stay calm in incidents and focus on fixing the process afterward
- You automate anything you have to do twice
- You work well inside a client team and build trust quickly
- You have prior experience in consulting
- Prior experience and connections in the Life Sciences industry is preferred
Thank you for your interest in a career with Appnovation Technologies! Please note that only those selected for an interview will be contacted.
At Appnovation, we recognize that diverse teams are the strongest teams. Diversity, Equity & Inclusion is not only something that we embrace - we celebrate it! We are proud to be an Equal Opportunity Employer and we encourage applicants from all backgrounds, lived experiences and industries to apply. Come join us at Appnovation, and learn more about how we stay true to our company values as we build better lives through better digital. Accommodations are available upon request throughout the recruitment process.
$110k - $130k per year
...others, and it defines our culture. About the job As a Site Reliability Engineer II on the Serving Platforms team within Infrastructure Engineering... ...Experience operating, scaling, and monitoring AI/ML or LLM-powered services and workloads in high-concurrency...SuggestedFull timeWork at officeLocal areaWorldwideFlexible hours2 days per week$150k - $185k per year
...just be in the right place! We’re looking for a Manager, Site Reliability Engineering to lead 2 teams of 4-6 platform engineers with a wide scope... ...Lead the design and implementation of scalable, reliable, and efficient infrastructure solutions for our global cloud...SuggestedLong term contractFull timeRemote workFlexible hours$110k - $120k per year
...recognition programs that celebrate your impact The Job: Site Reliability Engineer The Site Reliability Engineer is responsible for... ...UAT, production), ensuring changes are safe, repeatable, and observable. Design and maintain automated CI/CD pipelines and enforce...SuggestedFull timeTemporary workInternshipWork at officeRemote work$100k - $125k per year
...We are seeking an experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site Reliability Engineer,... ...in real time. Why join Tipalti? Tipalti is the AI-powered platform for finance automation, elevating how finance...SuggestedFull timeWork at officeFlexible hours$153.82k - $277k per year
...place where you can thrive, we can’t wait to meet you. Site Reliability Engineers (SREs) at Braze are responsible for keeping all internal-facing... ...stack, including Redis, Kafka, Postgres, or MongoDB Observability Systems: Experience using monitoring, profiling, and...SuggestedPermanent employmentFull timeInternshipWork at officeLocal areaRemote workFlexible hoursRotating shift$140k - $155k per year
...This is a hands-on senior engineering role focused on improving production... ...teams to ship secure, reliable, and scalable software with confidence... ...objectives (SLOs), improve observability, automate operational... ...on cloud-native technologies, site reliability engineering principles...Remote jobPermanent employmentFull timeFlexible hours$140k - $182k per year
...for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable... ..., Europe, Australia, and Japan. As one of our Senior Site Reliability Engineers, you will be 100% hands-on across infrastructure and software...Full time$154k - $200k per year
...for marketers through data-activated content generation and AI decisioning. The world’s most innovative brands rely on Movable... ...America, Europe, Australia, and Japan. As one of our Lead Site Reliability Engineers, you will combine hands-on technical expertise with strategic...Long term contractFull time$80 - $110 per hour
...AXON Networks delivers a robust AI-driven, analytics-based orchestration platform and a wide portfolio of next-gen high... ...Singapore and also operating in Denmark, Spain and Vietnam. The Site Reliability Engineer will improve the availability, performance, scalability and...Remote jobFull timeContract work- ...advantage. Job Description The SRE Role · SREs are engineers with the right mix of knowledge and skills in... ...applying scientific principles of experimentation and observation to entire systems to improve reliability, performance and operability). · We constantly...Full time
- ...belonging, collaboration, and accomplishment. Being a Senior Site Reliability Engineer at iManage Means… You are an engineer, a builder, and a... ...in on-call rotations. You’ll be a key voice in observability, change management, and service scalability, providing guidance...Full timeWork at officeLocal areaRemote workWorldwideMonday to fridayFlexible hours
$144k - $200k per year
**The Team** Platform Engineering is the department within SRE that... ...infrastructure, deployment machinery, and observability and alerting systems. The... ...and maintaining the reliable and globally connected multi-... ...We are seeking a talented Site Reliability Engineer (SRE) with...Full timeWork at officeRemote workWorldwideFlexible hours- ...and enterprises who are building AI systems to power magical... ...Cohere is a team of researchers, engineers, designers, and more, who are passionate... ...high-performance, scalable and reliable machine learning systems? Do... ...? We are looking for a Site Reliability Engineer to join the...Full timeWork at officeRemote workFlexible hours
$110k - $125k per year
...today and find plenty of reasons to SMILE! The Cloud Site Reliability Engineer (SRE) is responsible for ensuring the reliability, scalability... ..., integrates them into CI/CD pipelines, and uses observability tools to proactively detect and resolve bottlenecks. Working...Full timeRemote workFlexible hours- ...drives business value and inspires extraordinary results. Our AI-powered platform helps organizations modernize financial... ...visibility, and optimize spend across the enterprise. The Site Reliability Engineer III (SRE III) plays a critical role in ensuring Emburse’s systems...Full timeManual laborLocal areaFlexible hours
$130k - $180k per year
...Our Platform is growing and we are looking to hire a Senior Site Reliability Engineer (SRE) / Cloud Engineer Our main Cloud Platform is Azure... ...implement, and maintain CI/CD pipelines, enabling rapid and reliable software releases. Automate and optimize our infrastructure...Full timeRemote workVisa sponsorshipWork visaFlexible hours- ...combines Strategy, Experience & Design, Engineering and Managed Services. We build digital... ...build the OpenTelemetry instrumentation, observability integrations, tracing, SLOs, alerting,... ...~6+ years in SRE, platform reliability or observability engineering, with strong...Full time
$110k - $130k per year
...others, and it defines our culture. About the job As a Site Reliability Engineer II on the Serving Platforms team within Infrastructure Engineering... ...Experience operating, scaling, and monitoring AI/ML or LLM-powered services and workloads in high-concurrency...Work at officeLocal areaWorldwideFlexible hours2 days per week- ...work Job Description We are looking for an experienced Site Reliability Engineer to join a team responsible for the reliability,... ...improve scalability and resilience. Define and evolve observability practices, including dashboards, alerts, SLOs, and SLIs....Permanent employmentWork at officeLocal areaRemote work
$120k - $170k per year
...We are seeking a highly skilled and motivated Senior DevOps Engineer to join our dynamic team and play a key role in designing, implementing... ..., investigate incidents, and drive improvements that increase reliability and operational efficiency; Write and maintain automation,...Full timeWork at officeLocal areaFlexible hours$180k - $260k per year
...About the Role We’re looking for a AI Deployment Engineer with 3–6 years of experience to work... ...systems and continuously improve their reliability, performance, and level of autonomy.... ...locations and work closely with teams on-site. What We’re Looking For ~3–6 years...Long term contractFull timeTemporary workWork at office2 days per week3 days per week$135k - $210k per year
...Overview: Guidepoint seeks an experienced AI Engineer as an integral member of the Toronto-... ...like LLM Judges or MLflow, AI observability, and system monitoring. Evaluate and... ...# Live Coding & System Design Test (On-site, 2 hours) # Technical Leadership Interview...Full timeWorldwide$150k - $160k per year
...Requisition Number: 105712 AI Engineer Location You will have the flexibility to work fully remotely. Insight at a Glance... ...prompt strategies, guardrails, and evaluation loops to improve reliability, safety, and quality. Be AmbITious: This opportunity is not...Full timeLocal areaRemote work$130k - $150k per year
...engagement worldwide, and we're looking for an Intermediate Site Reliability Engineer to join our Engineering team. About the job We're looking... ...in code reviews, document changes, and support application, AI, and data engineering teams. About you Around 3–5 years...Full timeWork at officeRemote workWorldwide1 day per week- ...CoreFactor is searching for a Senior AI Engineer on a permanent/full-time basis. This position is hybrid and will require the successful... ...design, including evaluation methodology, guardrails, and reliability of LLM-powered systems. Drive continuous iteration and improvement...Permanent employmentFull timeWork at office
$135k - $170k per year
...General Information: Job Title: AI Engineer Location: Toronto, ON (Onsite/Hybrid... ...Optimize systems for cost, latency, and reliability Collaborate across teams where needed... ...experience (vLLM, TGI, llama.cpp) Observability tooling (Langfuse, LangSmith) Prior...Long term contractFull timeInternshipWork at officeImmediate startRemote workFlexible hours- ...is leading the industry on cutting-edge AI technology, revolutionizing performance... ...seniorities. Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy... ...of AI hardware by helping build highly reliable systems that power tomorrow's largest AI...Permanent employmentFull timeInternshipSecond job
- ...About the Role SJC Media is building the AI layer for its Digital Operations. As the founding Software Engineer, Applied AI, you own that layer end-to-end: from... ...within budget, and design agent systems that are reliable, secure, and auditable. Senior engineering ability...Full timeFlexible hours
- ...Nexxa is building the best AI systems for heavy industries — enabling machines... ...We're looking for Backend AI Engineers to design, build, and own the core AI... ...existing infrastructure. Own the reliability, performance, and observability of backend AI systems — logging, monitoring...Long term contractFull time
$122.3k - $170.7k per year
...everyone makes play happen. Production Infrastructure & Engineering (PI&E) organization provides the essential platforms and infrastructure... ...for players where and when they want to play. As a Site Reliability Engineer, your role covers the entire lifecycle of a product-...Full timeLocal area
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Site Reliability Engineer, AI Observability. Be the first to apply!
- senior site reliability engineer Toronto, ON
- site reliability engineer Toronto, ON
- site reliability engineer intern Toronto, ON
- website developer Toronto, ON
- site maintenance Toronto, ON
- site safety Toronto, ON
- site carpenter Toronto, ON
- senior site reliability engineer
- site reliability engineer
- site reliability engineer intern


