Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Application Reliability Engineer

$80k - $150k per year
Full-time

Innodata Inc.

Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.

Scope of the Role: 

We are looking for a hands-on Application Support Engineer to support, maintain, and enhance business-critical enterprise applications built and running on Google App Engine (GAE) and microservices.

The role is focused on application availability, production support, incident response, troubleshooting, and continuous feature enhancement rather than building a new application from the ground up. The ideal candidate can quickly understand an existing microservices-based application landscape, restore service when users are impacted, and deliver incremental enhancements safely across test, pre-production, and production environments.

Experience supporting large-scale, business-critical applications in complex enterprise technology environments is preferred.

What You’ll Own:


  • Provide production support and maintenance for enterprise applications hosted on Google App Engine, including standard and flexible environments.

  • Act as a first point of contact for user-impacting incidents: triage, diagnose, restore service, and drive issues to closure within agreed SLAs/SLOs.

  • Troubleshoot application errors, failed requests, latency and performance degradation, service-to-service failures, configuration issues, quota/scaling limits, and dependency or integration failures.

  • Design, develop, and deliver feature enhancements and functional improvements to existing applications based on user and business needs.

  • Support applications built on microservices architecture, including service boundaries, APIs/contracts, inter-service communication, authentication, and failure/retry behavior.

  • Own build, release, and deployment activities across development, test, pre-production, and production environments with appropriate validation, approvals, and rollback plans.

  • Manage App Engine deployments including versions, traffic splitting/migration, canary and staged rollouts, rollbacks, service configuration, and scaling settings.

  • Perform root-cause analysis for recurring production issues and implement sustainable fixes.

  • Build and maintain monitoring, alerting, logging, dashboards, and error reporting using Cloud Monitoring, Cloud Logging, Error Reporting, and Cloud Trace.

  • Support platform, framework, library, dependency, and runtime upgrades while maintaining stability, supportability, and compliance.

  • Support IAM, service accounts, access controls, secrets management, and operational governance.

  • Participate in change management, release-readiness reviews, and on-call/rotational support as required.

  • Collaborate with client and cross-functional Application Engineering, Product, QA, Data Engineering, Infrastructure, and Platform teams.

  • Adapt to established client-specific engineering, security, review, change-management, and operational processes.

  • Create and maintain technical documentation, operational runbooks, troubleshooting guides, deployment procedures, and support playbooks.

Microservices Expectations: 


  • Service decomposition and ownership: understand service boundaries, upstream/downstream dependencies, and ownership of behavior or data.

  • APIs and contracts: REST and gRPC/protobuf interfaces, versioning, backward compatibility, and contract-change impacts.

  • Inter-service communication: synchronous calls, asynchronous/event-driven messaging (Pub/Sub, Cloud Tasks), idempotency, retries, timeouts, backoff, and circuit breaking.

  • Service-to-service authentication and authorization using service accounts, identity/tokens, and least-privilege access.

  • Distributed troubleshooting using logs, distributed tracing, and correlation IDs to isolate failures across services.

  • Failure modes at scale, including cascading failures, partial outages, hot spots, quota exhaustion, and graceful degradation.

  • Independent deployability and coordination of multi-service releases when required.

  • Service-level observability, including dashboards, SLIs/SLOs, alerting, and error budgets.

You’ll Thrive in This Role If You Have:


  • 3-7 years of experience in Application Support, Application Engineering, Software Engineering, Cloud Engineering, or a related role.

  • Strong hands-on experience supporting, maintaining, and enhancing production applications on Google Cloud Platform (GCP).

  • Hands-on experience with Google App Engine, including deployment, configuration, scaling, versioning, and troubleshooting.

  • Solid understanding of microservices architecture, REST/gRPC APIs, service-to-service communication and authentication, distributed tracing, and cross-service debugging.

  • Practical experience with deployments and promotions across multiple environments, including release validation, rollback, and change control.

  • Strong programming skills in one or more of Python, Java, Node.js / JavaScript, or Go.

  • Working knowledge of SQL and application data stores such as Cloud SQL, Firestore/Datastore, Cloud Spanner, or BigQuery.

  • Good understanding of GCP IAM, service accounts, permissions, monitoring, logging, alerting, and production operations.

  • Experience with CI/CD pipelines and automated build/deployment tooling such as Cloud Build, Jenkins, GitHub Actions, or GitLab CI.

  • Experience troubleshooting complex production environments and performing root-cause analysis under time pressure.

  • Ability to quickly understand existing systems, codebases, services, configurations, and client-specific tools and workflows.

  • Strong communication and collaboration skills, including communication during user-impacting incidents.

Preferred Qualifications: 


  • Experience supporting large-scale internal or enterprise applications with demanding availability, reliability, and performance requirements.

  • Ability to quickly learn and operate within enterprise-specific application frameworks, deployment tooling, and support processes.

  • Cloud Run, GKE, Cloud Functions, or other GCP application services.

  • Apigee / API Gateway, load balancing, and API management.

  • Pub/Sub, Cloud Tasks, Cloud Scheduler, and asynchronous/event-driven patterns.

  • Infrastructure as Code, particularly Terraform.

  • SRE practices including SLIs/SLOs, error budgets, incident management, and blameless postmortems.

  • Containerization with Docker and Kubernetes fundamentals.

  • Frontend or full-stack experience supporting user-facing web applications.

  • Looker / Tableau / BI platforms or reporting integrations.

  • Supporting AI/ML, GenAI, or LLM-based applications running on GCP, including Vertex AI.

The expected salary range for this position is $80,000 - $150,000 CAD per year, based on experience, skills, and qualifications.

 

 

Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at  

If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at  View email address on jobs.jobcopilot.com and consider reporting it to the FTC at  ReportFraud.ftc.gov .

Vacancy posted 1 day ago
Similar jobs that could be interesting for youBased on the Application Reliability Engineer in Remote vacancy
  •  ...includes our scientific research & engineering division (Skynet Software)...  ...Avondale is looking for a Reliability Engineer with a background...  ...-performance, scalable, and reliable web systems. ~ We also...  ...Software welcomes and encourages applications from people with... 
    Suggested
    Full time

    Chelsea Avondale

    Remote
    1 day ago
  •  ...fit for you! About the Role We are looking for a Site Reliability Engineer to join our Network and Security Operations Center (NOC), a...  ...to deliver exceptional solutions. We welcome and encourage applicants from all backgrounds, experiences, and perspectives to join... 
    Suggested
    Long term contract
    Permanent employment
    Full time
    Work at office
    Remote work

    Tecsys Inc.

    Remote
    1 day ago
  •  ...to find out more. The role: We’re looking for a Site Reliability Engineer to manage, maintain, improve and provide support on our platform...  ...employer and we are determined to ensure that no applicant or employee receives less favourable treatment on the grounds... 
    Suggested
    Full time
    Remote work

    Tyk Technologies Limited

    Remote
    1 day ago
  • $110k - $120k per year

     ...Time, Permanent Reports to: Site Reliability Lead, Platform Engineering We're looking for Site...  ...metrics from operating systems and applications to support performance tuning and fault...  ...practices (logging standards, use of reliable libraries, SLA/SLO goals) Participate... 
    Suggested
    Permanent employment
    Full time

    Paybyphone

    Remote
    1 day ago
  • $100k - $125k per year

     ...experienced and motivated Software Engineer to join our dynamic Site Reliability Engineering (SRE) team. As a Site...  ...end-to-end lifecycle ownership of applications  Knowledge of architecture and application...  ...Tipalti  Our tech teams are the engine behind our business. Tipalti’s... 
    Suggested
    Full time
    Work at office
    Flexible hours

    Tipalti

    Remote
    1 day ago
  •  ...modern, IoT-enabled, cloud-based tool for reliability, safety, and operations of physical...  ...We’re looking for a Site Reliability Engineer to help advance MaintainX’s reliability...  ...creating a diverse environment. All qualified applicants will receive consideration for... 
    Full time
    Immediate start

    Maintainx

    Remote
    1 day ago
  •  ...Description Prioritize candidates with medical device or regulated hardware experience, strong background in reliability engineering, HALT/HASS/ALT, and statistical modeling (Weibull, lognormal). Technical Evaluation – Assess expertise in DFMEA/PFMEA, fault tree analysis... 
    Full time

    Sapsol Technologies Inc

    Remote
    1 day ago
  • $120k - $200k per year

     ...and many more.   ABOUT THE ROLE At LayerZero, our Site Reliability Engineering (SRE) team is at the intersection of software and systems...  ...internal systems to those external users interact with—are reliable, meet the uptime expectations of our users, and continuously... 
    Full time

    Layer Zero Labs Llc

    Remote
    1 day ago
  • $80k - $95k per year

     ...Running. Improve What's Next. We're partnering confidentially with a leading Canadian food manufacturer seeking a Mechanical Reliability Engineer to play a key role in improving equipment reliability, maintenance strategy, and plant performance within a highly automated manufacturing... 
    Long term contract
    Permanent employment
    Full time
    For contractors
    Relocation package

    Remote People

    Remote
    1 day ago
  •  ...visit .     The Role     CMG is looking for a Site Reliability Engineer (SRE) with a strong focus on monitoring, observability, and...  ...reliability, performance, and scalability of our infrastructure and applications. You will be responsible for designing, implementing, and... 
    Remote job
    Full time
    Local area

    Capital Markets Gateway

    Remote
    1 day ago
  • $20 per day

     ...careers page to see how you can grow with us! As a Site Reliability Engineer at Hiive, you will be responsible for ensuring the...  ...performance and system behavior, and ensuring these services are reliable, scalable, and cost-efficient in production. In this role... 
    Full time
    Summer holiday
    Relocation

    Hiive

    Remote
    1 day ago
  •  ...you’re relentless about the high quality, reliability, and security our customers demand. You...  ...deliverables will reach the entire engineering organization to enable product teams to...  ...craft. What You Bring ~6+ years of applicable software engineering or SRE experience... 
    Long term contract
    Full time
    Remote work
    Flexible hours

    Axon

    Remote
    1 day ago
  •  ...automation, to help customers choose the best solutions that meet their applications requirements. You’ll assist across all touch-points, including...  ...to develop and test new features and software. Applications Engineer Responsibilities Help customers choose the best solutions... 
    Permanent employment
    Full time
    Casual work
    Flexible hours

    Zaber Technologies

    Remote
    1 day ago
  • $145.9k per year

     ...Are you a seasoned production engineering professional who wants to...  ...looking for a Senior Site Reliability Engineer to be part of our...  ...software and help to ensure our application operates at a high level of...  ...are committed to working with applicants requesting accommodation at... 
    Long term contract
    Full time
    Internship
    Work at office
    Local area
    Worldwide

    Jobber

    Remote
    1 day ago
  •  ...Our mission is to deliver reliable, secure, and scalable data infrastructure...  ...lakehouse capabilities that engineering teams can depend on to build...  ...such, we strongly encourage applications from Indigenous peoples,...  ...of personal data from job applicants, please read our Candidate... 
    Full time
    Worldwide

    Appdirect

    Remote
    1 day ago
  • $197.5k - $225k per year

     ...About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving...  ..., multi-tenant, high-availability applications. Build and operate AI tooling infrastructure...  ...protected category in accordance with applicable law.  We also consider qualified... 
    Full time

    Securityscorecard

    Remote
    1 day ago
  • $145k - $185k per year

     ...simulation at scale possible. We're hiring a Senior Site Reliability Engineer to help build and operate that infrastructure. This role sits...  ...to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist... 
    Remote job
    Full time

    Parallel Domain

    Remote
    1 day ago
  • $123k - $160k per year

     ...here. About The Group We're hiring an AI DevOps & Reliability Engineer to own how software ships and runs at Branch. The role has...  ...qualify for relocation or visa sponsorship.  In accordance with applicable law, the following represents a reasonable estimated... 
    Remote job
    Full time
    Internship
    Relocation

    Branch Metrics

    Remote
    1 day ago
  •  ...installed in commercial HVAC, institutional, municipal, and industrial applications including parking facilities, refrigeration plants, commercial...  ...YOU: We are looking for a smart, articulate Application Engineer capable of designing Gas Detection systems for condo towers,... 
    Full time

    Critical Environment Technologies Canada Inc.

    Remote
    1 day ago
  • $101.2k - $136.9k per year

     ...member facing services. We are looking for a highly skilled Site Reliability Engineer (SRE) who will help build, operate and continuously improve...  ...at any stage of the recruitment process (including the application stage), we encourage you to let us know by contacting our Talent... 
    Permanent employment
    Full time
    Internship
    Work at office
    Immediate start
    Home office
    Flexible hours
    2 days per week

    Vancity

    Remote
    1 day ago
  • $260k - $275k per year

     ...enterprises • Solve complex reliability challenges at scale • Influence architecture and engineering culture at a company level...  ...platforms that our product and application teams will depend on   You...  ...focus on creating reusable, reliable, and scalable solutions that... 
    Full time

    Saviynt

    Remote
    1 day ago
  •  ...committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building...  ...employment opportunities to all employees and applicants and prohibits discrimination and harassment... 
    Full time
    Local area
    Remote work
    Home office
    Flexible hours

    Clickhouse

    Remote
    1 day ago
  •  ...few.  About the role We’re looking for a Senior Site Reliability Engineer (SRE) to help strengthen and scale our multi-cloud platform...  ...for those who are eligible to work in Canada. We thank all applicants for taking the time to apply, but only candidates who make it... 
    Full time
    Internship
    Remote work
    Work from home

    Scalepad

    Remote
    1 day ago
  • $130k - $150k per year

     ...Job Ti tle: Fixed Plant Reliabilit y Engineer   Reporting directly to the Mill Manager, the Fixed Plant Reliability Engineer is responsible for leading the development...  ...the Eskay Creek Site facilities. For those applicants located elsewhere in western Canada, flight... 
    Permanent employment
    Full time
    Local area
    Remote work

    Skeena Gold Silver

    Remote
    2 days ago
  • $110k - $160k per year

     ...Overview  We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within...  ...best practices. You'll work closely with Application, Platform, and Security teams to drive...  ...Opportunity Employer and considers applicants for employment without regard to race,... 
    Full time
    Internship
    Work at office
    Local area
    Flexible hours

    Magnet Forensics

    Remote
    1 day ago
  • $140k - $180k per year

     ...marketers, criminals, and surveillance dragnets. Our well-received applications have appeared on Lifehacker, Techradar, and CNET....  ...online privacy for all. Right now we are looking for a Site Reliability Engineer to help us tame DNS. About the Position Linux system... 
    Full time
    Direct hire
    Work at office
    Remote work

    Funded.club

    Remote
    1 day ago
  •  ...subsidiaries.   At this moment, we need an Applications Engineer to cover the province of Alberta ....  ...lathes. G-Drivers License, with reliable transportation Minimum: CNC...  ...employment opportunities to all employees and applicants for employment and prohibits... 
    Hourly pay
    Long term contract
    Full time
    Local area
    Night shift

    Mazak Corp

    Remote
    1 day ago
  • $153k - $187k per year

     ...businesses. We’re looking for an incredible Senior Site Reliability Engineer to join our SRE team. We aim to make reliability, security,...  ...speed reinforce one another so that the platform becomes the engine of Relay’s growth. Your love of making high-impact decisions... 
    Full time
    Internship

    Relay

    Remote
    1 day ago
  •  ...and many more.   ABOUT THE ROLE At LayerZero, our Site Reliability Engineering (SRE) team is at the intersection of software and systems...  ...internal systems to those external users interact with — are reliable, meet the uptime expectations of our users, and continuously... 
    Full time

    Layer Zero Labs Llc

    Remote
    1 day ago
  • $136.9k - $152.1k per year

     ...requirement for this role. Job summary: The Database Reliability Engineer (DBRE) is responsible for managing, building, maintaining, monitoring...  ...database infrastructure that our mission critical SaaS application runs on. The DBRE role is also responsible for improving... 
    Remote job
    Long term contract
    Full time
    Work at office
    Flexible hours

    Pointclickcare

    Remote
    1 day ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Application Reliability Engineer. Be the first to apply!