Application Reliability Engineer
$80k - $150k per yearInnodata Inc.
Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and Artificial Intelligence (AI) are inextricably linked. Our mission is to enable the responsible advancement of artificial intelligence by providing the data, evaluation frameworks, and human expertise required to build AI systems that can be trusted at scale. We provide a range of transferable solutions, platforms, and services for Generative AI / AI builders and adopters. In every relationship, we honor our 36+ year legacy delivering the highest quality data and outstanding outcomes for our customers.
Scope of the Role:
We are looking for a hands-on Application Support Engineer to support, maintain, and enhance business-critical enterprise applications built and running on Google App Engine (GAE) and microservices.
The role is focused on application availability, production support, incident response, troubleshooting, and continuous feature enhancement rather than building a new application from the ground up. The ideal candidate can quickly understand an existing microservices-based application landscape, restore service when users are impacted, and deliver incremental enhancements safely across test, pre-production, and production environments.
Experience supporting large-scale, business-critical applications in complex enterprise technology environments is preferred.
What You’ll Own:
- Provide production support and maintenance for enterprise applications hosted on Google App Engine, including standard and flexible environments.
- Act as a first point of contact for user-impacting incidents: triage, diagnose, restore service, and drive issues to closure within agreed SLAs/SLOs.
- Troubleshoot application errors, failed requests, latency and performance degradation, service-to-service failures, configuration issues, quota/scaling limits, and dependency or integration failures.
- Design, develop, and deliver feature enhancements and functional improvements to existing applications based on user and business needs.
- Support applications built on microservices architecture, including service boundaries, APIs/contracts, inter-service communication, authentication, and failure/retry behavior.
- Own build, release, and deployment activities across development, test, pre-production, and production environments with appropriate validation, approvals, and rollback plans.
- Manage App Engine deployments including versions, traffic splitting/migration, canary and staged rollouts, rollbacks, service configuration, and scaling settings.
- Perform root-cause analysis for recurring production issues and implement sustainable fixes.
- Build and maintain monitoring, alerting, logging, dashboards, and error reporting using Cloud Monitoring, Cloud Logging, Error Reporting, and Cloud Trace.
- Support platform, framework, library, dependency, and runtime upgrades while maintaining stability, supportability, and compliance.
- Support IAM, service accounts, access controls, secrets management, and operational governance.
- Participate in change management, release-readiness reviews, and on-call/rotational support as required.
- Collaborate with client and cross-functional Application Engineering, Product, QA, Data Engineering, Infrastructure, and Platform teams.
- Adapt to established client-specific engineering, security, review, change-management, and operational processes.
- Create and maintain technical documentation, operational runbooks, troubleshooting guides, deployment procedures, and support playbooks.
Microservices Expectations:
- Service decomposition and ownership: understand service boundaries, upstream/downstream dependencies, and ownership of behavior or data.
- APIs and contracts: REST and gRPC/protobuf interfaces, versioning, backward compatibility, and contract-change impacts.
- Inter-service communication: synchronous calls, asynchronous/event-driven messaging (Pub/Sub, Cloud Tasks), idempotency, retries, timeouts, backoff, and circuit breaking.
- Service-to-service authentication and authorization using service accounts, identity/tokens, and least-privilege access.
- Distributed troubleshooting using logs, distributed tracing, and correlation IDs to isolate failures across services.
- Failure modes at scale, including cascading failures, partial outages, hot spots, quota exhaustion, and graceful degradation.
- Independent deployability and coordination of multi-service releases when required.
- Service-level observability, including dashboards, SLIs/SLOs, alerting, and error budgets.
You’ll Thrive in This Role If You Have:
- 3-7 years of experience in Application Support, Application Engineering, Software Engineering, Cloud Engineering, or a related role.
- Strong hands-on experience supporting, maintaining, and enhancing production applications on Google Cloud Platform (GCP).
- Hands-on experience with Google App Engine, including deployment, configuration, scaling, versioning, and troubleshooting.
- Solid understanding of microservices architecture, REST/gRPC APIs, service-to-service communication and authentication, distributed tracing, and cross-service debugging.
- Practical experience with deployments and promotions across multiple environments, including release validation, rollback, and change control.
- Strong programming skills in one or more of Python, Java, Node.js / JavaScript, or Go.
- Working knowledge of SQL and application data stores such as Cloud SQL, Firestore/Datastore, Cloud Spanner, or BigQuery.
- Good understanding of GCP IAM, service accounts, permissions, monitoring, logging, alerting, and production operations.
- Experience with CI/CD pipelines and automated build/deployment tooling such as Cloud Build, Jenkins, GitHub Actions, or GitLab CI.
- Experience troubleshooting complex production environments and performing root-cause analysis under time pressure.
- Ability to quickly understand existing systems, codebases, services, configurations, and client-specific tools and workflows.
- Strong communication and collaboration skills, including communication during user-impacting incidents.
Preferred Qualifications:
- Experience supporting large-scale internal or enterprise applications with demanding availability, reliability, and performance requirements.
- Ability to quickly learn and operate within enterprise-specific application frameworks, deployment tooling, and support processes.
- Cloud Run, GKE, Cloud Functions, or other GCP application services.
- Apigee / API Gateway, load balancing, and API management.
- Pub/Sub, Cloud Tasks, Cloud Scheduler, and asynchronous/event-driven patterns.
- Infrastructure as Code, particularly Terraform.
- SRE practices including SLIs/SLOs, error budgets, incident management, and blameless postmortems.
- Containerization with Docker and Kubernetes fundamentals.
- Frontend or full-stack experience supporting user-facing web applications.
- Looker / Tableau / BI platforms or reporting integrations.
- Supporting AI/ML, GenAI, or LLM-based applications running on GCP, including Vertex AI.
The expected salary range for this position is $80,000 - $150,000 CAD per year, based on experience, skills, and qualifications.
Please be aware of recruitment scams involving individuals or organizations falsely claiming to represent employers. Innodata will never ask for payment, banking details, or sensitive personal information during the application process. To learn more on how to recognize job scams, please visit the Federal Trade Commission’s guide at
If you believe you’ve been targeted by a recruitment scam, please report it to Innodata at View email address on jobs.jobcopilot.com and consider reporting it to the FTC at ReportFraud.ftc.gov .
- ...includes our scientific research & engineering division (Skynet Software)... ...Avondale is looking for a Reliability Engineer with a background... ...-performance, scalable, and reliable web systems. ~ We also... ...Software welcomes and encourages applications from people with...SuggestedFull time
- ...Description Prioritize candidates with medical device or regulated hardware experience, strong background in reliability engineering, HALT/HASS/ALT, and statistical modeling (Weibull, lognormal). Technical Evaluation – Assess expertise in DFMEA/PFMEA, fault tree analysis...SuggestedFull time
- ...to find out more. The role: We’re looking for a Site Reliability Engineer to manage, maintain, improve and provide support on our platform... ...employer and we are determined to ensure that no applicant or employee receives less favourable treatment on the grounds...SuggestedFull timeRemote work
- ...modern, IoT-enabled, cloud-based tool for reliability, safety, and operations of physical... ...We’re looking for a Site Reliability Engineer to help advance MaintainX’s reliability... ...creating a diverse environment. All qualified applicants will receive consideration for...SuggestedFull timeImmediate start
- ...visit . The Role CMG is looking for a Site Reliability Engineer (SRE) with a strong focus on monitoring, observability, and... ...reliability, performance, and scalability of our infrastructure and applications. You will be responsible for designing, implementing, and...SuggestedRemote jobFull timeLocal area
$80k - $95k per year
...Running. Improve What's Next. We're partnering confidentially with a leading Canadian food manufacturer seeking a Mechanical Reliability Engineer to play a key role in improving equipment reliability, maintenance strategy, and plant performance within a highly automated manufacturing...Long term contractPermanent employmentFull timeFor contractorsRelocation package$145k - $185k per year
...simulation at scale possible. We're hiring a Senior Site Reliability Engineer to help build and operate that infrastructure. This role sits... ...to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses. These tools assist...Remote jobFull time$197.5k - $225k per year
...About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving... ..., multi-tenant, high-availability applications. Build and operate AI tooling infrastructure... ...protected category in accordance with applicable law. We also consider qualified...Full time$140k - $180k per year
...marketers, criminals, and surveillance dragnets. Our well-received applications have appeared on Lifehacker, Techradar, and CNET.... ...online privacy for all. Right now we are looking for a Site Reliability Engineer to help us tame DNS. About the Position Linux system...Full timeDirect hireWork at officeRemote work$110k - $160k per year
...Overview We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within... ...best practices. You'll work closely with Application, Platform, and Security teams to drive... ...Opportunity Employer and considers applicants for employment without regard to race,...Full timeInternshipWork at officeLocal areaFlexible hours- ...you’re relentless about the high quality, reliability, and security our customers demand. You... ...deliverables will reach the entire engineering organization to enable product teams to... ...craft. What You Bring ~6+ years of applicable software engineering or SRE experience...Long term contractFull timeRemote workFlexible hours
- ...committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building... ...employment opportunities to all employees and applicants and prohibits discrimination and harassment...Full timeLocal areaRemote workHome officeFlexible hours
- ...tiptoeing into it. We are rebuilding our engineering culture around a simple belief: AI... ...this role is at the center of keeping it reliable, fast, and scalable. As a Staff SRE, you... ...evaluation purposes in accordance with applicable privacy laws. By participating in an interview...Remote jobFull timeInternshipWork at officeLocal areaFlexible hoursShift workWeekend work
- ...among its board members. Role Overview The Senior Application Security Engineer II is a senior individual contributor responsible for strengthening... ...inclusive of annual base salary and annual bonus as applicable. For sales roles, the range provided is the role’s On...Long term contractFull timeWorldwide
$107k - $161k per year
...highly skilled and motivated Senior Site Reliability Engineer to join our Database Infrastructure... ...of databases supporting both corporate applications and customer hosting services. You'll... ...will consider for employment qualified applicants with criminal histories in a manner...Full timeSecond jobWork at officeLocal areaRemote workWork from home$120k - $140k per year
...position: Staff Software Systems Engineer Position type: Full time... ...to: Senior Manager of Applications Engineering Application... ...filled. Work authorization: Applicants must have legal... ...build system integrity and reliability - Distributed Team Ready:...Full timeLocal areaWork from homeFlexible hours- ...'ll join the Dedicated team as a Site Reliability Engineer focused on Environment Automation ,... ..., you'll help keep these environments reliable, scalable, secure, and consistent by treating... ..., cloud services, and GitLab applications, supporting early issue detection and...Full timeRemote workHome office
$75k - $105k per year
...connection. Reports to: Director of Engineering Location: Toronto, Canada or... ...the architecture, development, and reliability of the web applications powering our ecommerce business —... ...managers (Composer, npm), and templating engines (Blade, Twig). Experience with...Long term contractFull timeInternshipRemote work$163k - $194k per year
...daily. We run the platforms that every engineering team at Life360 depends on, including AWS... ...Life360 is hiring an AI-Native Site Reliability Engineer — a senior engineer who doesn’t... ...infrastructure in the future. As a Senior Site Reliable Engineer II - Infrastructure (AI Native)...Full timeSummer workRemote workFlexible hours$150k - $180k per year
...simplicity. It offers the simplest, most reliable, and most affordable way to move... ...what will your new role look like? As an Application Security Manager, you will be a hands-on... ...architecture and design discussions with engineering teams; Collaborating with Infrastructure...Full timeLocal area$95k - $110k per year
...build a better future. This is a Work-from-Home opportunity for applicants outside the commutable range of our office located at 2 Queen... .... You balance "fix it now" urgency with "fix it properly" engineering instincts — and you know which mode the moment calls for. You...Full timeInternshipWork at officeLocal areaRemote workWork from homeWeekend work2 days per week$70k - $95k per year
...AEM is looking for a motivated Business Applications Administrator to support the enterprise... ...OpenAI APIs, agentic workflows, prompt engineering). Experience in a multi-entity or... ...estimate has not been adjusted for the applicable geographic differential associated with...Remote jobFull timeLocal area$186.37k - $223.64k per year
...and we would be interested in applicants located in Canadian time... ...this time). Staff Backend Engineer - Application Core Services,... ...and Grafana teams to deliver reliable internal and customer-facing... ...includes maintaining the billing engine responsible for customer usage...Remote jobLong term contractFull timeContract workInternshipLocal areaImmediate start$120k - $155k per year
...is hiring a Senior Software Application Developer to join one of our... ...vendors. This role is for an engineer with a product-centric... ...relational data models to build reliable, performant backend services.... ...Assessment platforms and our applicant tracking systems may host data...Full timeTemporary workRemote work- ...a two year growth initiative to build our total workforce to 350. The workforce is highly skilled and consists of application consultants, software engineers, and networking engineers located throughout the U.S and Canada. Sage Intacct Application Consultant (Remote)...Remote jobFull timeWork from home
- Who We Are Welcome to TELUS Digital — where innovation drives impact at a global scale. As an award-winning digital product consultancy and the digital division of TELUS , one of Canada’s largest telecommunications providers, we design and deliver transformative customer...Full time
$207k - $272k per year
...workload migrations and modernization, cloud native application development, DevOps, data engineering, security and compliance, and everything in between.... ...application process, it will only be in accordance with applicable laws and regulations, and your information will be...Long term contractFull timeWork at officeRemote work- ...future of scientific research. We have created the world’s most comprehensive digital science platform – best-of-breed software applications already used by more than 2 million scientists, together in a single ecosystem united by a powerful, flexible enterprise data platform...Long term contractFull timeRemote workFlexible hours
- ...limited to, User Management: Provide technical support to Sales, Business Users, Institutional and Retail customer base with focus on application support, API support and network connectivity. Incident Management: Identification and resolution of production incidents. The...Full timeShift workWeekend work
$115k - $130k per year
...specific business management solution is engineered to address real-world needs—featuring... ...visit: Job Description The Manager, Application Support is a member of the Acumatica... ...Employer/Veterans/Disabled. All qualified applicants will receive consideration for employment...Long term contractFull timeTemporary workWorldwideFlexible hours
Do you want to receive more vacancies?
Subscribe and receive similar vacancies to Application Reliability Engineer. Be the first to apply!

