Sign up to access all features of our service.
  • Job search
  • Favorites
  • Create a CV
    New
  • Salaries
  • Subscriptions

Site Reliability Engineer

Full-time

Capital Markets Gateway

Canada
  • Remote job


The Company

 

Capital Markets Gateway LLC (CMG) is a capital markets-focused fintech transforming global equity capital markets (ECM) through data, technology, and connectivity. As the preferred source for ECM analytics and the first network connecting the buy-side and sell-side for ECM workflows, we are committed to reshaping how capital markets operate. Founded in 2017 by a team of ECM practitioners, CMG has completed three successful fundraising rounds and is backed by a group of the world’s most prestigious financial institutions. The CMG platform is currently relied upon by nearly 150 buy-side firms representing $40 trillion in AUM and 22 global investment banks. For more information, please visit .  

 

The Role  

 

CMG is looking for a Site Reliability Engineer (SRE) with a strong focus on monitoring, observability, and alerting to ensure the reliability, performance, and scalability of our infrastructure and applications. You will be responsible for designing, implementing, and maintaining monitoring solutions to provide visibility into system health and performance, proactively detect anomalies, and reduce incident response time.  

 

Our Engineering Team  

 

The CMG engineering team consists of domain experts who work collaboratively within a culture of cross-domain knowledge sharing. We value engineers who are passionate about modern technologies and best practices. 

Our engineers are encouraged to challenge the status quo and are constantly seeking improvement and efficiency in our code-base and platform. CMG engineers are empowered to explore solutions using bleeding edge technologies such as AI and bring recommendations to the table. We are in a period of making impactful engineering decisions. 

As part of our process, we believe in taking the time for research and prototyping - this is critical in making the right decisions. Given the experience of our team, we have naturally adopted best practices from local development, through code review and into production rollouts. Besides the standard pull requests, test automation, code coverage tracking, containerization, and one-click deployments we are constantly reviewing these foundational components to develop new best practices. 

 

Responsibilities  

Monitoring & Observability



  • Design, implement, and maintain monitoring and observability solutions using tools like Prometheus, Grafana Stack (Loki/Grafana/Tempo/Alert Manager), Datadog, and OpenTelemetry.  

  • Define and implement SLOs, SLIs, and error budgets to measure system reliability.  

  • Develop and optimize dashboards, alerts, and reports for system performance and business metrics. 

Alerting & Incident Management



  • Design actionable alerting strategies to minimize noise and improve MTTR.

  • Integrate alerting systems with Jira. 

  • Establish and refine runbooks for on-call teams to handle alerts efficiently. 

  • Empower teams to ensure observability coverage and incident response practices.  

Performance Optimization



  • Analyze system performance metrics, identify bottlenecks, and implement optimizations to improve system efficiency, scalability, and cost-effectiveness.

  • Help conduct load testing and capacity planning to ensure systems can handle peak traffic loads.  

Automation and Tooling



  • Identify opportunities for automation and develop tools to streamline operational processes, such as fail-over, configuration management, and monitoring.

  • Implement monitoring and alerting systems within automations to detect and resolve issues proactively.  

Collaboration and Communication



  • Collaborate closely with cross-functional teams, including software engineers, operations, and infrastructure teams, to understand system requirements, provide technical guidance, and drive solutions.

  • Communicate effectively to stakeholders about system changes, incidents, and improvements.  

  • Foment and spread SRE principles and practices across company.

Qualifications


  • Must be based in Latin America

  • English level - C1 or C2

  • Proven experience as a Site Reliability Engineer or similar role.  

  • Proficiency in logging, metrics, and tracing frameworks (DataDog, Loki, Prometheus, OpenTelemetry).  

  • Experience with cloud platforms (Azure preferred) and infrastructure-as-code tools (e.g., Terraform).  

  • Strong programming and scripting skills (Python, Bash).  

  • Proficiency in containerization technologies and orchestration tools (Docker, Kubernetes).

  • Understandingof Linux-based systems, networking, and security principles related to containerized applications.  

  • Strong problem-solving and troubleshooting skills, with a passion for identifying and resolving complex technical issues.  

  • Excellent communication and collaboration abilities.  

  • Ability to thrive in a fast-paced, constantly evolving environment.  

  • Experience with PostgreSQL monitoring and optimization (Optional/Nice to have).

If you're passionate about building resilient financial systems, optimizing observability at scale, and solving real-world reliability challenges in capital markets, we’d love to have you on our team!   

Our Tech Stack



  • Azure as an infrastructure provider. We are reviewing secondary cloud options.

  • Docker + Kubernetes for microservice orchestration using Istio service mesh. 

  • PostgreSQL for relational db, ElasticSearch for indexing, Redis for caching.

  • DataDog, Grafana and OpenTelemetry for observability. 

  • GitHub for our Version Control and CI (with our own runners). 

  • CD: Harness and FluxCD.

  • Terraform and Terragrunt as IaaC. 

  • Python and bash for scripting infrastructure. 

  • React - We’re all in on React – we maintain multiple single-page React apps.

  • TypeScript – 99% of our codebase is TypeScript.

  • Latest .NET version for our backend services.

  • GraphQL - Our standard for API communication is GraphQL served by our DotNet Back-End.

Our Values



  • We innovate with purpose  

  • We focus on outcomes vs. output  

  • We believe diverse and inclusive teams fuel innovation  

  • We are humble yet candid  

  • We do right by the customer 

What We Offer






  • Equity  





  • Unlimited PTO (15 days + bank holidays + unlimited additional paid leave)



  • Comprehensive benefits program managed by Globalization Partners  





  • Premium life and income protection  





  • Top private medical and dental insurance  





  • Employee Assistance Program (EAP) 





  • Pension contributions  





  • Remote work environment  






  • Education reimbursement  





  • Continuous learning opportunities  





  • Employee referral bonus  





  • Parental leave  


CMG embraces our ongoing commitment to building a culture reflecting the people, perspectives, and passions it represents. We will accept nothing less than equity, inclusion, and belonging for all. With the only constant in life being change, we will always listen, learn, and improve for the betterment of our teams, customers, and communities. CMG is proud to be an Equal Opportunity Employer. 

Vacancy posted 12 hours ago
Similar jobs that could be interesting for youBased on the Site Reliability Engineer in Canada vacancy
  •  ...like an environment that you believe could work for you then read on to find out more. The role: We’re looking for a Site Reliability Engineer to manage, maintain, improve and provide support on our platform. You will be curious by nature, always looking for ways to... 
    Suggested
    Full time
    Remote work

    Tyk Technologies Limited

    Canada
    12 hours ago
  •  ...environments. We are a modern, IoT-enabled, cloud-based tool for reliability, safety, and operations of physical equipment and...  ...valuing the company at $2.5 billion. We’re looking for a Site Reliability Engineer to help advance MaintainX’s reliability, observability, and... 
    Suggested
    Full time
    Immediate start

    Maintainx

    Canada
    12 hours ago
  • $145k - $185k per year

     ...environments, and the infrastructure underneath that platform is what makes simulation at scale possible. We're hiring a Senior Site Reliability Engineer to help build and operate that infrastructure. This role sits at the core of how we run large-scale, distributed simulation... 
    Suggested
    Remote job
    Full time

    Parallel Domain

    Canada
    12 hours ago
  •  ...a part of our journey! About the role We are committed to providing our customers with reliable and secure services so we are expanding our central Site Reliability Engineering team. You will be responsible for building and leading processes to ensure the reliability,... 
    Suggested
    Full time
    Local area
    Remote work
    Home office
    Flexible hours

    Clickhouse

    Canada
    12 hours ago
  •  ...you’re relentless about the high quality, reliability, and security our customers demand. You...  ...technical deliverables will reach the entire engineering organization to enable product teams to...  ...confidence. Exemplify cloud-native site reliability best practices with a strong... 
    Suggested
    Long term contract
    Full time
    Remote work
    Flexible hours

    Axon

    Canada
    12 hours ago
  •  ...made, and we are not tiptoeing into it. We are rebuilding our engineering culture around a simple belief: AI changes everything. How teams...  ...team builds on — and this role is at the center of keeping it reliable, fast, and scalable. As a Staff SRE, you'll own the infrastructure... 
    Remote job
    Full time
    Internship
    Work at office
    Local area
    Flexible hours
    Shift work
    Weekend work

    Babylist, Inc

    Canada
    12 hours ago
  • $197.5k - $225k per year

     ...investors including Silver Lake Waterman, Moody’s, Sequoia Capital, GV and Riverwood Capital. About the Team: As a Senior Site Reliability Engineer, you will be a key technical leader driving the design and optimization of our Kubernetes-based infrastructure and CI/CD... 
    Full time

    Securityscorecard

    Canada
    12 hours ago
  • $140k - $180k per year

     ...the world and our goal is preserving uncensored Internet access and online privacy for all. Right now we are looking for a Site Reliability Engineer to help us tame DNS. About the Position Linux system administration and troubleshooting Network configuration and... 
    Full time
    Direct hire
    Work at office
    Remote work

    Funded.club

    Canada
    12 hours ago
  • $110k - $160k per year

     ...to join our team working towards this goal, we would love to hear from you!  Role Overview  We're seeking a Senior Site Reliability Engineer to join our SaaS-Ops team within Shared Services Engineering. The team owns reliability and operational excellence for our highly... 
    Full time
    Internship
    Work at office
    Local area
    Flexible hours

    Magnet Forensics

    Canada
    12 hours ago
  •  ...overview of this role You'll join the Dedicated team as a Site Reliability Engineer focused on Environment Automation , where your work will...  .... In this role, you'll help keep these environments reliable, scalable, secure, and consistent by treating everything as... 
    Full time
    Remote work
    Home office

    Gitlab

    Canada
    12 hours ago
  • $107k - $161k per year

     ...to meet with your team for events or meetings. Join Our Team GoDaddy is seeking a highly skilled and motivated Senior Site Reliability Engineer to join our Database Infrastructure team. This role focuses on designing, developing, and deploying automated solutions for... 
    Full time
    Second job
    Work at office
    Local area
    Remote work
    Work from home

    GoDaddy

    Canada
    12 hours ago
  • $163k - $194k per year

     ...daily. We run the platforms that every engineering team at Life360 depends on, including AWS...  ...Life360 is hiring an AI-Native Site Reliability Engineer — a senior engineer who doesn’...  ...infrastructure in the future. As a Senior Site Reliable Engineer II - Infrastructure (AI Native... 
    Full time
    Summer work
    Remote work
    Flexible hours

    Life360

    Canada
    12 hours ago
  •  ...Description Prioritize candidates with medical device or regulated hardware experience, strong background in reliability engineering, HALT/HASS/ALT, and statistical modeling (Weibull, lognormal). Technical Evaluation – Assess expertise in DFMEA/PFMEA, fault tree analysis... 
    Full time

    Sapsol Technologies Inc

    Canada
    12 hours ago
  •  ...group includes our scientific research & engineering division (Skynet Software) and Canadian...  ...DESCRIPTION: Chelsea Avondale is looking for a Reliability Engineer with a background in...  ...maintaining high-performance, scalable, and reliable web systems. ~ We also encourage... 
    Full time

    Chelsea Avondale

    Canada
    12 hours ago
  • $80k - $150k per year

     ...Innodata (Nasdaq: INOD) is a global data engineering company. We believe that data and...  ...applications built and running on Google App Engine (GAE) and microservices. The role is...  ...applications with demanding availability, reliability, and performance requirements. Ability... 
    Full time
    Contract work
    Fixed term contract
    Internship
    Flexible hours

    Innodata Inc.

    Canada
    12 hours ago
  • $80k - $95k per year

     ...Running. Improve What's Next. We're partnering confidentially with a leading Canadian food manufacturer seeking a Mechanical Reliability Engineer to play a key role in improving equipment reliability, maintenance strategy, and plant performance within a highly automated manufacturing... 
    Long term contract
    Permanent employment
    Full time
    For contractors
    Relocation package

    Remote People

    Canada
    12 hours ago
  • $125k - $145k per year

     ...Customer Success Work the Problem Be Curious Own the Outcome Respect Each Other   YOUR IMPACT As a Senior Site Reliability Engineer (SRE) / DevOps Engineer you will be responsible for ensuring the stability, performance, scalability, and reliability of... 
    Local area

    Redwood Software

    Canada
    4 days ago
  •  ...Responsibilities: Coordinate mobilization and demobilization of site personnel, including flights and accommodation. Maintain...  ...solving skills and the ability to adapt to changing priorities. Reliable and punctual with a strong work ethic. Excellent communication... 
    Full time
    For contractors
    Work at office
    Remote work

    Thyssen Mining

    Canada
    12 hours ago
  • $23 - $27 per hour

     ...cooling to each data center in the Server Farm portfolio. This is complimented by providing remote monitoring services to both internal sites and external clients. The critical facilities site Administrative Assistant role performs functions critical to the day-to-day... 
    Full time
    Work at office
    Remote work

    Serverfarm

    Canada
    12 hours ago
  • $25 - $42.5 per hour

     ...Site Administrator Company Description WB Melback Corporation is a wholly Canadian owned company with headquarters based in Haileybury, Ontario with job sites all over the country. WB Melback Corporation provides Mechanical Maintenance and Industrial Construction Services... 
    Hourly pay
    Daily paid
    Full time
    Contract work
    Work at office
    Work from home
    Monday to friday
    Shift work

    Wb Melback

    Canada
    12 hours ago
  •  ...become part of a global team of over 50,000 planners, designers, engineers, scientists, digital innovators, program and construction managers...  .... Job Description AECOM is seeking a Health and Safety Site Representative to support a major hydroelectric infrastructure project... 
    Full time
    For contractors
    Work at office
    Local area
    Remote work
    Worldwide
    Relocation
    Flexible hours

    Aecom

    Canada
    12 hours ago
  • $95k - $135k per year

     ...The Senior Contaminated Sites Specialist is a senior technical and leadership role within the Brownfield Assessment & Development service...  .... Bachelor’s degree in hydrogeology, environmental science, engineering, or a related discipline. Eligible for registration with a... 
    Long term contract
    Full time
    Temporary work
    Part time
    Casual work
    Work at office
    Flexible hours

    Millennium EMS Solutions

    Canada
    12 hours ago
  • $40 - $55 per hour

     ...years, we've earned the trust of Calgary’s communities by delivering reliable, high-quality plumbing and HVAC services. We’re dedicated to...  ...The successful candidate will be responsible for supervising job sites, coordinating with project managers and estimators, and ensuring... 
    Hourly pay
    Full time
    For contractors
    Apprenticeship
    Local area
    Monday to friday
    Flexible hours

    Wiehler Mechanical Ltd

    Canada
    12 hours ago
  •  ...The Remote Site Technicians will safely and effectively diagnose, repair and or assemble all Caterpillar mining and earthmoving equipment...  ...education in the field of heavy equipment mechanics or heavy engine mechanics Minimum 3 years of experience as a technician in repair... 
    Remote job
    Permanent employment
    Full time
    Flexible hours

    Toromont Cat

    Canada
    12 hours ago
  •  ...located in Halifax and requires local residence. We are seeking a Site Technical Manager to oversee the daily operation and maintenance...  ...Government legislation BESO I or II, SMT, Fourth Class Power Engineer (NS certified), or equivalent. HVACR and refrigeration... 
    Full time
    Contract work
    For contractors
    Local area

    Dexterra Inc

    Canada
    12 hours ago
  • $22.54 - $27.05 per hour

     ...and contributes at meetings requested by management.    Other  ~ Perform other duties as required.  Accountabilities:  Reliable time and attendance at all scheduled shifts. Smooth operation of well-maintained station and stores while on shift. Maintain a... 
    Hourly pay
    Long term contract
    Full time
    Casual work
    Immediate start
    Monday to friday
    Flexible hours
    Shift work
    Day shift
    Afternoon shift

    Nchḵay̓

    Canada
    12 hours ago
  • $68.64k - $73.64k per year

     ...Field and Operations Team, the Safe & Secure Support Team plays an integral function in delivering our mission to every associate at every site. The Area Lead will be a part of the Field Safe & Secure Team providing execution of core Safe & Secure routines to deliver upon the... 
    Full time
    Immediate start
    Afternoon shift

    Carvana

    Canada
    12 hours ago
  •  ...flow expansion, we sit at the intersection of product/platform engineers and financial partners, connecting them to ensure that everyone...  ...ensure that data flowing through our partner and internal systems is reliable and correct Collaborate across the company, including... 

    Stripe

    Canada
    4 days ago
  • $242.54k - $302.84k per year

     ...skilled and diligent full-time Frontend Engineer to join our growing team. You will work as...  ...deployments, release ergonomics, and operational reliability Raise the quality bar for frontend...  ...company retreat, participate in team off-sites, and collaborate in person with teammates... 
    Long term contract
    Full time
    Internship
    Work at office
    Local area
    Remote work
    Home office
    Flexible hours

    Tailscale

    Canada
    12 hours ago
  •  ...Description Job Title: Principal Engineer Department: Engineering Work Location...  ...ensuring goals for performance, cost, and reliability are met.   Key Responsibilities: Technical...  ...when required for type testing on site or at external labs (heart run, SC tests,... 
    Full time

    Cam Tran

    Canada
    12 hours ago

Do you want to receive more vacancies?

Subscribe and receive similar vacancies to Site Reliability Engineer. Be the first to apply!