Lead Site Reliability Engineer

Company:  Mastercard
Location: O Fallon
Closing Date: 21/10/2026
Salary: $155,000 - $205,000 Per Annum
Hours: Full Time
Type: Permanent

Job Description

Mastercard is seeking a Lead Site Reliability Engineer to drive reliability, scalability, and security for mission-critical financial services platforms. You will design and optimize cloud-native, highly available systems, implement SRE best practices, and lead incident response and postmortems. Partnering with IT and Cybersecurity teams, you'll automate deployments, observability, and resilience testing while mentoring engineers. Ideal candidates bring deep experience with cloud, CI/CD, infrastructure-as-code, and securing large-scale, distributed systems in a regulated environment.

Responsibilities

  • Lead design and operation of highly available, secure, and scalable financial services platforms.
  • Define and implement SRE best practices, including SLOs, SLIs, and error budgets.
  • Architect and maintain cloud-native infrastructure using infrastructure-as-code and automation.
  • Own incident response, root cause analysis, and postmortems for critical production issues.
  • Drive observability across systems with robust monitoring, logging, and alerting solutions.
  • Collaborate closely with IT and Cybersecurity teams to embed security and compliance into the stack.
  • Optimize performance, capacity planning, and cost management for large-scale distributed systems.
  • Mentor and guide engineers on SRE principles, tooling, and operational excellence.
  • Continuously improve CI/CD pipelines and deployment strategies for safer, faster releases.
  • Champion a culture of reliability, innovation, and continuous improvement within the team.

Required Skills

  • Site Reliability Engineering (SRE)
  • Public cloud platforms (AWS, GCP, or Azure)
  • Kubernetes and container orchestration
  • Linux systems engineering and administration
  • Infrastructure as Code (Terraform, Cloud
  • Formation, or similar)
  • CI/CD pipelines (Jenkins, Git
  • Lab CI, Git
  • Hub Actions, or similar)
  • Monitoring and observability (Prometheus, Grafana, Datadog, Splunk, etc.)
  • Scripting/programming (Python, Go, or similar)
  • Security and compliance for financial/regulated environments
  • Incident management and on-call operations
Apply Now
Share this job
Mastercard
  • Similar Jobs

  • Senior Site Reliability Engineer

    O Fallon
    View Job
  • Manager, Site Reliability Engineering

    O Fallon
    View Job
  • Site Reliability Engineering Intern

    O Fallon
    View Job
  • Network Engineer

    Saint Peters
    View Job
  • BizOps Engineer II

    O Fallon
    View Job
An unhandled error has occurred. Reload 🗙