Jobs / M&T***

Principal Site Reliability Engineer

M&T*** · Buffalo, NY, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Buffalo, NY, United States139,700-232,900 USD/yearlyOnsite
Remuneration
139,700-232,900 USD/yearly
Location
Buffalo, NY, United States
Visa sponsorship
Sponsors visa

Job summary

Overview: Responsible for designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise. Operates as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering practices, operational excellence, and automation across the Software Development Lifecycle.

Qualifications

  • Serves as a mentor and technical leader for less experienced engineers across Technology.
  • Lead incident management practices, including detection, response, escalation, and recovery processes.
  • Drive problem management and root cause analysis to prevent systemic issues.
  • Develop and promote observability strategies, including logging, monitoring, alerting, and tracing.
  • Lead automation initiatives for self-healing systems and operational workflows.
  • Contribute to and review technical roadmaps with reliability and performance considerations.
  • Partner with development, infrastructure, cybersecurity, and architecture teams.
  • Serve as a technical authority for performance, resilience, and capacity planning.
  • Drive production readiness practices including performance testing and failover capabilities.
  • Lead cross-team reliability improvement initiatives.
  • Participate in and lead post-incident reviews ensuring actionable outcomes.
  • Mentor engineers on reliability engineering and best practices.

Responsibilities

  • Responsible for designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise.
  • Accountable for defining and driving service reliability standards, including SLOs, SLAs, and error budgets across platforms.
  • Design and implement highly available, fault-tolerant architectures aligned with enterprise scalability and resiliency
  • Applies expert-level SRE practices across multiple platforms.
  • Drives enterprise-wide reliability improvements and influences technical direction without direct authority.
  • Supervisory/Managerial
  • No supervisory
  • Education and Experience Required:
  • In lieu of a degree, a combined minimum of 11 years’ education and/or relevant work experience, including a minimum of 7 years’ systems analysis and/ or application development work experience.
  • Expert experience in system design, reliability engineering, and production operations.
  • Advanced proficiency in at least one programming or scripting language.
  • Education and Experience Preferred:

Skills

Communication

Degrees

AssociateBachelorDegree

Industry

AutomotiveBankingEducation

Company size

EnterpriseSmb

Contract length

9 years