Jobs / M&T***
Principal Site Reliability Engineer
M&T*** · Buffalo, NY, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Buffalo, NY, United States139,700-232,900 USD/yearlyOnsite
Remuneration
139,700-232,900 USD/yearly
Location
Buffalo, NY, United States
Visa sponsorship
Sponsors visa
Job summary
Overview: Responsible for designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise. Operates as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering practices, operational excellence, and automation across the Software Development Lifecycle.
Qualifications
- Serves as a mentor and technical leader for less experienced engineers across Technology.
- Lead incident management practices, including detection, response, escalation, and recovery processes.
- Drive problem management and root cause analysis to prevent systemic issues.
- Develop and promote observability strategies, including logging, monitoring, alerting, and tracing.
- Lead automation initiatives for self-healing systems and operational workflows.
- Contribute to and review technical roadmaps with reliability and performance considerations.
- Partner with development, infrastructure, cybersecurity, and architecture teams.
- Serve as a technical authority for performance, resilience, and capacity planning.
- Drive production readiness practices including performance testing and failover capabilities.
- Lead cross-team reliability improvement initiatives.
- Participate in and lead post-incident reviews ensuring actionable outcomes.
- Mentor engineers on reliability engineering and best practices.
Responsibilities
- Responsible for designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise.
- Accountable for defining and driving service reliability standards, including SLOs, SLAs, and error budgets across platforms.
- Design and implement highly available, fault-tolerant architectures aligned with enterprise scalability and resiliency
- Applies expert-level SRE practices across multiple platforms.
- Drives enterprise-wide reliability improvements and influences technical direction without direct authority.
- Supervisory/Managerial
- No supervisory
- Education and Experience Required:
- In lieu of a degree, a combined minimum of 11 years’ education and/or relevant work experience, including a minimum of 7 years’ systems analysis and/ or application development work experience.
- Expert experience in system design, reliability engineering, and production operations.
- Advanced proficiency in at least one programming or scripting language.
- Education and Experience Preferred:
Skills
Communication
Degrees
AssociateBachelorDegree
Industry
AutomotiveBankingEducation
Company size
EnterpriseSmb
Contract length
9 years