Jobs / Dev***
Staff Site Reliability Engineer
Dev*** · United Kingdom
Visa sponsorship details are locked. Unlock company name and apply link with .
United KingdomRemote
Remuneration
Not specified
Location
United Kingdom
Visa sponsorship
Sponsors visa
Job summary
WHO WE ARE AI is changing how software gets built. Code production is becoming a commodity. The focus is shifting from writing code to orchestrating, verifying, and governing change – and the toolchain is the new constraint.
Benefits
A ground-floor role in a new SRE team - you'll shape how we do things, not inherReal ownership of production systems used by engineers at companies you've heardDirect interaction with customers when things go wrong (and when they go right).A culture that values automation over heroics.In-person meetings, such as our annual company offsite and team meetings.Work from home in a remote-first environment.Competitive salaries and equity grants.LOCATIONRemote from anywhere in Europe (GMT).While our team works remotely and is spread across the globe, we deeply value daFind more English Speaking Jobs in United Kingdom on Arbeitnow
Qualifications
- WHO YOU ARE
- 7+ years in SRE, DevOps, or an equivalent role operating production services at scale.
- Experience leading reliability initiatives across multiple teams or services.
- Demonstrated ability to influence technical direction without direct authority.
- Experience designing and operating systems with SLOs and error budgets, and exercising strong judgment in balancing reliability, velocity, and cost.
- Strong Kubernetes experience in production environments.
- Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
- Proficiency with observability
- Experience as a founding or early SRE establishing practices in a growing SaaS organization.
- Familiarity with Dev***.
- JVM language experience (Java, Kotlin).
- Experience with customer-facing and executive-level incident communications.
Responsibilities
- When we execute, we take responsibility for our decisions, measure the success of our innovations, and learn from the results.
- As a Lead SRE, you'll be a technical and operational leader for reliability across Dev***.
- You'll help define our SRE vision, set standards for how we operate production services, and mentor other SREs as the team grows.
- You'll work on our internally-built Cloud Application Platform, Kubernetes on AWS, and develop deep expertise in it.
- When incidents happen, you'll troubleshoot issues across the stack, from application to infrastructure.
- You'll collaborate with the Cloud Platform team to improve the tooling you depend on, and with engineering teams to build reliability into how we ship software.
- If you like automating things and hate doing the same task twice, you'll fit in well.
- You'll be part of a distributed, remote-first team that values asynchronous communication and written documentation.
- Operate and maintain all Dev*** instances and supporting services in production.
- Define and evolve SRE standards, practices, and operating models, including on-call, incident response, postmortems, and SLOs.
- Participate in a follow-the-sun on-call rotation, acting as a technical escalation point for complex or high-severity incidents.
- Lead incident response and blameless retrospectives, ensuring learnings result in measurable reliability improvements.
Skills
CommunicationEnglishLeadershipSAP
Degrees
Associate
Languages
English
Work schedule
24/7NightOn-callRotationShift
Industry
AutomotiveBankingEnergySaas