Jobs / Dev***
Senior Site Reliability Engineer
Dev*** · United Kingdom
Visa sponsorship details are locked. Unlock company name and apply link with .
United KingdomRemote
Remuneration
Not specified
Location
United Kingdom
Visa sponsorship
Sponsors visa
Job summary
WHO WE ARE AI is changing how software gets built. Code production is becoming a commodity. The focus is shifting from writing code to orchestrating, verifying, and governing change – and the toolchain is the new constraint.
Benefits
A ground-floor role in a new SRE team—you'll shape how we do things, not inheritReal ownership of production systems used by engineers at companies you've heardDirect interaction with customers when things go wrong (and when they go right).A culture that values automation over heroics.In-person meetings, such as our annual company offsite and team meetings.Work from home in a remote-first environment.Competitive salaries and equity grants.LOCATIONRemote from anywhere in Europe in the GMT timezone.While our team works remotely and is spread across the globe, we deeply value daFind more English Speaking Jobs in United Kingdom on Arbeitnow
Qualifications
- WHO YOU ARE
- 5+ years in SRE, DevOps, or equivalent role operating production services at scale.
- Strong Kubernetes experience in production environments.
- Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2).
- Proficiency with observability
- Experience operating SaaS platforms at scale.
- Familiarity with Dev***.
- JVM language experience (Java, Kotlin).
- Disaster recovery planning and execution experience.
- Customer-facing incident communication
Responsibilities
- When we execute, we take responsibility for our decisions, measure the success of our innovations, and learn from the results.
- You'll work on our internally-built Cloud Application Platform, Kubernetes on AWS, and develop deep expertise in it.
- When incidents happen, you'll troubleshoot issues across the stack, from application to infrastructure.
- You'll collaborate with the Cloud Platform team to improve the tooling you depend on, and with engineering teams to build reliability into how we ship software.
- If you like automating things and hate doing the same task twice, you'll fit in well.
- You'll be part of a distributed, remote-first team that values asynchronous communication and written documentation.
- Operate and maintain all Dev*** instances and supporting services.
- Participate in a follow-the-sun on-call rotation, owning incident response and troubleshooting issues across the stack.
- Drive automation across application deployment, upgrades, monitoring, self-healing, and recovery.
- Build and maintain observability for all managed services (logging, metrics, tracing, and alerting).
- Work with engineering teams to build reliability into features from the start.
- Run incident response and retrospectives, and make sure we learn from them.
Skills
CommunicationEnglishSAP
Degrees
Associate
Languages
English
Work schedule
24/7NightOn-callRotationShift
Industry
AutomotiveBankingEnergySaas