Jobs / Ora***

Senior Site Reliability Engineer

Ora*** · United States
Visa sponsorship details are locked. Unlock company name and apply link with .
United StatesRemote
Remuneration
Not specified
Location
United States
Visa sponsorship
Sponsors visa

Job summary

OCI Incident Response is the first line of defense in maintaining the high availability of Ora***’s cloud. We minimize customer-impacting events by making them shorter, less frequent, and less impactful through large-scale incident management.

Benefits

We minimize customer-impacting events by making them shorter, less frequent, and

Qualifications

  • We are at the forefront of reducing event duration by leveraging our operational experience, knowledge of best practices, and ability to develop
  • for product roadmaps.
  • Articulate the technical characteristics of services and technology areas, and guide development teams to engineer and add premier capabilities to the Ora*** Cloud service portfolio.
  • Act as the ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs).
  • Minimum
  • Bachelor’s degree or higher in Computer Science or relevant work experience..
  • 3+ years’ experience in Site Reliability Engineering, DevOps, or System Engineering.
  • Must have public cloud operations experience (e.g., AWS, Azure, GCP, OCI).
  • Extensive experience with Major Incident Management in a cloud-based environment.
  • Demonstrate clear understanding of automation and orchestration principles.
  • Experience having worked in at least one modern object-oriented programming language.
  • Experience with professional software engineering standard methodologies such as Agile project management, coding standards, code reviews, source control management, build processes, testing, and operations.

Responsibilities

  • Solve complex problems related to infrastructure cloud services and automate common tasks to ensure continuous availability with minimal human intervention.
  • Command and coordinate SMEs and service leaders to restore services as quickly as possible during major incidents, while keeping accurate and timely data on the progress of such incidents.
  • Utilize a deep understanding of cloud computing design patterns and their dependencies to mitigate complex major incidents.
  • Embed a methodical approach to troubleshoot large, complex, interconnected systems used in incident detection and orchestration.
  • Document pertinent information related to incidents that aids process improvement, identifies deviations, and enables the creation of an incident knowledge base.
  • Monitor and evaluate high-level service and infrastructure dashboards, taking action to address identified anomalies.
  • Identify opportunities and take ownership of automation and/or continuous improvement of incident management process steps and best practices.
  • Define and document the technical architecture of large-scale distributed systems.
  • Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services.
  • Be responsible for the design and delivery of the mission-critical stack, with a focus on security, resiliency, scalability, and performance.
  • Partner with development teams to define operational

Skills

Culinary ArtsProject Management

Degrees

AssociateDegree

Work schedule

RotationShiftWeekend

Industry

AutomotiveDefense

Company size

Smb