Jobs / Ora***

Principal Site Reliability Engineer - Exadata Cloud Service

Ora*** · United States
Visa sponsorship details are locked. Unlock company name and apply link with .
United StatesRemote
Remuneration
Not specified
Location
United States
Visa sponsorship
Sponsors visa

Job summary

Job Description Are you interested in solving the complex challenges involved in building and operating large-scale, distributed cloud infrastructure? Ora*** Cloud Infrastructure is developing the next generation of cloud technologies operating in highly available, scalable, secure, distributed, and multi-tenant environments.

Qualifications

  • Mentor engineers and provide technical leadership during complex projects, incidents, and architectural discussions.
  • Establish and promote standard engineering practices, operational procedures, and reliability principles across the organization.
  • Participate in an on-call rotation and provide escalation support for critical production incidents.
  • Required
  • Bachelor’s degree in computer science, Computer Engineering, Information Systems, Management Information Systems, or another relevant technical field, or equivalent practical experience.
  • Typically, 8 or more years of experience in software engineering, site reliability engineering, systems engineering, database engineering, cloud operations, system administration, or a related technical discipline.
  • Advanced programming and scripting
  • Preferred
  • Experience with Ora*** Database

Responsibilities

  • Design, develop, test, and deliver software and automation that improve the availability, scalability, latency, security, operability, and efficiency of Ora*** Database as a Service offering.
  • Lead the investigation and resolution of complex technical issues spanning Exadata Cloud Service, Autonomous Database, Ora*** Database, operating systems, virtualization, storage, networking, and cloud infrastructure.
  • Coordinate response to high-severity incidents, including technical diagnosis, mitigation, stakeholder communication, recovery, and post-incident review.
  • Perform detailed root-cause analysis and develop corrective and preventive solutions that reduce the likelihood and impact of recurrence.
  • Build automation to eliminate repetitive operational work, reduce human error, accelerate incident response, and improve fleet-management efficiency.
  • Apply AI-assisted engineering and operations techniques to improve anomaly detection, incident correlation, troubleshooting, knowledge retrieval, capacity forecasting, and operational decision-making.
  • Evaluate and integrate generative AI, machine learning, and large language model capabilities into appropriate SRE workflows while maintaining security, privacy, accuracy, and human oversight.
  • Develop

Skills

CommunicationLeadership

Degrees

AssociateDegree

Work schedule

On-callRotationShift

Industry

Automotive

Company size

EnterpriseSmb