Jobs / Ora***
Principal Site Reliability Engineer - Exadata Cloud Service
Ora*** · United States
Visa sponsorship details are locked. Unlock company name and apply link with .
United StatesRemote
Remuneration
Not specified
Location
United States
Visa sponsorship
Sponsors visa
Job summary
Job Description Are you interested in solving the complex challenges involved in building and operating large-scale, distributed cloud infrastructure? Ora*** Cloud Infrastructure is developing the next generation of cloud technologies operating in highly available, scalable, secure, distributed, and multi-tenant environments.
Qualifications
- Mentor engineers and provide technical leadership during complex projects, incidents, and architectural discussions.
- Establish and promote standard engineering practices, operational procedures, and reliability principles across the organization.
- Participate in an on-call rotation and provide escalation support for critical production incidents.
- Required
- Bachelor’s degree in computer science, Computer Engineering, Information Systems, Management Information Systems, or another relevant technical field, or equivalent practical experience.
- Typically, 8 or more years of experience in software engineering, site reliability engineering, systems engineering, database engineering, cloud operations, system administration, or a related technical discipline.
- Advanced programming and scripting
- Preferred
- Experience with Ora*** Database
Responsibilities
- Design, develop, test, and deliver software and automation that improve the availability, scalability, latency, security, operability, and efficiency of Ora*** Database as a Service offering.
- Lead the investigation and resolution of complex technical issues spanning Exadata Cloud Service, Autonomous Database, Ora*** Database, operating systems, virtualization, storage, networking, and cloud infrastructure.
- Coordinate response to high-severity incidents, including technical diagnosis, mitigation, stakeholder communication, recovery, and post-incident review.
- Perform detailed root-cause analysis and develop corrective and preventive solutions that reduce the likelihood and impact of recurrence.
- Build automation to eliminate repetitive operational work, reduce human error, accelerate incident response, and improve fleet-management efficiency.
- Apply AI-assisted engineering and operations techniques to improve anomaly detection, incident correlation, troubleshooting, knowledge retrieval, capacity forecasting, and operational decision-making.
- Evaluate and integrate generative AI, machine learning, and large language model capabilities into appropriate SRE workflows while maintaining security, privacy, accuracy, and human oversight.
- Develop
Skills
CommunicationLeadership
Degrees
AssociateDegree
Work schedule
On-callRotationShift
Industry
Automotive
Company size
EnterpriseSmb