Jobs / App***
Site Reliability Engineer, Apple Data Platform / Big Data Platform
App*** · Austin, TX, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Austin, TX, United StatesOnsite
Remuneration
Not specified
Location
Austin, TX, United States
Visa sponsorship
Sponsors visa
Job summary
The App*** Services Engineering team (ASE) is one of the most exciting examples of App***'s long-held passion for combining art and technology. These are the people who power the App Store, App*** TV, App*** Music, App*** Podcasts, and App*** Books - at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.
Qualifications
- Experience with REST Catalog services (e.g., Glue Catalog) and data governance frameworks.
- Prior experience in a customer-facing or technical support role, with a demonstrated passion for customer success.
- Familiarity with observability tooling: Prometheus, Grafana, Splunk, PagerDuty.
- Working knowledge of CI/CD pipelines and deployment workflows.
- Experience with S3 and cloud storage/networking fundamentals.
- Familiarity with data pipeline orchestration and workflow scheduling patterns.
- A track record of automating manual operations through scripting or tooling.
- Intellectual curiosity and a drive to keep learning - for yourself, your team, and the org.
- Minimum
- Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
- 1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.
- Proficient in Python; working knowledge of Golang a plus.
Responsibilities
- Operate, monitor, and triage production and non-production environments across the ADP portfolio - data processing, ML/AI, and multi-cloud infrastructure.
- Participate in a rotating on-call schedule across supported services, including occasional weekday and weekend coverage.
- Own the operational health of big data platform services as SME - driving reliability, support, and customer guidance for Spark, Flink, Airflow, Trino, Notebooks, REST Catalog, and governance tooling.
- Serve as a primary point of contact for internal customers via Slack - clearly communicating status, root cause, and next steps during active issues.
- Screen, triage, and resolve customer-reported service issues and support tickets, prioritizing based on customer impact and urgency.
- Partner with dev teams across time zones to onboard new services - understanding architecture, then designing monitoring, alerting, and dashboards (Prometheus, Grafana, Splunk).
- Build automation and self-healing tooling that reduces manual toil and scales the team's operational capacity.
- Identify, escalate, and resolve production issues to protect platform reliability and customer experience.
- Champion customer success by helping internal teams understand platform capabilities and adopt
Skills
Communication
Degrees
AssociateBachelorDegree
Work schedule
On-callRotationShiftWeekend
Industry
AutomotiveEnergyOil-gas
Contract length
4 years