Jobs / App***

Site Reliability Engineer, Apple Data Platform / Big Data Platform

App*** · Austin, TX, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Austin, TX, United StatesOnsite
Remuneration
Not specified
Location
Austin, TX, United States
Visa sponsorship
Sponsors visa

Job summary

The App*** Services Engineering team (ASE) is one of the most exciting examples of App***'s long-held passion for combining art and technology. These are the people who power the App Store, App*** TV, App*** Music, App*** Podcasts, and App*** Books - at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.

Qualifications

  • Experience with REST Catalog services (e.g., Glue Catalog) and data governance frameworks.
  • Prior experience in a customer-facing or technical support role, with a demonstrated passion for customer success.
  • Familiarity with observability tooling: Prometheus, Grafana, Splunk, PagerDuty.
  • Working knowledge of CI/CD pipelines and deployment workflows.
  • Experience with S3 and cloud storage/networking fundamentals.
  • Familiarity with data pipeline orchestration and workflow scheduling patterns.
  • A track record of automating manual operations through scripting or tooling.
  • Intellectual curiosity and a drive to keep learning - for yourself, your team, and the org.
  • Minimum
  • Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
  • 1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.
  • Proficient in Python; working knowledge of Golang a plus.

Responsibilities

  • Operate, monitor, and triage production and non-production environments across the ADP portfolio - data processing, ML/AI, and multi-cloud infrastructure.
  • Participate in a rotating on-call schedule across supported services, including occasional weekday and weekend coverage.
  • Own the operational health of big data platform services as SME - driving reliability, support, and customer guidance for Spark, Flink, Airflow, Trino, Notebooks, REST Catalog, and governance tooling.
  • Serve as a primary point of contact for internal customers via Slack - clearly communicating status, root cause, and next steps during active issues.
  • Screen, triage, and resolve customer-reported service issues and support tickets, prioritizing based on customer impact and urgency.
  • Partner with dev teams across time zones to onboard new services - understanding architecture, then designing monitoring, alerting, and dashboards (Prometheus, Grafana, Splunk).
  • Build automation and self-healing tooling that reduces manual toil and scales the team's operational capacity.
  • Identify, escalate, and resolve production issues to protect platform reliability and customer experience.
  • Champion customer success by helping internal teams understand platform capabilities and adopt

Skills

Communication

Degrees

AssociateBachelorDegree

Work schedule

On-callRotationShiftWeekend

Industry

AutomotiveEnergyOil-gas

Contract length

4 years