Jobs / App***

Site Reliability Engineer, Apple Data Platform / Multi-Cloud Infrastructure

App*** · Austin, TX, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Austin, TX, United StatesOnsite
Remuneration
Not specified
Location
Austin, TX, United States
Visa sponsorship
Sponsors visa

Job summary

The App*** Services Engineering team (ASE) is one of the most exciting examples of App***'s long-held passion for combining art and technology. These are the people who power the App Store, App*** TV, App*** Music, App*** Podcasts, and App*** Books - at extensive scale, meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 countries.

Qualifications

  • Experience with Infrastructure-as-Code (Crossplane and/or Terraform), including debugging state drift and composition/controller issues.
  • Experience with GitOps workflows (Flux or similar) - HelmRepository/reconciliation troubleshooting and Helm chart deployment.
  • Multi-cloud exposure (GCP) - parity and migration scenarios are emerging areas of focus.
  • Experience with Splunk for log pipeline debugging (e.g., fluent-bit).
  • Familiarity with Spark/Flink running on Kubernetes (executor scheduling, node affinity).
  • Comfort with GitHub PR review workflows in an infrastructure-as-code / GitOps context.
  • A track record of automating manual operations through scripting or tooling.
  • Intellectual curiosity and a drive to keep learning - for yourself, your team, and the org.
  • Minimum
  • Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
  • 1-4 years in a Site Reliability Engineering, DevOps, or Infrastructure-focused role.

Responsibilities

  • Operate, monitor, and triage production and non-production environments across the ADP portfolio - data processing, ML/AI, and multi-cloud infrastructure.
  • Participate in a rotating on-call schedule across supported services, including occasional weekday and weekend coverage.
  • Own the operational health of multi-cloud infrastructure as SME - driving reliability, support, and customer guidance for AWS services, EKS clusters, and cross-cloud networking.
  • Provide Slack-based support to internal customers; screen, triage, and resolve service related issues.
  • Debug production incidents involving IAM permission errors, storage quota limits, control-plane/data-plane namespace separation, and cluster-wide disruptions.
  • Partner with dev teams across time zones to onboard new services - understanding architecture, then designing monitoring, alerting, and dashboards (Prometheus, Grafana, Splunk).
  • Maintain and evolve Infrastructure-as-Code (Crossplane, Terraform) and GitOps (Flux) workflows, troubleshooting state drift and reconciliation issues.
  • Build automation and self-healing tooling that reduces manual toil and scales the team's operational capacity.
  • Identify, escalate, and resolve production issues to protect platform reliability and customer experience.
  • Collaborate with SRE and dev partner teams, engineering, and program management to align execution with team and org goals.
  • Preferred

Skills

Communication

Degrees

AssociateBachelorDegree

Work schedule

On-callRotationShiftWeekend

Industry

AutomotiveEnergyOil-gas

Contract length

4 years