Jobs / App***

Senior Software Engineer

App*** · Cupertino, CA, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Cupertino, CA, United States184,700-324,800 USD/yearlyOnsite
Remuneration
184,700-324,800 USD/yearly
Location
Cupertino, CA, United States
Visa sponsorship
Sponsors visa

Job summary

As part of the Siri organization, you will build the systems and tooling that make evaluation a first-class part of how Siri is developed - not an after-the-fact check - spanning human evaluation, real user feedback, reward and alignment signals, and data science rigor across iOS, iPadOS, macOS, watchOS, and visionOS.

Benefits

At Apple, base pay is one part of our total compensation package and is determinThis provides the opportunity to progress as you grow and develop within a role.The base pay range for this role is between $184,700 and $324,800, and your baseIncluding: Comprehensive medical and dental coverage, retirementAdditionally, this role might be eligible for discretionary bonuses or commissioLearn more about AppleNote: Apple benefit, compensation and employee stock programs are subject to eli

Qualifications

  • toward concrete, automated coverage
  • Turning product goals into measurable system behavior - instrumenting the product, building eval harnesses, and creating test datasets grounded in real user workflows
  • Building and improving reward models and alignment signals that measure whether Siri responses meet user needs
  • Designing and shipping evaluation tooling, pipelines, and architecture end-to-end - from data ingestion through scoring to monitoring in production - with an eye toward observability, logging, and reproducibility
  • Preferred
  • Experience evaluating ML, LLM, or agent-based systems, including familiarity with metrics, scoring methodology, trajectory and outcome analysis, and techniques like prompting, RAG, or LLM as judge
  • Understanding of reinforcement learning and the underlying techniques behind modern LLMs (e.g.
  • transformer architectures, RLHF/RLAIF, fine-tuning, reward modeling ) and frameworks such as PyTorch or Hugging Face, as applied to evaluation and reward signal design
  • Familiarity with eval-driven development - defining success criteria and test cases from product goals and real user workflows rather than abstract benchmarks
  • Experience with data science methods applied to quality measurement - defining ground truth, measuring inter-rater agreement (e.g.
  • Cohen's/Fleiss' kappa), and validating automated scorers using basic statistical techniques (e.g.
  • correlation, confidence intervals, hypothesis testing)

Responsibilities

  • In this role you'll contribute across several interconnected work streams spanning evaluation quality, reward/alignment signals, and data science.
  • This is a largely unexplored space with few established playbooks, so being self-driven is a must - you'll define your own path as much as execute one.
  • Supporting the evaluation of new Siri features and interaction modalities, working from ambiguous early

Skills

Communication

Degrees

AssociateBachelor

Work schedule

Shift

Industry

AutomotiveEducation

Company size

Smb

Relocation

Yes