Jobs / App***
ML Software Engineer
App*** · Seattle, WA, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Seattle, WA, United States142,300-263,300 USD/yearlyRemote
Remuneration
142,300-263,300 USD/yearly
Location
Seattle, WA, United States
Visa sponsorship
Sponsors visa
Job summary
Our team builds the ML-inference stack that powers generative AI for App*** Intelligence's Private Cloud Compute - running on App*** Silicon in the datacenter, distributing work across on-SoC acceleration hardware and multi-node clusters. Built on Private Cloud Compute's privacy guarantees, we're growing the team to scale across more platforms and support a widening set of features.
Benefits
At Apple, base pay is one part of our total compensation package and is determinThis provides the opportunity to progress as you grow and develop within a role.The base pay range for this role is between $142,300 and $263,300, and your baseIncluding: Comprehensive medical and dental coverage, retirementAdditionally, this role might be eligible for discretionary bonuses or commissioLearn more about AppleNote: Apple benefit, compensation and employee stock programs are subject to eli
Qualifications
- We're a collection of highly skilled and friendly engineers who value each other's opinions and experience.
- We are a team of domain experts, each specializing in specific core subject areas, with broad collective experience across cloud software services and platforms.
- Low-level or close-to-the-metal work - systems programming, performance, or hardware/SoC bring-up.
- Server-side Swift, or RPC/networking stacks (gRPC, Protocol Buffers).
- Depth in ML inference serving optimizations - quantization, sparsity, batching, KV cache, tokenization, GPU acceleration - enough to optimize the system and reason about the trade-offs and
- it places on models (you won't be training them).
- Solid grasp of concurrency, async/streaming, resource lifecycle, and error/cancellation handling; bonus for App*** platform experience (XPC, Instruments, Swift Concurrency).
- Minimum
- 2 Years practical experience plus Bachelor's degree in Computer Science, Computer Engineering, or a related field - or equivalent practical experience.
- Experience building large-scale distributed systems that serve ML inference, reasoning across processes, hosts, and service tiers as well as model behavior under load.
- A ML performance-centric mindset - able to reason about latency/throughput trade-offs and to distinguish what can be solved at the system level from what requires model-system co-design.
- Strong in a systems or server language - Python, Go, Rust, Java, C++, or similar.
Responsibilities
- You will integrate inference code into a full service stack so that user traffic is served reliably and performantly, with a strong focus on code that is easy and safe to develop, update, and monitor in production.
Degrees
AssociateBachelorDegree
Work schedule
On-call
Industry
EducationEnergyMedia
Contract length
2 years
Relocation
Yes