Jobs / Cam***
Principal Site Reliability Engineer, Machine Learning
Cam*** · Cambridge, MA, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Cambridge, MA, United States142,000-177,600 USD/yearlyRemote
Remuneration
142,000-177,600 USD/yearly
Location
Cambridge, MA, United States
Visa sponsorship
Sponsors visa
Job summary
CMT is looking for a Principal Site Reliability Engineer I, Machine Learning to help us change the world. CMT has helped protect over 65 million drivers and prevent over 126,000 crashes worldwide. We build AI to solve some of the most difficult challenges in mobility — understanding and reducing risk, detecting crashes, and getting people life-saving help.
Benefits
Fair and competitive salary based onWork on a mission with real impact: crashes prevented, injuries avoided, lives pJoin an industry leader — 65 million drivers protected, powering 140+ programs aCMT is also Great Place to Work CertifiedBe part of the team inventing the future of mobility and road safetyMove fast, own outcomes, do work that mattersHigh ownership, small teams, and direct access to leadership — no layers betweenUnlimited PTO, flexible scheduling, competitive salary, annual performance bonusIncluding medical, dental, vision, and 401k matchSummer Fridays provide team members with half days to rechargeJoin one of our employee resource groups: Black, AAPI, LGBTQIA+, Women, Book CluComprehensive wellness, education, and employee assistance programs
Qualifications
- Bachelor's degree or equivalent years of experience and/or certification in a related field
- 7+ years working in Site Reliability Engineering or Information Technology
- Design and document systems, including writing and reviewing code, to automate away problems within your team's domain
- Intermediate to expert experience deploying and maintaining AWS services such as EC2, ECS, EKS, SQS, Lambda, Dynamo, RDS/Aurora, S3, and IAM
- Intermediate to expert experience monitoring services and applications using
Responsibilities
- Use independent judgment and discretion to own SLOs, error budgets, and the operational health of Ray clusters running on AWS EKS and Databricks workloads on AWS EC2 across multiple accounts and regions
- Maintain the observability of uptime, availability, and scalability of EKS Ray and Databricks workloads using CloudWatch and Datadog, including defining alerting that maps to SLOs
- Operate and tune EKS Ray workloads at scale including autoscaling, GPU scheduling, and automated failure recovery
- Manage Databricks on AWS including workspace administration, cluster policies, Unity Catalog, job orchestration, and IAM Roles and Policies
- Maintain ongoing cost visibility, cost optimization, and capacity planning across EC2 and EKS workloads, including through the use of On Demand Capacity Reservations and Spot lifecycle
- Perform ongoing maintenance of the underlying EC2 and EKS infrastructure, including regular security updates and operating system upgrades
- Codify everything as infrastructure-as-code using Terraform and CI/CD pipelines, enabling updates through Pull Requests with approval workflows, while also automating maintenance tasks to reduce toil
- Lead incident response for Data Science and Machine Learning platform outages, run blameless postmortems, and drive systemic remediation, including participating in an on-call rotation
- Complete any additional tasks as they arise
- Base Salary Range
- The base salary range for this position is: $142,000 to $177,600.
- This range is specifically for Cambridge, MA
Skills
Leadership
Degrees
AssociateBachelorDegree
Work schedule
On-callRotationShift
Industry
AutomotiveEducationEnergyInsuranceMediaOil-gas