Jobs / Tre***
Senior Site Reliability and Infrastructure Engineer
Tre*** · New York, NY, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
New York, NY, United States160,000-220,000 USD/yearlyHybrid
Remuneration
160,000-220,000 USD/yearly
Location
New York, NY, United States
Visa sponsorship
Sponsors visa
Job summary
In the face of rising threats, increasing pressure on affordability, and unprecedented demand for power, Tre*** empowers energy companies to modernize their field work to meet the growth and challenges ahead.
Benefits
Our data pipeline, machine learning training platform, and web app could all benOffered; and the potential future value of any long-term incentives.This information is provided per the New York City Human Rights Law.Please note that the range provided is applicable only to New York City-based apBase compensation may vary if the work location is outside of New York City.Treeswift is proud to be an equal opportunity employer.
Qualifications
- To tackle this challenge, we are bringing together a team of mission-driven experts with deep industry experience in robotics (Penn, Caltech, CMU) and enterprise software development (Palantir, Stripe, Oracle, MongoDB).
- from a critical industry into high-quality data products, so understanding the business holistically is key.
- We take pride in managing complexity and providing high-fidelity data that our customers can use to make better-informed decisions.
- relevant work experience, location, and other factors.
- This salary estimate excludes the value of any potential bonuses; the value of any
Responsibilities
- You will build the observability and reliability foundations that let us run this system confidently as customer data volume grows: monitoring, alerting, performance/cost visibility, and clear operational practices.
- Design and implement reliability and observability for high-volume pipeline operations, including:
- actionable monitoring/alerting for DAG/task failures and reruns
- visibility into operational workflows like flight orchestration (including DLQ/failed-message alerting and notification pathways)
- dashboards and SLO/SLI definitions focused on correctness, throughput, and pipeline health
- Make machine learning inference operations more reliable and observable:
- instrument inference runs executed inside pipeline runners (model checkpoint resolution, S3 sync behavior, thresholds and fallback behavior, and output correctness)
- add operational visibility for inference outcomes (e.g., unknown classification rates, fallback usage, and failure modes)
- Create operational tooling and continuously improve systems (‘leave it better than you found it’), including:
- runbooks, incident learnings, and engineering standards for debugging at scale
- automate away toil in deployment and operations workflows as we learn what hurts most
- On-call / incident response
Skills
CommunicationLeadership
Degrees
Associate
Work schedule
On-callRotationShift
Industry
AutomotiveEnergyOil-gas
Company size
EnterpriseSmbStartup
Contract length
10 years