Jobs / Tre***

Senior Site Reliability and Infrastructure Engineer

Tre*** · New York, NY, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
New York, NY, United States160,000-220,000 USD/yearlyHybrid
Remuneration
160,000-220,000 USD/yearly
Location
New York, NY, United States
Visa sponsorship
Sponsors visa

Job summary

In the face of rising threats, increasing pressure on affordability, and unprecedented demand for power, Tre*** empowers energy companies to modernize their field work to meet the growth and challenges ahead.

Benefits

Our data pipeline, machine learning training platform, and web app could all benOffered; and the potential future value of any long-term incentives.This information is provided per the New York City Human Rights Law.Please note that the range provided is applicable only to New York City-based apBase compensation may vary if the work location is outside of New York City.Treeswift is proud to be an equal opportunity employer.

Qualifications

  • To tackle this challenge, we are bringing together a team of mission-driven experts with deep industry experience in robotics (Penn, Caltech, CMU) and enterprise software development (Palantir, Stripe, Oracle, MongoDB).
  • from a critical industry into high-quality data products, so understanding the business holistically is key.
  • We take pride in managing complexity and providing high-fidelity data that our customers can use to make better-informed decisions.
  • relevant work experience, location, and other factors.
  • This salary estimate excludes the value of any potential bonuses; the value of any

Responsibilities

  • You will build the observability and reliability foundations that let us run this system confidently as customer data volume grows: monitoring, alerting, performance/cost visibility, and clear operational practices.
  • Design and implement reliability and observability for high-volume pipeline operations, including:
  • actionable monitoring/alerting for DAG/task failures and reruns
  • visibility into operational workflows like flight orchestration (including DLQ/failed-message alerting and notification pathways)
  • dashboards and SLO/SLI definitions focused on correctness, throughput, and pipeline health
  • Make machine learning inference operations more reliable and observable:
  • instrument inference runs executed inside pipeline runners (model checkpoint resolution, S3 sync behavior, thresholds and fallback behavior, and output correctness)
  • add operational visibility for inference outcomes (e.g., unknown classification rates, fallback usage, and failure modes)
  • Create operational tooling and continuously improve systems (‘leave it better than you found it’), including:
  • runbooks, incident learnings, and engineering standards for debugging at scale
  • automate away toil in deployment and operations workflows as we learn what hurts most
  • On-call / incident response

Skills

CommunicationLeadership

Degrees

Associate

Work schedule

On-callRotationShift

Industry

AutomotiveEnergyOil-gas

Company size

EnterpriseSmbStartup

Contract length

10 years