Jobs / Int***

Senior Engineer, Platform & Site Reliability

Int*** · Jacksonville, FL, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Jacksonville, FL, United StatesOnsite
Remuneration
Not specified
Location
Jacksonville, FL, United States
Visa sponsorship
Sponsors visa

Job summary

Overview: Job Purpose As a key player within Int***'s (ICE) innovative servicing technology division, our team is dedicated to delivering cutting-edge mortgage processing solutions on a resilient, cloud-native platform.

Qualifications

  • Writes technical specifications and operational runbooks based on conceptual design and stated business and reliability
  • Develops and/or reviews automated tests and reliability validation before release, with an emphasis on Unit, Component, and Scenario tests.
  • Troubleshoots operational failures in both test and production environments and leads root cause analysis.
  • Mentors or guides the work of less experienced site reliability and software engineers.
  • Remains current on industry standards in cloud, DevOps, SRE, and web development disciplines.
  • Performs additional related

Responsibilities

  • By joining our team, you will directly shape the reliability, performance, and operability of the platform that powers our business.
  • Provisions, upgrades, and operates Amazon EKS (Kubernetes) clusters across multiple AWS regions and environments (UAT, stable, production, and chaos), ensuring a secure, scalable, and highly available platform.
  • Manages cloud infrastructure declaratively through Kubernetes using Crossplane, and implements GitOps practices with ArgoCD to deliver all infrastructure as code.
  • Operates and upgrades the Istio service mesh, including canary rollouts, along with Envoy and ingress, for traffic management, routing, and mutual TLS.
  • Designs and operates monitoring, logging, and observability solutions using Prometheus, Grafana, Jaeger, OpenTelemetry (OTEL), Kiali, and Fluent Bit, and defines SLOs, alerting, and dashboards.
  • Improves system reliability through capacity planning, autoscaling, incident response, on-call practices, and chaos engineering.
  • Administers platform services such as cert-manager, external-dns, external-secrets, sealed-secrets, and the AWS Load Balancer Controller, and integrates with AWS services including DMS and MSK.
  • Builds and maintains CI/CD pipelines (Azure DevOps) and automated security scanning (for example, SonarQube) to deliver changes safely and repeatably.
  • Provides full-stack Java (Spring) and React (TypeScript) development for platform tooling and for the microservices and micro frontends that run on the platform.
  • Designs and develops APIs and automation that support platform capabilities and self-service for product teams.
  • Participates in reliability and architecture design ceremonies and analyzes system needs to determine technical
  • as assigned.

Degrees

AssociateBachelorDegree

Work schedule

On-call

Industry

AutomotiveEducationEnergy

Company size

Smb

Security clearance

Secret