Jobs / Int***
Senior Engineer, Platform & Site Reliability
Int*** · Jacksonville, FL, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Jacksonville, FL, United StatesOnsite
Remuneration
Not specified
Location
Jacksonville, FL, United States
Visa sponsorship
Sponsors visa
Job summary
Overview: Job Purpose As a key player within Int***'s (ICE) innovative servicing technology division, our team is dedicated to delivering cutting-edge mortgage processing solutions on a resilient, cloud-native platform.
Qualifications
- Writes technical specifications and operational runbooks based on conceptual design and stated business and reliability
- Develops and/or reviews automated tests and reliability validation before release, with an emphasis on Unit, Component, and Scenario tests.
- Troubleshoots operational failures in both test and production environments and leads root cause analysis.
- Mentors or guides the work of less experienced site reliability and software engineers.
- Remains current on industry standards in cloud, DevOps, SRE, and web development disciplines.
- Performs additional related
Responsibilities
- By joining our team, you will directly shape the reliability, performance, and operability of the platform that powers our business.
- Provisions, upgrades, and operates Amazon EKS (Kubernetes) clusters across multiple AWS regions and environments (UAT, stable, production, and chaos), ensuring a secure, scalable, and highly available platform.
- Manages cloud infrastructure declaratively through Kubernetes using Crossplane, and implements GitOps practices with ArgoCD to deliver all infrastructure as code.
- Operates and upgrades the Istio service mesh, including canary rollouts, along with Envoy and ingress, for traffic management, routing, and mutual TLS.
- Designs and operates monitoring, logging, and observability solutions using Prometheus, Grafana, Jaeger, OpenTelemetry (OTEL), Kiali, and Fluent Bit, and defines SLOs, alerting, and dashboards.
- Improves system reliability through capacity planning, autoscaling, incident response, on-call practices, and chaos engineering.
- Administers platform services such as cert-manager, external-dns, external-secrets, sealed-secrets, and the AWS Load Balancer Controller, and integrates with AWS services including DMS and MSK.
- Builds and maintains CI/CD pipelines (Azure DevOps) and automated security scanning (for example, SonarQube) to deliver changes safely and repeatably.
- Provides full-stack Java (Spring) and React (TypeScript) development for platform tooling and for the microservices and micro frontends that run on the platform.
- Designs and develops APIs and automation that support platform capabilities and self-service for product teams.
- Participates in reliability and architecture design ceremonies and analyzes system needs to determine technical
- as assigned.
Degrees
AssociateBachelorDegree
Work schedule
On-call
Industry
AutomotiveEducationEnergy
Company size
Smb
Security clearance
Secret