Jobs / JPM***
Software Engineer III - DevOps/ML Ops AWS Streaming
JPM*** · Columbus, OH, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Columbus, OH, United StatesRemote
Remuneration
Not specified
Location
Columbus, OH, United States
Visa sponsorship
Sponsors visa
Job summary
JOB DESCRIPTION You're ready to gain the skills and experience needed to grow within your role and advance your career — and we have the perfect software engineering opportunity for you.
Benefits
And programs to meet employee needs, based on eligibility.Additional details about total compensation andWill be provided during the hiring process.We recognize that our people are our strength and the diverse talents they bringWe are an equal opportunity employer and place a high value on diversity and incVisit our FAQs for more information about requesting an accommodation.Equal Opportunity Employer/Disability/VeteransABOUT THE TEAMWe're proud to lead the U.In credit card sales and deposit growth and have the most-used digital solutions
Qualifications
- Capabilities, and
Responsibilities
- Provision and manage Amazon Web Services Kubernetes and container environments, including networking, access controls, cluster configuration, and autoscaling, using Terraform to support consistent, repeatable delivery.
- Build reusable infrastructure-as-code modules, define standards, and reduce configuration drift through strong environment hygiene and remote state management practices.
- Enable end-to-end machine learning operations workflows across build, validation, packaging, deployment, monitoring, and retraining to support production model lifecycle needs.
- Implement robust deployment and release patterns for machine learning-enabled services, including versioning, progressive delivery, and rollback strategies aligned to reliability goals.
- Operate and troubleshoot Apache Kafka streaming integrations, addressing throughput, latency, resiliency, and consumer lag to sustain real-time decisioning workloads.
- Deploy, operate, and scale Apache Flink on Kubernetes, managing job lifecycle, state, checkpoints, upgrades, recovery, and performance tuning.
- Deploy, maintain, and scale Ray Serve on Kubernetes to provide low-latency inference and service orchestration, improving resource efficiency and runtime stability.
- Implement and continuously improve observability (logs, metrics, tracing), dashboards, and alerting, and lead incident response and corrective actions to reduce mean time to recovery.
- Leverages enterprise-authorized AI coding assist
Degrees
Associate
Industry
AutomotiveBankingHealthcareMedia
Company size
EnterpriseSmb
Security clearance
Secret