Jobs / JPM***

Software Engineer III - DevOps/ML Ops AWS Streaming

JPM*** · Columbus, OH, United States
Visa sponsorship details are locked. Unlock company name and apply link with .
Columbus, OH, United StatesRemote
Remuneration
Not specified
Location
Columbus, OH, United States
Visa sponsorship
Sponsors visa

Job summary

JOB DESCRIPTION You're ready to gain the skills and experience needed to grow within your role and advance your career — and we have the perfect software engineering opportunity for you.

Benefits

And programs to meet employee needs, based on eligibility.Additional details about total compensation andWill be provided during the hiring process.We recognize that our people are our strength and the diverse talents they bringWe are an equal opportunity employer and place a high value on diversity and incVisit our FAQs for more information about requesting an accommodation.Equal Opportunity Employer/Disability/VeteransABOUT THE TEAMWe're proud to lead the U.In credit card sales and deposit growth and have the most-used digital solutions

Qualifications

  • Capabilities, and

Responsibilities

  • Provision and manage Amazon Web Services Kubernetes and container environments, including networking, access controls, cluster configuration, and autoscaling, using Terraform to support consistent, repeatable delivery.
  • Build reusable infrastructure-as-code modules, define standards, and reduce configuration drift through strong environment hygiene and remote state management practices.
  • Enable end-to-end machine learning operations workflows across build, validation, packaging, deployment, monitoring, and retraining to support production model lifecycle needs.
  • Implement robust deployment and release patterns for machine learning-enabled services, including versioning, progressive delivery, and rollback strategies aligned to reliability goals.
  • Operate and troubleshoot Apache Kafka streaming integrations, addressing throughput, latency, resiliency, and consumer lag to sustain real-time decisioning workloads.
  • Deploy, operate, and scale Apache Flink on Kubernetes, managing job lifecycle, state, checkpoints, upgrades, recovery, and performance tuning.
  • Deploy, maintain, and scale Ray Serve on Kubernetes to provide low-latency inference and service orchestration, improving resource efficiency and runtime stability.
  • Implement and continuously improve observability (logs, metrics, tracing), dashboards, and alerting, and lead incident response and corrective actions to reduce mean time to recovery.
  • Leverages enterprise-authorized AI coding assist

Degrees

Associate

Industry

AutomotiveBankingHealthcareMedia

Company size

EnterpriseSmb

Security clearance

Secret