Job Description
We have an opportunity to impact your career and provide an adventure where you can push the limits of what's possible.
As a Software Engineer III at JPMorganChase within the Firmwide LLM Serving Platform team, you are an integral part of an agile team that designs, builds, and operates the services that make large language models usable at scale. This is an infrastructure-meets-ML role: you don't need to be an ML researcher, but you should be excited to learn how model architectures and inference constraints translate into real production systems. You will contribute to a living platform where we optimize performance — pushing down latency, increasing throughput, maximizing GPU utilization, and eliminating waste across the request lifecycle.
Job responsibilities
Build core backend services for LLM inference, including request routing, batching, scheduling, streaming responses, and quota/limits.
Implement and maintain APIs and SDKs used b...