TTheTopTechJobs
HomePricingMCPJobsTrackerInsights
HomePricingMCPJobsTrackerInsights
TTheTopTechJobs
HomePricingMCPJobsTrackerInsights
HomePricingMCPJobsTrackerInsights
Loading…

TheTopTechJobs

The insider board: fresh roles you won't find on LinkedIn. Curated niche markets for product managers, engineering managers, software engineers, and product designers.

Browse JobsUnlock Fresh

Job Markets

  • US Product Manager Jobs
  • US Engineering Manager Jobs
  • US Software Engineer Jobs
  • US Product Designer Jobs

Jobs by Location

  • San Francisco
  • New York
  • Seattle
  • Austin
  • Chicago
  • Los Angeles
  • Virginia
  • Ohio
  • New Jersey
  • Massachusetts
  • Florida
  • Illinois
  • Colorado
  • North Carolina
  • All locations →

Jobs by Type

  • Remote
  • Hybrid
  • On-Site
  • Senior
  • Staff
  • Director
  • All types →

Product

  • Home
  • Pricing
  • MCP Automations
  • H-1B Sponsors
  • Insights
  • About & Methodology
  • Privacy
  • Terms

Account

Insider board for niche US roles
Privacy·Terms·© 2026 Boolean and Bean Pty Ltd · TheTopTechJobs is a product of Boolean and Bean Pty Ltd
Loading…
Loading…
← All jobs
JPMorgan Chase logo

Software Engineer III - Machine Learning Platform

Posted in last 14 days
JPMorgan Chase
ON SITE
FULL TIME

Palo Alto, CA, United States

Salary context

Typical pay for US Software Engineer Jobs roles: $140k–$195k (median $160k).

Based on 783 live listings with disclosed salary.

Dept: Corporate Sector

Related searches

US Software Engineer Jobs in Palo Alto

Research this role and its H-1B context

Use these guides to compare role fit, employer history, and public sponsorship evidence. An employer's past filing activity does not guarantee sponsorship for this opening or for a particular candidate.

  • career guideCareer guideA practical, data-led report for comparing recent employer H-1B filing history with live product, engineering, software, and design openings.
Apply Now

Description

We have an exciting and rewarding opportunity for you to take your software engineering career to the next level. As a Software Engineer III at JPMorganChase within the AI/ML data platform team you serve as a seasoned member of an agile team to build and operate scalable, reliable ML training systems and pipelines on AWS and other cloud platforms. You will productionize training workloads (often GPU-based), improve performance and cost efficiency, and enable repeatable, well-governed training across environments Job Responsibilities - Design, build, and maintain end-to-end ML training platform. - Run and optimize GPU training workloads (single-node and distributed), improving throughput, utilization and reproducibility. - Build and operate training infrastructure on Kubernetes (e.g., EKS and other manage Kubernetes platforms), including resource management and workload troubleshooting. - Enable Gen AI/LLM training and fine-tuning workflows (e.g., supervised fine-tuning), including evaluation harnesses, artifact/version governance, and scalable GPU execution patterns aligned to enterprise controls. - Implement observability for training systems: metrics, logs, dashboards, alerting, and operational runbooks. - Partner with data engineering and platform teams to define interfaces, standards, and guardrails (security, access, cost controls) - Improve developer experience for training: standardized containers, CI/CD, templates, documentation, and self-service workflow - Leverages enterprise-authorized AI coding assist tools within the work environment to improve code quality, delivery speed, and productivity across complex deliverables (e.g., code generation/refactoring, unit test creation, documentation), while validating outputs through peer review, automated testing, and secure coding standards; contributes learnings and reusable patterns to improve broader team effectiveness. - Applies knowledge of tools within the Software Development Life Cycle toolchain, including enterprise-authorized AI-assisted development and automation capabilities, to improve the value realized by automation. Required qualifications, capabilities, and skills - Formal training or certification on software engineering concepts and 3+ years applied experience - Demonstrated experience running ML training in cloud environments and debugging issues across infrastructure & code. - Strong Python skills with solid engineering practices (testing, code reviews, modular design, dependency management). - Experience building automation/CI for ML codebases (build, test, release, deployment/promotion workflows). - Hands on experience with deep learning training workflows and at least one major framework (eg., PyTorch or TensorFlow). - Understanding of training performance and stability: data loading bottlenecks, mixed precision, checkpointing, reproducibility, and evaluation methodology. - Experience with distributed training and related concepts (e.g., DDP/FSDP/DeepSpeed concepts, collective communication basics, scaling and bottleneck analysis). - Ability to profile and optimize training systems (CPU/GPU utilization, memory, I/O throughput, networking, scheduling). - Experience with Kubernetes fundamentals for running compute-intensive workloads and AWS (eg., EKS/ECR, S3, IAM, VPC/networking, Cloudwatch, EC2) - Hands-on experience using enterprise-authorized AI-assisted software development tools within the work environment (e.g., for coding, test creation, troubleshooting, or documentation) with demonstrated ability to critically evaluate, validate, and refine AI-generated outputs for correctness, performance, and security. - Understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; ability to guide peers on safe and effective usage within team practices. Preferred qualifications, capabilities, and skills - Experience running training workloads across multiple cloud platforms and managing portability, performance, and governance across environments. - Familiarity with cloud-native networking/storage patterns for high-throughput training and artifact management. - Experience optimizing training input pipelines (sharding, prefetching, caching, format choices such as Parquet/WebDataset) and working with large datasets. - Familiarity with distributed compute frameworks (Spark, Ray, Dask) for feature/dataset generation. - Familiarity with workflow orchestration tools (Airflow-like systems, Argo Workflows-like patterns) and model registry concepts. - Experience optimizing training cost/performance (right-sizing, scheduling policies, interruptible capacity strategies where applicable budge guardrails, quota planning). - Strong observability practice for training systems: metrics/logs/traces, GPU telemetry, dashboards, and alert tuning. FEDERAL DEPOSIT INSURANCE ACT: This position is subject to Section 19 of the Federal Deposit Insurance Act. As such, an employment offer for this position is contingent on JPMorganChase’s review of criminal conviction history, including pretrial diversions or program entries.
ATS: oracle hcmPosted: Aug 20, 2026Updated: Aug 24, 2026View original posting

Similar jobs

JPMorgan Chase logo
JPMorgan Chase
5 days ago

Software Engineer III - Machine Learning Platform

New
Palo Alto, CA, United StatesFull Time
ON SITE
JPMorgan Chase logo
JPMorgan Chase
2 days ago

Java Full stack Software Engineer III - React/Python

New
Jersey City, NJ, United StatesFull Time
ON SITE
JPMorgan Chase logo
JPMorgan Chase
4 days ago

Full-Stack Java/Python React Software Engineer III- Trading applications

New
Jersey City, NJ, United StatesFull Time
ON SITE
JPMorgan Chase logo
JPMorgan Chase
4 days ago

Director of Software Engineering - Linux OS

New
Columbus, OH, United StatesFull Time
ON SITE

Software Engineer III - Machine Learning Platform

Apply Now