Staff AI Infrastructure Engineer

Engineering · AI Platform
San Francisco, United States
Staff
On-site

Job Overview

Join the AI Platform team to design GPU cluster scheduling, distributed training pipelines, and low-latency inference infrastructure. You will work closely with applied ML teams to make model training and serving fast, reliable, and cost-efficient.

Responsibilities

  • Operate and scale GPU training clusters for internal ML teams
  • Build tooling for distributed training job scheduling and checkpointing
  • Optimize inference serving latency and throughput
  • Establish best practices for reproducible ML infrastructure

Qualifications

  • 7+ years of experience in infrastructure or ML systems engineering
  • Hands-on experience with CUDA and GPU cluster operations
  • Experience with distributed training frameworks such as PyTorch
  • Comfort operating Kubernetes-based infrastructure at scale

Compensation & Benefits

This position has an estimated salary range of $245,000.00 - $300,000.00 per year, plus potential equity and bonus.

Key skills

Python
CUDA
Kubernetes
PyTorch
distributed training

Compliance Information

Equal Opportunity: Northwind is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws.

Hiring Disclosure: We are committed to providing equal employment opportunities to all employees and applicants for employment.

Labor Law Reference: Fair Labor Standards Act (FLSA), Title VII of the Civil Rights Act of 1964

Jobs