Staff AI Infrastructure Engineer
Engineering · AI Platform
San Francisco, United States
Staff
On-site
Job Overview
Join the AI Platform team to design GPU cluster scheduling, distributed training pipelines, and low-latency inference infrastructure. You will work closely with applied ML teams to make model training and serving fast, reliable, and cost-efficient.
Responsibilities
- Operate and scale GPU training clusters for internal ML teams
- Build tooling for distributed training job scheduling and checkpointing
- Optimize inference serving latency and throughput
- Establish best practices for reproducible ML infrastructure
Qualifications
- 7+ years of experience in infrastructure or ML systems engineering
- Hands-on experience with CUDA and GPU cluster operations
- Experience with distributed training frameworks such as PyTorch
- Comfort operating Kubernetes-based infrastructure at scale
Compensation & Benefits
This position has an estimated salary range of $245,000.00 - $300,000.00 per year, plus potential equity and bonus.
Key skills
Python
CUDA
Kubernetes
PyTorch
distributed training
Compliance Information
Equal Opportunity: Northwind is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws.
Hiring Disclosure: We are committed to providing equal employment opportunities to all employees and applicants for employment.
Labor Law Reference: Fair Labor Standards Act (FLSA), Title VII of the Civil Rights Act of 1964