Job Overview
Large training runs fail in boring ways: a bad NIC, a straggling node, a checkpoint that will not restore. You will own the reliability and throughput of our training fleet so that research time is spent on research.
Responsibilities
- Own scheduling, health checking, and failure recovery across large GPU clusters
- Drive up goodput on multi-week training runs and cut time lost to stragglers
- Build the observability researchers use to understand what their run is doing
- Work with hardware and networking partners on fleet-level issues
Qualifications
- 7+ years in infrastructure, distributed systems, or high-performance computing
- Direct experience operating large GPU or accelerator fleets
- Fluency with collective communication libraries and interconnect performance
- Strong debugging instincts under time pressure
Compensation & Benefits
This position has an estimated salary range of $320,000.00 - $450,000.00 per year, plus potential equity and bonus.
Key skills
Kubernetes
NCCL
Go
Python
InfiniBand
Compliance Information
Equal Opportunity: Northwind is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws.
Hiring Disclosure: We are committed to providing equal employment opportunities to all employees and applicants for employment.
Labor Law Reference: Fair Labor Standards Act (FLSA), Title VII of the Civil Rights Act of 1964