Training Cluster Engineer

Engineering · AI Platform
Seattle, United States
Staff
Hybrid

Job Overview

Large training runs fail in boring ways: a bad NIC, a straggling node, a checkpoint that will not restore. You will own the reliability and throughput of our training fleet so that research time is spent on research.

Responsibilities

  • Own scheduling, health checking, and failure recovery across large GPU clusters
  • Drive up goodput on multi-week training runs and cut time lost to stragglers
  • Build the observability researchers use to understand what their run is doing
  • Work with hardware and networking partners on fleet-level issues

Qualifications

  • 7+ years in infrastructure, distributed systems, or high-performance computing
  • Direct experience operating large GPU or accelerator fleets
  • Fluency with collective communication libraries and interconnect performance
  • Strong debugging instincts under time pressure

Compensation & Benefits

This position has an estimated salary range of $320,000.00 - $450,000.00 per year, plus potential equity and bonus.

Key skills

Kubernetes
NCCL
Go
Python
InfiniBand

Compliance Information

Equal Opportunity: Northwind is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws.

Hiring Disclosure: We are committed to providing equal employment opportunities to all employees and applicants for employment.

Labor Law Reference: Fair Labor Standards Act (FLSA), Title VII of the Civil Rights Act of 1964

Jobs