Research Engineer, Inference

Engineering · Frontier Systems
San Francisco, United States
Staff
On-site

Job Overview

Northwind runs frontier-scale models for millions of daily requests. As a Research Engineer on Inference you will work at the boundary of research and systems: profiling and optimizing serving kernels, designing batching and caching strategies, and shipping the latency wins directly into production.

Responsibilities

  • Profile and optimize inference throughput and latency across large model deployments
  • Design continuous batching, KV-cache and quantization strategies for production serving
  • Work with researchers to make new architectures servable without regressing quality
  • Own serving benchmarks and defend them against regressions release over release

Qualifications

  • Strong systems background with a track record of shipping performance work
  • Hands-on experience with GPU programming, CUDA kernels, or Triton
  • Deep familiarity with transformer inference and modern serving stacks
  • Comfortable reading current research and turning it into production systems

Compensation & Benefits

This position has an estimated salary range of $360,000.00 - $530,000.00 per year, plus potential equity and bonus.

Key skills

CUDA
Triton
PyTorch
Python
distributed systems

Compliance Information

Equal Opportunity: Northwind is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws.

Hiring Disclosure: We are committed to providing equal employment opportunities to all employees and applicants for employment.

Labor Law Reference: Fair Labor Standards Act (FLSA), Title VII of the Civil Rights Act of 1964

Jobs