Research Engineer, Inference
Job Overview
Northwind runs frontier-scale models for millions of daily requests. As a Research Engineer on Inference you will work at the boundary of research and systems: profiling and optimizing serving kernels, designing batching and caching strategies, and shipping the latency wins directly into production.
Responsibilities
- Profile and optimize inference throughput and latency across large model deployments
- Design continuous batching, KV-cache and quantization strategies for production serving
- Work with researchers to make new architectures servable without regressing quality
- Own serving benchmarks and defend them against regressions release over release
Qualifications
- Strong systems background with a track record of shipping performance work
- Hands-on experience with GPU programming, CUDA kernels, or Triton
- Deep familiarity with transformer inference and modern serving stacks
- Comfortable reading current research and turning it into production systems
Compensation & Benefits
This position has an estimated salary range of $360,000.00 - $530,000.00 per year, plus potential equity and bonus.
Key skills
Compliance Information
Equal Opportunity: Northwind is an equal opportunity employer and complies with all applicable federal, state, and local fair employment practices laws.
Hiring Disclosure: We are committed to providing equal employment opportunities to all employees and applicants for employment.
Labor Law Reference: Fair Labor Standards Act (FLSA), Title VII of the Civil Rights Act of 1964