Skip to main content
C

Member of Technical Staff, Inference Systems

Confidential

Location

California, United States

Salary

Not specified

Type

fulltime

Posted

Today

via linkedin

Job Description

Member of Technical Staff, Inference Systems (On-site)

Location: On-site (5 days/week in-office - Palo Alto, CA)

Salary: $230K-$350K \+ competitive equity

Sponsorship: open to visa transfers (OPT, H1B) and new sponsorships (H1B, TN)

We're hiring a Member of Technical Staff, Inference Systems at a well-funded, founding-stage team building a high-performance AI inference platform from the ground up. You'll build a new inference runtime in Rust, owning batching, scheduling, request routing, and the full serving stack.

This is a from-scratch build, not a wrapper around existing tools. The team is architecting the entire runtime with latency, throughput, and cost per token as first-order concerns. Every core architectural decision is still open, and you'll be one of the people making them.

The problem you'd help solve:

Serving LLMs at scale is a systems problem, not a model problem. Throughput and cost per token are decided by scheduling, batching, KV cache management, and how well the runtime uses GPUs across nodes. Most teams inherit these decisions from a general-purpose engine. This team is building the runtime itself, with no legacy constraints, for engineers who want to work on inference internals rather than around them.

You'll build the serving runtime from scratch in Rust, design KV cache management and prefix caching that cut latency and cost per token, scale serving across multi-GPU and multi-node setups, profile and benchmark the full inference pipeline, and work directly with the founding team on the architecture that defines the platform.

You'll likely be a fit if you have

:

  • 2-10 years of experience as a backend, systems, or distributed systems engineer
  • Deep familiarity with inference internals: attention, KV cache, batching, scheduling
  • Hands-on time inside engines like vLLM, SGLang, or TensorRT-LLM, or experience building LLM serving infrastructure
  • Strong systems programming skills in Rust, C\+\+, or similar, and a genuine willingness to work in Rust day to day
  • Experience profiling and optimizing performance-critical systems
  • Comfort owning an entire stack rather than a narrow slice

Nice to have:

  • Multi-GPU and multi-node serving experience
  • CUDA, Triton, or NCCL experience
  • Production Rust
  • Early-stage startup experience

What you won't find here:

This role won't suit you if you want remote or hybrid work, or if you prefer a defined scope with clear boundaries. The team is small, the pace is high, and you'll be shaping architecture rather than picking up well-specified tickets. You'll be on-site five days a week in Palo Alto.

Looking for more opportunities?

Browse thousands of graduate jobs and entry-level positions.

Browse All Jobs