Location
Palo Alto, CA
Salary
$230,000 - $350,000 /yearly
Type
fulltime
Posted
Today
Job Description
About the Company
We are partnering with an early-stage AI infrastructure startup building a next-generation inference platform from the ground up, with a relentless focus on performance. The founding team has deep roots in AI infrastructure, and the company is taking on the hardest problems in LLM serving: scheduling, KV cache management, request routing, and the runtime systems behind inference at scale.
The entire stack is being engineered in Rust, with latency, throughput, and cost efficiency as first-order concerns. The company is backed by notable investors and is currently in stealth ahead of a public announcement.
This is an opportunity to shape the architecture of a high-performance inference system with no legacy constraints, working alongside a small, elite team of systems engineers who care about the craft of building fast, reliable infrastructure for AI.
The Role
You will build a new LLM inference system from scratch, with a clean-slate architecture and no legacy code to work around. You will join an incredible founding team working on the hard problems in LLM serving, including KV cache management and complex request routing. This role suits performance-minded engineers who want their work to have a direct impact on latency and cost per token at the core of the platform.
You know inference internals well (attention, KV cache, batching, scheduling), and you want to own the whole serving stack rather than one narrow component.
What You'll Do
- Build a new inference runtime from scratch in Rust, owning batching, scheduling, request routing, and the full serving stack.
- Design and implement KV cache management, prefix caching, and other optimizations that reduce latency and cost per token.
- Scale serving across GPUs and nodes, and solve multi-GPU and multi-node challenges.
- Profile, benchmark, and ship performance improvements across the inference pipeline.
- Work closely with the founding team on the core architectural decisions that define the platform.
What We're Looking For
- 2-10 years of experience in backend distributed systems and/or LLM inference infrastructure.
- Hands-on experience with inference engines such as vLLM, SGLang, or TensorRT-LLM.
- Strong grounding in LLM serving concepts: KV cache, continuous batching, scheduling, and request routing.
- Rust: experience is strongly preferred.
- Comfort with intense, high-ownership work at an early-stage startup, and a collaborative working style.
- A CS background from a top college program.
Tech Stack
Rust, Python, PyTorch, C\+\+, Go, vLLM, SGLang, TensorRT-LLM, CUDA, Triton, NCCL
Compensation
Base salary is $230K–$350K, depending on experience and depth in inference systems. Equity is competitive, with details discussed during the interview process.
*Only those selected to move forward in the interview process will be contacted.
Looking for more opportunities?
Browse thousands of graduate jobs and entry-level positions.