Skip to main content
T

Research Engineer - AI-Optimized Inference

Touring Capital

Location

Remote

Salary

Not specified

Type

fulltime

Posted

Today

via linkedin

Job Description

Infinity Artificial Intelligence Institute San Francisco Bay Area

Research Engineer - AI-Optimized Inference

Infinity Artificial Intelligence Institute San Francisco Bay Area

6 days ago 130 applicants

See who Infinity Artificial Intelligence Institute has hired for this role

Save

  • Report this job

Company

: Infinity

  • Team: Systems / AI Infrastructure Location: San Francisco (on-site)
  • Type: Full-time

The Mission

The fastest way we've found to make a kernel faster is to let an AI rewrite it and prove, empirically, that the rewrite actually won. Systems like AlphaEvolve made the shape of this loop clear: propose a change, evaluate it against the version it replaces, keep it only when it's measurably better, and repeat that thousands of times. What comes out the other side is code no person sat down and wrote, and it beats the code a person did.

This role points that loop directly at inference. The code being rewritten is the kernels that implement the operations running on an AI accelerator, along with the batching, data movement, and scheduling around them that determine how much of the chip's peak performance you actually get to keep. The mandate is concrete and unambiguous: serve an inference stack that is at least 50% faster than the inference libraries people already use today, measured end to end on real workloads rather than on a microbenchmark built to flatter the result.

We've run this play before. On Qwen3-8B, the loop took throughput from roughly 1,400 tokens per second to over 20,000 in a single day and beat vLLM by more than 13%, documented in our published research. Every kernel it produces feeds directly into the Infinity Kernel Registry, so a win found on one model or one chip compounds instead of disappearing. This role exists because that same discipline, an AI proposing changes with a human accountable for the evaluation that decides what counts as a win, is how Infinity intends to stay ahead of every hand-tuned inference library on the market, not just match one once.

What You'll Work On

You'll own the loop that rewrites inference until it beats the incumbent, end to end. Depending on your strengths:

  • The optimization loop itself. Generating a candidate rewrite, evaluating it against the current best on real inference workloads, and keeping it only when it wins outright. Making that evaluation fast, fair, and resistant to gaming is most of the actual difficulty here, not the code generation.
  • The kernels themselves. Matmul, attention, normalization, and collective operations running on the accelerator, rewritten and rewritten again until they close in on the chip's measured peak rather than its spec-sheet number.
  • The layers wrapped around the kernels. Batching, data movement, and scheduling, which is usually where the real gap between a kernel's individual peak and the stack's actual delivered throughput is hiding.
  • The 50% bar itself. Benchmarking honestly against the libraries people are actually serving with today, end to end, so a claimed win survives contact with a production workload instead of evaporating on the next model or batch size.
  • Correctness underneath all of it. A faster kernel that returns a different answer isn't faster, it's wrong, so every candidate gets checked against a reference implementation before its speed is allowed to count for anything.
  • Feedback into the kernel registry, so a win discovered on one model or one chip gets reused on the next instead of being rediscovered from scratch.

What we're looking for

We care about depth and range more than a checklist, but strong candidates will have most of the following:

  • Performance engineering on real inference or accelerator code. You've made an attention kernel, a matmul, or a serving path meaningfully faster before, and you can walk through exactly why the fix worked.
  • Familiarity with search- or evolution-based optimization. AlphaEvolve-style loops, superoptimization, autotuning, or a genuine appetite to build one of these systems from scratch if you haven't yet.
  • A solid grasp of inference internals: kernels, batching, KV cache, scheduling, and the specific places where throughput quietly leaks away.
  • Measurement discipline that refuses to be flattered by a benchmark chosen because it makes the number look good.
  • Python and a systems language, plus real hands-on kernel work in CUDA, ROCm/HIP, Triton, or something comparable.

Nice to have

  • Written high-performance attention or matmul kernels by hand, not just called into someone else's.
  • Worked on code evolution, superoptimizers, or autotuners such as Ansor or Triton's autotuning stack.
  • Contributed to vLLM, SGLang, TensorRT-LLM, or a comparable inference serving stack.
  • Built the evaluation harness at the center of an optimization loop, and learned firsthand how those harnesses get gamed.

Who we are

Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.

  • Seniority level Entry level
  • Employment type Full-time
  • Job function Engineering and Information Technology
  • Industries Software Development

Referrals increase your chances of interviewing at Infinity Artificial Intelligence Institute by 2x

See who you know

Get notified about new Research Engineer jobs in

San Francisco Bay Area

.

Sign in to create job alert

Similar jobs

  • ML Research Engineer - Hardware Codesign

ML Research Engineer - Hardware Codesign

OpenAI

San Francisco, CA $185,000 - $455,000 2 weeks ago

  • AI Inference Performance Engineer

AI Inference Performance Engineer

NVIDIA

Santa Clara, CA 2 weeks ago

  • Fellow, AI Workload Optimization

Fellow, AI Workload Optimization

AMD

San Jose, CA $256,000 - $384,000 20 hours ago

  • AI Systems, Model Optimization

AI Systems, Model Optimization

Unconventional AI

Palo Alto, CA 1 day ago

  • ML Architect, Hardware Software Co-Design (All Levels)

ML Architect, Hardware Software Co-Design (All Levels)

Rivian

Palo Alto, CA 2 weeks ago

  • Research Scientist / Engineer – Performance Optimization

Research Scientist / Engineer – Performance Optimization

Luma

San Francisco Bay Area 14 hours ago

  • Research Engineer [ Performance Engineering ]

Research Engineer [ Performance Engineering ]

Metamorphic

Palo Alto, CA 1 week ago

  • Research Engineer

Research Engineer

Lightning AI

San Francisco, CA 2 weeks ago

  • Staff AI Inference and Acceleration Engineer

Staff AI Inference and Acceleration Engineer

Figure

San Jose, CA 3 weeks ago

  • Performance Research Engineer (multiple levels)

Performance Research Engineer (multiple levels)

Efficient Computer

San Jose, CA 2 weeks ago

  • CoDesign \& NextGen Performance Engineer

CoDesign \& NextGen Performance Engineer

Cerebras

Sunnyvale, CA 1 week ago

  • Applied AI Research Engineer

Applied AI Research Engineer

Beam

San Francisco, CA

$140,000\.00

$200,000\.00

1 week ago

  • Hardware / Software CoDesign Engineer - 3P

Hardware / Software CoDesign Engineer - 3P

OpenAI

San Francisco, CA

$342,000\.00

$555,000\.00

2 weeks ago

  • Senior High-Performance AI Training Engineer

Senior High-Performance AI Training Engineer

NVIDIA

Santa Clara, CA 2 weeks ago

  • Member of Technical Staff - Research Engineer

Member of Technical Staff - Research Engineer

Black Forest Labs

San Francisco, CA 3 weeks ago

  • Research Engineer, Training \& Inference

Research Engineer, Training \& Inference

Harmonic

Palo Alto, CA 1 week ago

  • HPE Labs - Principal AI and Machine Learning Research Engineer

HPE Labs - Principal AI and Machine Learning Research Engineer

Hewlett Packard Enterprise

Milpitas, CA 5 months ago

  • AI Infra Engineer

AI Infra Engineer

Black Sesame Technologies Inc

San Jose, CA 2 weeks ago

  • Inference Optimization Engineer (local / edge runtime)

Inference Optimization Engineer (local / edge runtime)

Intel

Santa Clara, CA 2 weeks ago

  • Performance Modeling Engineer

Performance Modeling Engineer

Etched

San Jose, CA

$175,000\.00

$275,000\.00

6 days ago

  • Research Engineer - AI Performance \& Kernel Optimization

Research Engineer - AI Performance \& Kernel Optimization

Zyphra

San Francisco, CA 4 months ago

  • Member of Technical Staff - Edge Inference Engineer

Member of Technical Staff - Edge Inference Engineer

Liquid AI

San Francisco, CA 2 weeks ago

  • Member of Technical Staff, ML Systems

Member of Technical Staff, ML Systems

Netpreme

Santa Clara, CA 7 months ago

  • GPU AI Compute Architecture Engineer

GPU AI Compute Architecture Engineer

AMD

San Jose, CA

$172,000\.00

$258,000\.00

2 weeks ago

  • Applied Research Engineer

Applied Research Engineer

Zep AI

San Francisco, CA

$180,000\.00

$250,000\.00

2 months ago

  • ML Accelerator Architect

ML Accelerator Architect

Waymo

Mountain View, CA 2 months ago

  • Senior Performance Engineer

Senior Performance Engineer

Samsung Semiconductor

San Jose, CA 15 hours ago

People also viewed

  • Staff Engineer, TPU Co-Design

Staff Engineer, TPU Co-Design

Google

Sunnyvale, CA 1 week ago

  • Research Engineer, Core ML

Research Engineer, Core ML

Together AI

San Francisco, CA 4 days ago

  • AI Hardware Architecture

AI Hardware Architecture

Unconventional AI

Palo Alto, CA 2 weeks ago

  • Sr. AI Inference Systems Engineer

Sr. AI Inference Systems Engineer

Tencent

Palo Alto, CA 2 months ago

  • Senior Engineer, Performance Architecture

Senior Engineer, Performance Architecture

Samsung Semiconductor

San Jose, CA 2 days ago

  • Senior Developer Technology Engineer - AI

Senior Developer Technology Engineer - AI

NVIDIA AI

Santa Clara, CA 1 day ago

  • Member of Technical Staff (AI Inference Engineer)

Member of Technical Staff (AI Inference Engineer)

Perplexity

San Francisco, CA $220,000 - $485,000 2 weeks ago

  • AI Performance Engineer

AI Performance Engineer

The Biological Computing Co. (TBC)

San Francisco, CA 1 week ago

  • Senior Developer Technology Engineer - AI

Senior Developer Technology Engineer - AI

NVIDIA

Santa Clara, CA 2 weeks ago

  • Senior Deep Learning Performance Architect

Senior Deep Learning Performance Architect

NVIDIA

Santa Clara, CA 1 week ago

Similar Searches

  • Staff Research Engineer jobs

1,530 open jobs

  • Senior Research Engineer jobs

40,266 open jobs

  • Principal Research Engineer jobs

1,425 open jobs

  • Chemistry Physics Teacher jobs

10,440 open jobs

  • Senior System Test Engineer jobs

40,916 open jobs

  • Associate Research Engineer jobs

103 open jobs

  • Graduate Student Researcher jobs

6,176 open jobs

  • Standards Engineer jobs

43,317 open jobs

  • Optimization Engineer jobs

16,870 open jobs

  • Research And Development Engineer jobs

232,935 open jobs

  • Wireless System Engineer jobs

17,999 open jobs

  • Research Scientist jobs

24,159 open jobs

  • Academic Researcher jobs

3,644 open jobs

  • Senior Simulation Engineer jobs

2,082 open jobs

  • Bioinformatics Engineer jobs

228,351 open jobs

  • Medical Engineer jobs

13,791 open jobs

  • Embedded System Developer jobs

4,924 open jobs

  • Innovation Engineer jobs

15,615 open jobs

  • Algorithm Developer jobs

1,195 open jobs

  • Fisheries Biologist jobs

1,092 open jobs

  • Senior Materials Engineer jobs

3,555 open jobs

  • Senior Research And Development Engineer jobs

2,365 open jobs

  • Materials Engineer jobs

25,965 open jobs

  • Biological Science Technician jobs

16,603 open jobs

  • Modeling Engineer jobs

228,311 open jobs

Explore top content on LinkedIn

Find curated posts and insights for relevant topics all in one place.

View top content

Looking for more opportunities?

Browse thousands of graduate jobs and entry-level positions.

Browse All Jobs