Location
Remote
Salary
Not specified
Type
fulltime
Posted
Today
Job Description
Infinity Artificial Intelligence Institute San Francisco Bay Area
Research Engineer - AI-Optimized Inference
Infinity Artificial Intelligence Institute San Francisco Bay Area
6 days ago 130 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
- Report this job
Company
: Infinity
- Team: Systems / AI Infrastructure Location: San Francisco (on-site)
- Type: Full-time
The Mission
The fastest way we've found to make a kernel faster is to let an AI rewrite it and prove, empirically, that the rewrite actually won. Systems like AlphaEvolve made the shape of this loop clear: propose a change, evaluate it against the version it replaces, keep it only when it's measurably better, and repeat that thousands of times. What comes out the other side is code no person sat down and wrote, and it beats the code a person did.
This role points that loop directly at inference. The code being rewritten is the kernels that implement the operations running on an AI accelerator, along with the batching, data movement, and scheduling around them that determine how much of the chip's peak performance you actually get to keep. The mandate is concrete and unambiguous: serve an inference stack that is at least 50% faster than the inference libraries people already use today, measured end to end on real workloads rather than on a microbenchmark built to flatter the result.
We've run this play before. On Qwen3-8B, the loop took throughput from roughly 1,400 tokens per second to over 20,000 in a single day and beat vLLM by more than 13%, documented in our published research. Every kernel it produces feeds directly into the Infinity Kernel Registry, so a win found on one model or one chip compounds instead of disappearing. This role exists because that same discipline, an AI proposing changes with a human accountable for the evaluation that decides what counts as a win, is how Infinity intends to stay ahead of every hand-tuned inference library on the market, not just match one once.
What You'll Work On
You'll own the loop that rewrites inference until it beats the incumbent, end to end. Depending on your strengths:
- The optimization loop itself. Generating a candidate rewrite, evaluating it against the current best on real inference workloads, and keeping it only when it wins outright. Making that evaluation fast, fair, and resistant to gaming is most of the actual difficulty here, not the code generation.
- The kernels themselves. Matmul, attention, normalization, and collective operations running on the accelerator, rewritten and rewritten again until they close in on the chip's measured peak rather than its spec-sheet number.
- The layers wrapped around the kernels. Batching, data movement, and scheduling, which is usually where the real gap between a kernel's individual peak and the stack's actual delivered throughput is hiding.
- The 50% bar itself. Benchmarking honestly against the libraries people are actually serving with today, end to end, so a claimed win survives contact with a production workload instead of evaporating on the next model or batch size.
- Correctness underneath all of it. A faster kernel that returns a different answer isn't faster, it's wrong, so every candidate gets checked against a reference implementation before its speed is allowed to count for anything.
- Feedback into the kernel registry, so a win discovered on one model or one chip gets reused on the next instead of being rediscovered from scratch.
What we're looking for
We care about depth and range more than a checklist, but strong candidates will have most of the following:
- Performance engineering on real inference or accelerator code. You've made an attention kernel, a matmul, or a serving path meaningfully faster before, and you can walk through exactly why the fix worked.
- Familiarity with search- or evolution-based optimization. AlphaEvolve-style loops, superoptimization, autotuning, or a genuine appetite to build one of these systems from scratch if you haven't yet.
- A solid grasp of inference internals: kernels, batching, KV cache, scheduling, and the specific places where throughput quietly leaks away.
- Measurement discipline that refuses to be flattered by a benchmark chosen because it makes the number look good.
- Python and a systems language, plus real hands-on kernel work in CUDA, ROCm/HIP, Triton, or something comparable.
Nice to have
- Written high-performance attention or matmul kernels by hand, not just called into someone else's.
- Worked on code evolution, superoptimizers, or autotuners such as Ansor or Triton's autotuning stack.
- Contributed to vLLM, SGLang, TensorRT-LLM, or a comparable inference serving stack.
- Built the evaluation harness at the center of an optimization loop, and learned firsthand how those harnesses get gamed.
Who we are
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.
- Seniority level Entry level
- Employment type Full-time
- Job function Engineering and Information Technology
- Industries Software Development
Referrals increase your chances of interviewing at Infinity Artificial Intelligence Institute by 2x
See who you know
Get notified about new Research Engineer jobs in
San Francisco Bay Area
.
Sign in to create job alert
Similar jobs
- ML Research Engineer - Hardware Codesign
ML Research Engineer - Hardware Codesign
OpenAI
San Francisco, CA $185,000 - $455,000 2 weeks ago
- AI Inference Performance Engineer
AI Inference Performance Engineer
NVIDIA
Santa Clara, CA 2 weeks ago
- Fellow, AI Workload Optimization
Fellow, AI Workload Optimization
AMD
San Jose, CA $256,000 - $384,000 20 hours ago
- AI Systems, Model Optimization
AI Systems, Model Optimization
Unconventional AI
Palo Alto, CA 1 day ago
- ML Architect, Hardware Software Co-Design (All Levels)
ML Architect, Hardware Software Co-Design (All Levels)
Rivian
Palo Alto, CA 2 weeks ago
- Research Scientist / Engineer – Performance Optimization
Research Scientist / Engineer – Performance Optimization
Luma
San Francisco Bay Area 14 hours ago
- Research Engineer [ Performance Engineering ]
Research Engineer [ Performance Engineering ]
Metamorphic
Palo Alto, CA 1 week ago
- Research Engineer
Research Engineer
Lightning AI
San Francisco, CA 2 weeks ago
- Staff AI Inference and Acceleration Engineer
Staff AI Inference and Acceleration Engineer
Figure
San Jose, CA 3 weeks ago
- Performance Research Engineer (multiple levels)
Performance Research Engineer (multiple levels)
Efficient Computer
San Jose, CA 2 weeks ago
- CoDesign \& NextGen Performance Engineer
CoDesign \& NextGen Performance Engineer
Cerebras
Sunnyvale, CA 1 week ago
- Applied AI Research Engineer
Applied AI Research Engineer
Beam
San Francisco, CA
$140,000\.00
$200,000\.00
1 week ago
- Hardware / Software CoDesign Engineer - 3P
Hardware / Software CoDesign Engineer - 3P
OpenAI
San Francisco, CA
$342,000\.00
$555,000\.00
2 weeks ago
- Senior High-Performance AI Training Engineer
Senior High-Performance AI Training Engineer
NVIDIA
Santa Clara, CA 2 weeks ago
- Member of Technical Staff - Research Engineer
Member of Technical Staff - Research Engineer
Black Forest Labs
San Francisco, CA 3 weeks ago
- Research Engineer, Training \& Inference
Research Engineer, Training \& Inference
Harmonic
Palo Alto, CA 1 week ago
- HPE Labs - Principal AI and Machine Learning Research Engineer
HPE Labs - Principal AI and Machine Learning Research Engineer
Hewlett Packard Enterprise
Milpitas, CA 5 months ago
- AI Infra Engineer
AI Infra Engineer
Black Sesame Technologies Inc
San Jose, CA 2 weeks ago
- Inference Optimization Engineer (local / edge runtime)
Inference Optimization Engineer (local / edge runtime)
Intel
Santa Clara, CA 2 weeks ago
- Performance Modeling Engineer
Performance Modeling Engineer
Etched
San Jose, CA
$175,000\.00
$275,000\.00
6 days ago
- Research Engineer - AI Performance \& Kernel Optimization
Research Engineer - AI Performance \& Kernel Optimization
Zyphra
San Francisco, CA 4 months ago
- Member of Technical Staff - Edge Inference Engineer
Member of Technical Staff - Edge Inference Engineer
Liquid AI
San Francisco, CA 2 weeks ago
- Member of Technical Staff, ML Systems
Member of Technical Staff, ML Systems
Netpreme
Santa Clara, CA 7 months ago
- GPU AI Compute Architecture Engineer
GPU AI Compute Architecture Engineer
AMD
San Jose, CA
$172,000\.00
$258,000\.00
2 weeks ago
- Applied Research Engineer
Applied Research Engineer
Zep AI
San Francisco, CA
$180,000\.00
$250,000\.00
2 months ago
- ML Accelerator Architect
ML Accelerator Architect
Waymo
Mountain View, CA 2 months ago
- Senior Performance Engineer
Senior Performance Engineer
Samsung Semiconductor
San Jose, CA 15 hours ago
People also viewed
- Staff Engineer, TPU Co-Design
Staff Engineer, TPU Co-Design
Sunnyvale, CA 1 week ago
- Research Engineer, Core ML
Research Engineer, Core ML
Together AI
San Francisco, CA 4 days ago
- AI Hardware Architecture
AI Hardware Architecture
Unconventional AI
Palo Alto, CA 2 weeks ago
- Sr. AI Inference Systems Engineer
Sr. AI Inference Systems Engineer
Tencent
Palo Alto, CA 2 months ago
- Senior Engineer, Performance Architecture
Senior Engineer, Performance Architecture
Samsung Semiconductor
San Jose, CA 2 days ago
- Senior Developer Technology Engineer - AI
Senior Developer Technology Engineer - AI
NVIDIA AI
Santa Clara, CA 1 day ago
- Member of Technical Staff (AI Inference Engineer)
Member of Technical Staff (AI Inference Engineer)
Perplexity
San Francisco, CA $220,000 - $485,000 2 weeks ago
- AI Performance Engineer
AI Performance Engineer
The Biological Computing Co. (TBC)
San Francisco, CA 1 week ago
- Senior Developer Technology Engineer - AI
Senior Developer Technology Engineer - AI
NVIDIA
Santa Clara, CA 2 weeks ago
- Senior Deep Learning Performance Architect
Senior Deep Learning Performance Architect
NVIDIA
Santa Clara, CA 1 week ago
Similar Searches
- Staff Research Engineer jobs
1,530 open jobs
- Senior Research Engineer jobs
40,266 open jobs
- Principal Research Engineer jobs
1,425 open jobs
- Chemistry Physics Teacher jobs
10,440 open jobs
- Senior System Test Engineer jobs
40,916 open jobs
- Associate Research Engineer jobs
103 open jobs
- Graduate Student Researcher jobs
6,176 open jobs
- Standards Engineer jobs
43,317 open jobs
- Optimization Engineer jobs
16,870 open jobs
- Research And Development Engineer jobs
232,935 open jobs
- Wireless System Engineer jobs
17,999 open jobs
- Research Scientist jobs
24,159 open jobs
- Academic Researcher jobs
3,644 open jobs
- Senior Simulation Engineer jobs
2,082 open jobs
- Bioinformatics Engineer jobs
228,351 open jobs
- Medical Engineer jobs
13,791 open jobs
- Embedded System Developer jobs
4,924 open jobs
- Innovation Engineer jobs
15,615 open jobs
- Algorithm Developer jobs
1,195 open jobs
- Fisheries Biologist jobs
1,092 open jobs
- Senior Materials Engineer jobs
3,555 open jobs
- Senior Research And Development Engineer jobs
2,365 open jobs
- Materials Engineer jobs
25,965 open jobs
- Biological Science Technician jobs
16,603 open jobs
- Modeling Engineer jobs
228,311 open jobs
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content
Looking for more opportunities?
Browse thousands of graduate jobs and entry-level positions.