Research Engineer - AI-Optimized Inference
Software Engineering, Data Science
San Francisco, CA, USA
Posted on Jul 23, 2026
Infinity Artificial Intelligence Institute San Francisco Bay Area
Research Engineer - AI-Optimized Inference
Infinity Artificial Intelligence Institute San Francisco Bay Area
6 days ago 141 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
The fastest way we've found to make a kernel faster is to let an AI rewrite it and prove, empirically, that the rewrite actually won. Systems like AlphaEvolve made the shape of this loop clear: propose a change, evaluate it against the version it replaces, keep it only when it's measurably better, and repeat that thousands of times. What comes out the other side is code no person sat down and wrote, and it beats the code a person did.
This role points that loop directly at inference. The code being rewritten is the kernels that implement the operations running on an AI accelerator, along with the batching, data movement, and scheduling around them that determine how much of the chip's peak performance you actually get to keep. The mandate is concrete and unambiguous: serve an inference stack that is at least 50% faster than the inference libraries people already use today, measured end to end on real workloads rather than on a microbenchmark built to flatter the result.
We've run this play before. On Qwen3-8B, the loop took throughput from roughly 1,400 tokens per second to over 20,000 in a single day and beat vLLM by more than 13%, documented in our published research. Every kernel it produces feeds directly into the Infinity Kernel Registry, so a win found on one model or one chip compounds instead of disappearing. This role exists because that same discipline, an AI proposing changes with a human accountable for the evaluation that decides what counts as a win, is how Infinity intends to stay ahead of every hand-tuned inference library on the market, not just match one once.
What You'll Work On
You'll own the loop that rewrites inference until it beats the incumbent, end to end. Depending on your strengths:
We care about depth and range more than a checklist, but strong candidates will have most of the following:
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.
See who you know
Get notified about new Research Engineer jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
OpenAI
San Francisco, CA $185,000 - $455,000 2 weeks ago
NVIDIA
Santa Clara, CA 3 weeks ago
AMD
San Jose, CA $256,000 - $384,000 1 day ago
Unconventional AI
Palo Alto, CA 1 day ago
Rivian
Palo Alto, CA 2 weeks ago
Luma
San Francisco Bay Area 1 day ago
Metamorphic
Palo Alto, CA 1 week ago
Lightning AI
San Francisco, CA 3 weeks ago
Figure
San Jose, CA 3 weeks ago
Efficient Computer
San Jose, CA 2 weeks ago
Cerebras
Sunnyvale, CA 1 week ago
Beam
San Francisco, CA
$140,000.00
$200,000.00
1 week ago
OpenAI
San Francisco, CA
$342,000.00
$555,000.00
2 weeks ago
NVIDIA
Santa Clara, CA 3 weeks ago
Black Forest Labs
San Francisco, CA 3 weeks ago
Harmonic
Palo Alto, CA 1 week ago
Hewlett Packard Enterprise
Milpitas, CA 5 months ago
Black Sesame Technologies Inc
San Jose, CA 2 weeks ago
Intel
Santa Clara, CA 2 weeks ago
Etched
San Jose, CA
$175,000.00
$275,000.00
1 week ago
Zyphra
San Francisco, CA 4 months ago
Liquid AI
San Francisco, CA 2 weeks ago
Netpreme
Santa Clara, CA 7 months ago
AMD
San Jose, CA
$172,000.00
$258,000.00
2 weeks ago
Zep AI
San Francisco, CA
$180,000.00
$250,000.00
2 months ago
Waymo
Mountain View, CA 2 months ago
Samsung Semiconductor
San Jose, CA 1 day ago
People also viewed
Google
Sunnyvale, CA 2 weeks ago
Together AI
San Francisco, CA 5 days ago
Unconventional AI
Palo Alto, CA 2 weeks ago
Tencent
Palo Alto, CA 2 months ago
Samsung Semiconductor
San Jose, CA 3 days ago
NVIDIA AI
Santa Clara, CA 1 day ago
Perplexity
San Francisco, CA $220,000 - $485,000 2 weeks ago
The Biological Computing Co. (TBC)
San Francisco, CA 1 week ago
NVIDIA
Santa Clara, CA 21 minutes ago
NVIDIA
Santa Clara, CA 2 weeks ago
Similar Searches
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content
Research Engineer - AI-Optimized Inference
Infinity Artificial Intelligence Institute San Francisco Bay Area
6 days ago 141 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
- Report this job
- Team: Systems / AI Infrastructure Location: San Francisco (on-site)
- Type: Full-time
The fastest way we've found to make a kernel faster is to let an AI rewrite it and prove, empirically, that the rewrite actually won. Systems like AlphaEvolve made the shape of this loop clear: propose a change, evaluate it against the version it replaces, keep it only when it's measurably better, and repeat that thousands of times. What comes out the other side is code no person sat down and wrote, and it beats the code a person did.
This role points that loop directly at inference. The code being rewritten is the kernels that implement the operations running on an AI accelerator, along with the batching, data movement, and scheduling around them that determine how much of the chip's peak performance you actually get to keep. The mandate is concrete and unambiguous: serve an inference stack that is at least 50% faster than the inference libraries people already use today, measured end to end on real workloads rather than on a microbenchmark built to flatter the result.
We've run this play before. On Qwen3-8B, the loop took throughput from roughly 1,400 tokens per second to over 20,000 in a single day and beat vLLM by more than 13%, documented in our published research. Every kernel it produces feeds directly into the Infinity Kernel Registry, so a win found on one model or one chip compounds instead of disappearing. This role exists because that same discipline, an AI proposing changes with a human accountable for the evaluation that decides what counts as a win, is how Infinity intends to stay ahead of every hand-tuned inference library on the market, not just match one once.
What You'll Work On
You'll own the loop that rewrites inference until it beats the incumbent, end to end. Depending on your strengths:
- The optimization loop itself. Generating a candidate rewrite, evaluating it against the current best on real inference workloads, and keeping it only when it wins outright. Making that evaluation fast, fair, and resistant to gaming is most of the actual difficulty here, not the code generation.
- The kernels themselves. Matmul, attention, normalization, and collective operations running on the accelerator, rewritten and rewritten again until they close in on the chip's measured peak rather than its spec-sheet number.
- The layers wrapped around the kernels. Batching, data movement, and scheduling, which is usually where the real gap between a kernel's individual peak and the stack's actual delivered throughput is hiding.
- The 50% bar itself. Benchmarking honestly against the libraries people are actually serving with today, end to end, so a claimed win survives contact with a production workload instead of evaporating on the next model or batch size.
- Correctness underneath all of it. A faster kernel that returns a different answer isn't faster, it's wrong, so every candidate gets checked against a reference implementation before its speed is allowed to count for anything.
- Feedback into the kernel registry, so a win discovered on one model or one chip gets reused on the next instead of being rediscovered from scratch.
We care about depth and range more than a checklist, but strong candidates will have most of the following:
- Performance engineering on real inference or accelerator code. You've made an attention kernel, a matmul, or a serving path meaningfully faster before, and you can walk through exactly why the fix worked.
- Familiarity with search- or evolution-based optimization. AlphaEvolve-style loops, superoptimization, autotuning, or a genuine appetite to build one of these systems from scratch if you haven't yet.
- A solid grasp of inference internals: kernels, batching, KV cache, scheduling, and the specific places where throughput quietly leaks away.
- Measurement discipline that refuses to be flattered by a benchmark chosen because it makes the number look good.
- Python and a systems language, plus real hands-on kernel work in CUDA, ROCm/HIP, Triton, or something comparable.
- Written high-performance attention or matmul kernels by hand, not just called into someone else's.
- Worked on code evolution, superoptimizers, or autotuners such as Ansor or Triton's autotuning stack.
- Contributed to vLLM, SGLang, TensorRT-LLM, or a comparable inference serving stack.
- Built the evaluation harness at the center of an optimization loop, and learned firsthand how those harnesses get gamed.
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.
- Seniority level Entry level
- Employment type Full-time
- Job function Engineering and Information Technology
- Industries Software Development
See who you know
Get notified about new Research Engineer jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
- ML Research Engineer - Hardware Codesign
OpenAI
San Francisco, CA $185,000 - $455,000 2 weeks ago
- AI Inference Performance Engineer
NVIDIA
Santa Clara, CA 3 weeks ago
- Fellow, AI Workload Optimization
AMD
San Jose, CA $256,000 - $384,000 1 day ago
- AI Systems, Model Optimization
Unconventional AI
Palo Alto, CA 1 day ago
- ML Architect, Hardware Software Co-Design (All Levels)
Rivian
Palo Alto, CA 2 weeks ago
- Research Scientist / Engineer – Performance Optimization
Luma
San Francisco Bay Area 1 day ago
- Research Engineer [ Performance Engineering ]
Metamorphic
Palo Alto, CA 1 week ago
- Research Engineer
Lightning AI
San Francisco, CA 3 weeks ago
- Staff AI Inference and Acceleration Engineer
Figure
San Jose, CA 3 weeks ago
- Performance Research Engineer (multiple levels)
Efficient Computer
San Jose, CA 2 weeks ago
- CoDesign & NextGen Performance Engineer
Cerebras
Sunnyvale, CA 1 week ago
- Applied AI Research Engineer
Beam
San Francisco, CA
$140,000.00
$200,000.00
1 week ago
- Hardware / Software CoDesign Engineer - 3P
OpenAI
San Francisco, CA
$342,000.00
$555,000.00
2 weeks ago
- Senior High-Performance AI Training Engineer
NVIDIA
Santa Clara, CA 3 weeks ago
- Member of Technical Staff - Research Engineer
Black Forest Labs
San Francisco, CA 3 weeks ago
- Research Engineer, Training & Inference
Harmonic
Palo Alto, CA 1 week ago
- HPE Labs - Principal AI and Machine Learning Research Engineer
Hewlett Packard Enterprise
Milpitas, CA 5 months ago
- AI Infra Engineer
Black Sesame Technologies Inc
San Jose, CA 2 weeks ago
- Inference Optimization Engineer (local / edge runtime)
Intel
Santa Clara, CA 2 weeks ago
- Performance Modeling Engineer
Etched
San Jose, CA
$175,000.00
$275,000.00
1 week ago
- Research Engineer - AI Performance & Kernel Optimization
Zyphra
San Francisco, CA 4 months ago
- Member of Technical Staff - Edge Inference Engineer
Liquid AI
San Francisco, CA 2 weeks ago
- Member of Technical Staff, ML Systems
Netpreme
Santa Clara, CA 7 months ago
- GPU AI Compute Architecture Engineer
AMD
San Jose, CA
$172,000.00
$258,000.00
2 weeks ago
- Applied Research Engineer
Zep AI
San Francisco, CA
$180,000.00
$250,000.00
2 months ago
- ML Accelerator Architect
Waymo
Mountain View, CA 2 months ago
- Senior Performance Engineer
Samsung Semiconductor
San Jose, CA 1 day ago
People also viewed
- Co-Design Engineer, Google Cloud TPU
Sunnyvale, CA 2 weeks ago
- Research Engineer, Core ML
Together AI
San Francisco, CA 5 days ago
- AI Hardware Architecture
Unconventional AI
Palo Alto, CA 2 weeks ago
- Sr. AI Inference Systems Engineer
Tencent
Palo Alto, CA 2 months ago
- Senior Engineer, Performance Architecture
Samsung Semiconductor
San Jose, CA 3 days ago
- Senior Developer Technology Engineer - AI
NVIDIA AI
Santa Clara, CA 1 day ago
- Member of Technical Staff (AI Inference Engineer)
Perplexity
San Francisco, CA $220,000 - $485,000 2 weeks ago
- AI Performance Engineer
The Biological Computing Co. (TBC)
San Francisco, CA 1 week ago
- Senior Developer Technology Engineer - AI
NVIDIA
Santa Clara, CA 21 minutes ago
- Senior Deep Learning Performance Architect
NVIDIA
Santa Clara, CA 2 weeks ago
Similar Searches
- Staff Research Engineer jobs
- Senior Research Engineer jobs
- Principal Research Engineer jobs
- Chemistry Physics Teacher jobs
- Senior System Test Engineer jobs
- Associate Research Engineer jobs
- Graduate Student Researcher jobs
- Standards Engineer jobs
- Optimization Engineer jobs
- Research And Development Engineer jobs
- Wireless System Engineer jobs
- Research Scientist jobs
- Academic Researcher jobs
- Senior Simulation Engineer jobs
- Bioinformatics Engineer jobs
- Medical Engineer jobs
- Embedded System Developer jobs
- Innovation Engineer jobs
- Algorithm Developer jobs
- Fisheries Biologist jobs
- Senior Materials Engineer jobs
- Senior Research And Development Engineer jobs
- Materials Engineer jobs
- Biological Science Technician jobs
- Modeling Engineer jobs
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content