Member of Technical Staff - AI-Optimized Inference
Software Engineering, IT, Data Science
San Francisco, CA, USA
USD 175k-350k / year
Posted on Aug 8, 2026
Infinity Artificial Intelligence Institute San Francisco Bay Area
Member of Technical Staff - AI-Optimized Inference
Infinity Artificial Intelligence Institute San Francisco Bay Area
22 hours ago 29 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
Company: Infinity
The fastest way we've found to make a kernel faster is to let an AI rewrite it and prove, empirically, that the rewrite actually won. Systems like AlphaEvolve made the shape of this loop clear: propose a change, evaluate it against the version it replaces, keep it only when it's measurably better, and repeat that thousands of times. What comes out the other side is code no person sat down and wrote, and it beats the code a person did.
This role points that loop directly at inference. The code being rewritten is the kernels that implement the operations running on an AI accelerator, along with the batching, data movement, and scheduling around them that determine how much of the chip's peak performance you actually get to keep. The mandate is concrete and unambiguous: serve an inference stack that is at least 50% faster than the inference libraries people already use today, measured end to end on real workloads rather than on a microbenchmark built to flatter the result.
We've run this play before. On Qwen3-8B, the loop took throughput from roughly 1,400 tokens per second to over 20,000 in a single day and beat vLLM by more than 13%, documented in our published research. Every kernel it produces feeds directly into the Infinity Kernel Registry, so a win found on one model or one chip compounds instead of disappearing. This role exists because that same discipline, an AI proposing changes with a human accountable for the evaluation that decides what counts as a win, is how Infinity intends to stay ahead of every hand-tuned inference library on the market, not just match one once.
What You'll Work On
You'll own the loop that rewrites inference until it beats the incumbent, end to end. Depending on your strengths:
We care about depth and range more than a checklist, but strong candidates will have most of the following:
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.
See who you know
Get notified about new Member of Technical Staff jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
Perplexity
Palo Alto, CA $220,000 - $485,000 1 week ago
Intel
Santa Clara, CA 1 day ago
Morph
San Francisco, CA $175,000 - $350,000 6 days ago
Touring Capital
San Francisco Bay Area 11 hours ago
Inferact
San Francisco, CA $200,000 - $400,000 1 month ago
NVIDIA
Santa Clara, CA 1 week ago
NVIDIA
Santa Clara, CA 2 weeks ago
Inception
San Francisco Bay Area
$200,000.00
$350,000.00
2 weeks ago
Infinity Artificial Intelligence Institute
San Francisco Bay Area 1 day ago
SpaceXAI
Palo Alto, CA 1 week ago
Reactor
San Francisco, CA 3 months ago
Black Forest Labs
San Francisco, CA 1 month ago
Cohere
San Francisco, CA 2 weeks ago
RadixArk
Palo Alto, CA 1 month ago
Microsoft
Mountain View, CA 2 weeks ago
Figure
San Jose, CA 1 month ago
Liquid AI
San Francisco, CA 1 week ago
OpenAI
San Francisco, CA
$295,000.00
$555,000.00
2 weeks ago
Postman
San Francisco, CA 4 days ago
Black Forest Labs
San Francisco, CA 2 weeks ago
Samsung Semiconductor
San Jose, CA 2 weeks ago
Eventual
San Francisco, CA 15 hours ago
Gimlet Labs
San Francisco, CA
$150,000.00
$350,000.00
3 months ago
RadixArk
Palo Alto, CA 2 weeks ago
DigitalOcean
San Francisco, CA 2 weeks ago
AMD
San Jose, CA
$268,000.00
$402,000.00
1 week ago
LinkedIn
Mountain View, CA 1 week ago
People also viewed
Magic
San Francisco, CA $225,000 - $550,000 1 week ago
Prime Intellect
San Francisco, CA 2 months ago
Odyssey
Palo Alto, CA 4 months ago
AMD
San Jose, CA $172,000 - $258,000 2 weeks ago
Databricks
San Francisco, CA 4 days ago
Gimlet Labs
San Francisco, CA $150,000 - $350,000 3 months ago
OpenAI
San Francisco, CA $293,000 - $445,000 3 hours ago
Adobe
San Jose, CA 5 days ago
Fireworks AI
San Mateo, CA 6 days ago
General Motors
Sunnyvale, CA 3 days ago
Similar Searches
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content
Member of Technical Staff - AI-Optimized Inference
Infinity Artificial Intelligence Institute San Francisco Bay Area
22 hours ago 29 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
- Report this job
Company: Infinity
- Team: Systems / AI Infrastructure Location: San Francisco (on-site)
- Type: Full-time
The fastest way we've found to make a kernel faster is to let an AI rewrite it and prove, empirically, that the rewrite actually won. Systems like AlphaEvolve made the shape of this loop clear: propose a change, evaluate it against the version it replaces, keep it only when it's measurably better, and repeat that thousands of times. What comes out the other side is code no person sat down and wrote, and it beats the code a person did.
This role points that loop directly at inference. The code being rewritten is the kernels that implement the operations running on an AI accelerator, along with the batching, data movement, and scheduling around them that determine how much of the chip's peak performance you actually get to keep. The mandate is concrete and unambiguous: serve an inference stack that is at least 50% faster than the inference libraries people already use today, measured end to end on real workloads rather than on a microbenchmark built to flatter the result.
We've run this play before. On Qwen3-8B, the loop took throughput from roughly 1,400 tokens per second to over 20,000 in a single day and beat vLLM by more than 13%, documented in our published research. Every kernel it produces feeds directly into the Infinity Kernel Registry, so a win found on one model or one chip compounds instead of disappearing. This role exists because that same discipline, an AI proposing changes with a human accountable for the evaluation that decides what counts as a win, is how Infinity intends to stay ahead of every hand-tuned inference library on the market, not just match one once.
What You'll Work On
You'll own the loop that rewrites inference until it beats the incumbent, end to end. Depending on your strengths:
- The optimization loop itself. Generating a candidate rewrite, evaluating it against the current best on real inference workloads, and keeping it only when it wins outright. Making that evaluation fast, fair, and resistant to gaming is most of the actual difficulty here, not the code generation.
- The kernels themselves. Matmul, attention, normalization, and collective operations running on the accelerator, rewritten and rewritten again until they close in on the chip's measured peak rather than its spec-sheet number.
- The layers wrapped around the kernels. Batching, data movement, and scheduling, which is usually where the real gap between a kernel's individual peak and the stack's actual delivered throughput is hiding.
- The 50% bar itself. Benchmarking honestly against the libraries people are actually serving with today, end to end, so a claimed win survives contact with a production workload instead of evaporating on the next model or batch size.
- Correctness underneath all of it. A faster kernel that returns a different answer isn't faster, it's wrong, so every candidate gets checked against a reference implementation before its speed is allowed to count for anything.
- Feedback into the kernel registry, so a win discovered on one model or one chip gets reused on the next instead of being rediscovered from scratch.
We care about depth and range more than a checklist, but strong candidates will have most of the following:
- Performance engineering on real inference or accelerator code. You've made an attention kernel, a matmul, or a serving path meaningfully faster before, and you can walk through exactly why the fix worked.
- Familiarity with search- or evolution-based optimization. AlphaEvolve-style loops, superoptimization, autotuning, or a genuine appetite to build one of these systems from scratch if you haven't yet.
- A solid grasp of inference internals: kernels, batching, KV cache, scheduling, and the specific places where throughput quietly leaks away.
- Measurement discipline that refuses to be flattered by a benchmark chosen because it makes the number look good.
- Python and a systems language, plus real hands-on kernel work in CUDA, ROCm/HIP, Triton, or something comparable.
- Written high-performance attention or matmul kernels by hand, not just called into someone else's.
- Worked on code evolution, superoptimizers, or autotuners such as Ansor or Triton's autotuning stack.
- Contributed to vLLM, SGLang, TensorRT-LLM, or a comparable inference serving stack.
- Built the evaluation harness at the center of an optimization loop, and learned firsthand how those harnesses get gamed.
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.
- Seniority level Mid-Senior level
- Employment type Full-time
- Job function Engineering and Information Technology
- Industries Software Development
See who you know
Get notified about new Member of Technical Staff jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
- Member of Technical Staff (AI Inference Engineer)
Perplexity
Palo Alto, CA $220,000 - $485,000 1 week ago
- AI Infrastructure Engineer
Intel
Santa Clara, CA 1 day ago
- Member of Technical Staff
Morph
San Francisco, CA $175,000 - $350,000 6 days ago
- Member of Technical Staff - Universal Inference
Touring Capital
San Francisco Bay Area 11 hours ago
- Member of Technical Staff, Inference
Inferact
San Francisco, CA $200,000 - $400,000 1 month ago
- Senior Engineer - AI Agents and Systems
NVIDIA
Santa Clara, CA 1 week ago
- Senior Software Engineer, AI Inference Systems
NVIDIA
Santa Clara, CA 2 weeks ago
- Member of Technical Staff, Kernels
Inception
San Francisco Bay Area
$200,000.00
$350,000.00
2 weeks ago
- Member of Technical Staff - Universal Inference
Infinity Artificial Intelligence Institute
San Francisco Bay Area 1 day ago
- Member of Technical Staff - RL Inference
SpaceXAI
Palo Alto, CA 1 week ago
- Member of Technical Staff, Inference
Reactor
San Francisco, CA 3 months ago
- Member of Technical Staff - Research Engineer
Black Forest Labs
San Francisco, CA 1 month ago
- Member of Technical Staff, Model Efficiency
Cohere
San Francisco, CA 2 weeks ago
- Member of Technical Staff — Inference Palo Alto, CA
RadixArk
Palo Alto, CA 1 month ago
- Senior Software Engineer - Performance
Microsoft
Mountain View, CA 2 weeks ago
- Staff AI Inference and Acceleration Engineer
Figure
San Jose, CA 1 month ago
- Member of Technical Staff - Edge Inference Engineer
Liquid AI
San Francisco, CA 1 week ago
- Software Engineer, Model Inference
OpenAI
San Francisco, CA
$295,000.00
$555,000.00
2 weeks ago
- Member of Technical Staff, AI Platform & Architecture (Infrastructure)
Postman
San Francisco, CA 4 days ago
- Member of Technical Staff - Model Serving / API Backend Engineer
Black Forest Labs
San Francisco, CA 2 weeks ago
- Senior Staff Engineer, AI Software
Samsung Semiconductor
San Jose, CA 2 weeks ago
- Member of Technical Staff, Specialized Focus
Eventual
San Francisco, CA 15 hours ago
- Member of Technical Staff - ML Systems & Inference
Gimlet Labs
San Francisco, CA
$150,000.00
$350,000.00
3 months ago
- Member of Technical Staff — Developer Technology Palo Alto, CA
RadixArk
Palo Alto, CA 2 weeks ago
- Staff Engineer, Inference Optimizations
DigitalOcean
San Francisco, CA 2 weeks ago
- Fellow Software Engineer — AI Performance & Reliability
AMD
San Jose, CA
$268,000.00
$402,000.00
1 week ago
- Staff Software Engineer - AI Platform
Mountain View, CA 1 week ago
People also viewed
- Member of Technical Staff, Inference & RL Systems
Magic
San Francisco, CA $225,000 - $550,000 1 week ago
- Member of Technical Staff - Inference
Prime Intellect
San Francisco, CA 2 months ago
- Member of Technical Staff, ML Performance
Odyssey
Palo Alto, CA 4 months ago
- Opensource Al workload Software Engineer
AMD
San Jose, CA $172,000 - $258,000 2 weeks ago
- Staff Software Engineer - GenAI inference
Databricks
San Francisco, CA 4 days ago
- Member of Technical Staff - Applied AI Research
Gimlet Labs
San Francisco, CA $150,000 - $350,000 3 months ago
- Systems Generalist, GPT Infrastructure
OpenAI
San Francisco, CA $293,000 - $445,000 3 hours ago
- Senior AI Platform Engineer
Adobe
San Jose, CA 5 days ago
- Member of Technical Staff, Performance Optimization
Fireworks AI
San Mateo, CA 6 days ago
- Staff ML Compiler Engineer
General Motors
Sunnyvale, CA 3 days ago
Similar Searches
- Senior Member of Technical Staff jobs
- Member Technical jobs
- Senior Wireless Engineer jobs
- Physical Design Engineer jobs
- Staff Test Engineer jobs
- Principal Firmware Engineer jobs
- Line Technician jobs
- Lead Infrastructure Engineer jobs
- Support Team Manager jobs
- Vice President Software jobs
- Senior Lead Software Engineer jobs
- Principal Researcher jobs
- Switch Engineer jobs
- Staff Software Engineer jobs
- Lead Quality Engineer jobs
- Control Coordinator jobs
- Market Maker jobs
- Yield Engineer jobs
- Computer Scientist jobs
- Lead Test Engineer jobs
- House Supervisor jobs
- Core Engineer jobs
- Cable Technician jobs
- Principal Software Engineer jobs
- Logic Design Engineer jobs
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content