Member of Technical Staff - AI-Optimized Inference

Touring Capital
Touring Capital

Software Engineering, IT, Data Science

San Francisco, CA, USA

USD 175k-350k / year

Posted on Aug 8, 2026
Infinity Artificial Intelligence Institute San Francisco Bay Area

Member of Technical Staff - AI-Optimized Inference

Infinity Artificial Intelligence Institute San Francisco Bay Area

22 hours ago 29 applicants

See who Infinity Artificial Intelligence Institute has hired for this role

Save

  • Report this job

Member of Technical - Staff AI-Optimized Inference

Company: Infinity

  • Team: Systems / AI Infrastructure Location: San Francisco (on-site)
  • Type: Full-time

The Mission

The fastest way we've found to make a kernel faster is to let an AI rewrite it and prove, empirically, that the rewrite actually won. Systems like AlphaEvolve made the shape of this loop clear: propose a change, evaluate it against the version it replaces, keep it only when it's measurably better, and repeat that thousands of times. What comes out the other side is code no person sat down and wrote, and it beats the code a person did.

This role points that loop directly at inference. The code being rewritten is the kernels that implement the operations running on an AI accelerator, along with the batching, data movement, and scheduling around them that determine how much of the chip's peak performance you actually get to keep. The mandate is concrete and unambiguous: serve an inference stack that is at least 50% faster than the inference libraries people already use today, measured end to end on real workloads rather than on a microbenchmark built to flatter the result.

We've run this play before. On Qwen3-8B, the loop took throughput from roughly 1,400 tokens per second to over 20,000 in a single day and beat vLLM by more than 13%, documented in our published research. Every kernel it produces feeds directly into the Infinity Kernel Registry, so a win found on one model or one chip compounds instead of disappearing. This role exists because that same discipline, an AI proposing changes with a human accountable for the evaluation that decides what counts as a win, is how Infinity intends to stay ahead of every hand-tuned inference library on the market, not just match one once.

What You'll Work On

You'll own the loop that rewrites inference until it beats the incumbent, end to end. Depending on your strengths:

  • The optimization loop itself. Generating a candidate rewrite, evaluating it against the current best on real inference workloads, and keeping it only when it wins outright. Making that evaluation fast, fair, and resistant to gaming is most of the actual difficulty here, not the code generation.
  • The kernels themselves. Matmul, attention, normalization, and collective operations running on the accelerator, rewritten and rewritten again until they close in on the chip's measured peak rather than its spec-sheet number.
  • The layers wrapped around the kernels. Batching, data movement, and scheduling, which is usually where the real gap between a kernel's individual peak and the stack's actual delivered throughput is hiding.
  • The 50% bar itself. Benchmarking honestly against the libraries people are actually serving with today, end to end, so a claimed win survives contact with a production workload instead of evaporating on the next model or batch size.
  • Correctness underneath all of it. A faster kernel that returns a different answer isn't faster, it's wrong, so every candidate gets checked against a reference implementation before its speed is allowed to count for anything.
  • Feedback into the kernel registry, so a win discovered on one model or one chip gets reused on the next instead of being rediscovered from scratch.

What We're Looking For

We care about depth and range more than a checklist, but strong candidates will have most of the following:

  • Performance engineering on real inference or accelerator code. You've made an attention kernel, a matmul, or a serving path meaningfully faster before, and you can walk through exactly why the fix worked.
  • Familiarity with search- or evolution-based optimization. AlphaEvolve-style loops, superoptimization, autotuning, or a genuine appetite to build one of these systems from scratch if you haven't yet.
  • A solid grasp of inference internals: kernels, batching, KV cache, scheduling, and the specific places where throughput quietly leaks away.
  • Measurement discipline that refuses to be flattered by a benchmark chosen because it makes the number look good.
  • Python and a systems language, plus real hands-on kernel work in CUDA, ROCm/HIP, Triton, or something comparable.

Nice to have

  • Written high-performance attention or matmul kernels by hand, not just called into someone else's.
  • Worked on code evolution, superoptimizers, or autotuners such as Ansor or Triton's autotuning stack.
  • Contributed to vLLM, SGLang, TensorRT-LLM, or a comparable inference serving stack.
  • Built the evaluation harness at the center of an optimization loop, and learned firsthand how those harnesses get gamed.

Who We Are

Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.

  • Seniority level Mid-Senior level
  • Employment type Full-time
  • Job function Engineering and Information Technology
  • Industries Software Development

Referrals increase your chances of interviewing at Infinity Artificial Intelligence Institute by 2x

See who you know

Get notified about new Member of Technical Staff jobs in San Francisco Bay Area.

Sign in to create job alert

Similar jobs

  • Member of Technical Staff (AI Inference Engineer)

Member of Technical Staff (AI Inference Engineer)

Perplexity

Palo Alto, CA $220,000 - $485,000 1 week ago

  • AI Infrastructure Engineer

AI Infrastructure Engineer

Intel

Santa Clara, CA 1 day ago

  • Member of Technical Staff

Member of Technical Staff

Morph

San Francisco, CA $175,000 - $350,000 6 days ago

  • Member of Technical Staff - Universal Inference

Member of Technical Staff - Universal Inference

Touring Capital

San Francisco Bay Area 11 hours ago

  • Member of Technical Staff, Inference

Member of Technical Staff, Inference

Inferact

San Francisco, CA $200,000 - $400,000 1 month ago

  • Senior Engineer - AI Agents and Systems

Senior Engineer - AI Agents and Systems

NVIDIA

Santa Clara, CA 1 week ago

  • Senior Software Engineer, AI Inference Systems

Senior Software Engineer, AI Inference Systems

NVIDIA

Santa Clara, CA 2 weeks ago

  • Member of Technical Staff, Kernels

Member of Technical Staff, Kernels

Inception

San Francisco Bay Area

$200,000.00

$350,000.00

2 weeks ago

  • Member of Technical Staff - Universal Inference

Member of Technical Staff - Universal Inference

Infinity Artificial Intelligence Institute

San Francisco Bay Area 1 day ago

  • Member of Technical Staff - RL Inference

Member of Technical Staff - RL Inference

SpaceXAI

Palo Alto, CA 1 week ago

  • Member of Technical Staff, Inference

Member of Technical Staff, Inference

Reactor

San Francisco, CA 3 months ago

  • Member of Technical Staff - Research Engineer

Member of Technical Staff - Research Engineer

Black Forest Labs

San Francisco, CA 1 month ago

  • Member of Technical Staff, Model Efficiency

Member of Technical Staff, Model Efficiency

Cohere

San Francisco, CA 2 weeks ago

  • Member of Technical Staff — Inference Palo Alto, CA

Member of Technical Staff — Inference Palo Alto, CA

RadixArk

Palo Alto, CA 1 month ago

  • Senior Software Engineer - Performance

Senior Software Engineer - Performance

Microsoft

Mountain View, CA 2 weeks ago

  • Staff AI Inference and Acceleration Engineer

Staff AI Inference and Acceleration Engineer

Figure

San Jose, CA 1 month ago

  • Member of Technical Staff - Edge Inference Engineer

Member of Technical Staff - Edge Inference Engineer

Liquid AI

San Francisco, CA 1 week ago

  • Software Engineer, Model Inference

Software Engineer, Model Inference

OpenAI

San Francisco, CA

$295,000.00

$555,000.00

2 weeks ago

  • Member of Technical Staff, AI Platform & Architecture (Infrastructure)

Member of Technical Staff, AI Platform & Architecture (Infrastructure)

Postman

San Francisco, CA 4 days ago

  • Member of Technical Staff - Model Serving / API Backend Engineer

Member of Technical Staff - Model Serving / API Backend Engineer

Black Forest Labs

San Francisco, CA 2 weeks ago

  • Senior Staff Engineer, AI Software

Senior Staff Engineer, AI Software

Samsung Semiconductor

San Jose, CA 2 weeks ago

  • Member of Technical Staff, Specialized Focus

Member of Technical Staff, Specialized Focus

Eventual

San Francisco, CA 15 hours ago

  • Member of Technical Staff - ML Systems & Inference

Member of Technical Staff - ML Systems & Inference

Gimlet Labs

San Francisco, CA

$150,000.00

$350,000.00

3 months ago

  • Member of Technical Staff — Developer Technology Palo Alto, CA

Member of Technical Staff — Developer Technology Palo Alto, CA

RadixArk

Palo Alto, CA 2 weeks ago

  • Staff Engineer, Inference Optimizations

Staff Engineer, Inference Optimizations

DigitalOcean

San Francisco, CA 2 weeks ago

  • Fellow Software Engineer — AI Performance & Reliability

Fellow Software Engineer — AI Performance & Reliability

AMD

San Jose, CA

$268,000.00

$402,000.00

1 week ago

  • Staff Software Engineer - AI Platform

Staff Software Engineer - AI Platform

LinkedIn

Mountain View, CA 1 week ago

People also viewed

  • Member of Technical Staff, Inference & RL Systems

Member of Technical Staff, Inference & RL Systems

Magic

San Francisco, CA $225,000 - $550,000 1 week ago

  • Member of Technical Staff - Inference

Member of Technical Staff - Inference

Prime Intellect

San Francisco, CA 2 months ago

  • Member of Technical Staff, ML Performance

Member of Technical Staff, ML Performance

Odyssey

Palo Alto, CA 4 months ago

  • Opensource Al workload Software Engineer

Opensource Al workload Software Engineer

AMD

San Jose, CA $172,000 - $258,000 2 weeks ago

  • Staff Software Engineer - GenAI inference

Staff Software Engineer - GenAI inference

Databricks

San Francisco, CA 4 days ago

  • Member of Technical Staff - Applied AI Research

Member of Technical Staff - Applied AI Research

Gimlet Labs

San Francisco, CA $150,000 - $350,000 3 months ago

  • Systems Generalist, GPT Infrastructure

Systems Generalist, GPT Infrastructure

OpenAI

San Francisco, CA $293,000 - $445,000 3 hours ago

  • Senior AI Platform Engineer

Senior AI Platform Engineer

Adobe

San Jose, CA 5 days ago

  • Member of Technical Staff, Performance Optimization

Member of Technical Staff, Performance Optimization

Fireworks AI

San Mateo, CA 6 days ago

  • Staff ML Compiler Engineer

Staff ML Compiler Engineer

General Motors

Sunnyvale, CA 3 days ago

Similar Searches

  • Senior Member of Technical Staff jobs

8,067 open jobs

  • Member Technical jobs

62,993 open jobs

  • Senior Wireless Engineer jobs

35,731 open jobs

  • Physical Design Engineer jobs

6,753 open jobs

  • Staff Test Engineer jobs

2,677 open jobs

  • Principal Firmware Engineer jobs

1,780 open jobs

  • Line Technician jobs

109,890 open jobs

  • Lead Infrastructure Engineer jobs

16,227 open jobs

  • Support Team Manager jobs

83,600 open jobs

  • Vice President Software jobs

49,146 open jobs

  • Senior Lead Software Engineer jobs

49,381 open jobs

  • Principal Researcher jobs

4,530 open jobs

  • Switch Engineer jobs

9,391 open jobs

  • Staff Software Engineer jobs

64,945 open jobs

  • Lead Quality Engineer jobs

10,548 open jobs

  • Control Coordinator jobs

39,876 open jobs

  • Market Maker jobs

1,432 open jobs

  • Yield Engineer jobs

9,445 open jobs

  • Computer Scientist jobs

49,477 open jobs

  • Lead Test Engineer jobs

13,921 open jobs

  • House Supervisor jobs

29,485 open jobs

  • Core Engineer jobs

33,936 open jobs

  • Cable Technician jobs

11,124 open jobs

  • Principal Software Engineer jobs

73,845 open jobs

  • Logic Design Engineer jobs

1,858 open jobs

Explore top content on LinkedIn

Find curated posts and insights for relevant topics all in one place.

View top content