Member of Technical Staff - Universal Inference

Touring Capital
Touring Capital

IT

San Francisco, CA, USA

USD 200k-400k / year

Posted on Aug 8, 2026
Infinity Artificial Intelligence Institute San Francisco Bay Area

Member of Technical Staff - Universal Inference

Infinity Artificial Intelligence Institute San Francisco Bay Area

1 day ago 33 applicants

See who Infinity Artificial Intelligence Institute has hired for this role

Save

  • Report this job

Member of Technical Staff - Universal Inference

Company: Infinity

  • Team: Systems / AI Infrastructure
  • Location: San Francisco (on-site)
  • Type: Full-time

The Mission

Every accelerator that comes online needs an inference stack, and today that stack is hand-built per chip, per model, per optimization - a permanent, growing backlog of engineering work that scales linearly with the number of chips and models in the world.

We're building Infy: a global inference library, generated and continuously maintained by AI, that targets every major chip - NVIDIA, AMD, Trainium, TPU, Maia, Cerebras, Groq, Tenstorrent, and dozens more - across every model category, from LLMs to vision, audio, and robotics. Think vLLM, but for every accelerator on the market, kept automatically up to date as chips, models, and optimization techniques change.

What You'll Work On

Depending on your strengths, you'll own one or more layers of the system:

  • Optimization agents - the strategy → analyze → generate → curate loop that turns a chip + model pair into an optimized kernel or serving path, and the orchestrator that runs this across many chips and models in parallel.
  • Kernel optimization loop - the hierarchical agent system that iterates on individual kernels (matmul, attention variants, normalization, RoPE, MoE, collectives) against reference implementations and measured-peak performance gates, drawing on and growing a registry of thousands of known kernels.
  • Library generator - the system that takes a validated set of optimized components and assembles them into a real, installable inference library per chip, with a standards-compatible serving interface.
  • Hardware probe - agents that build a structured representation of a chip's memory hierarchy, compute regions, and inter-component communication characteristics, so optimization strategy can branch correctly on the chip's actual architecture.
  • Coverage tracking - the enablement and optimization tables that track, per chip and per model category, what's implemented and how fast it runs, and that drive prioritization of what to build next.
  • SOTA paper replication - agents that read inference-optimization papers and automatically implement and validate the techniques they describe, so the library's optimizations keep pace with published research rather than lagging behind it.
  • Continuous tracking - the pipeline that watches for new papers, new hardware, and new models, and feeds that into what Infy builds and rebuilds next.
  • Benchmark harness - the metrics layer (TTFT, TPOT, ITL, E2EL, goodput, and model-category-specific equivalents like image-generation latency) that every optimization is measured against.
  • Inference service - the gateway, router, and billing layers that turn the library into a running, AI-optimized inference service people can actually call.

What We're Looking For

We care more about depth and range than a specific checklist, but strong candidates will have most of:

  • Real experience with ML inference internals - vLLM or similar serving stacks, attention kernel variants, KV cache management, quantization, continuous batching.
  • Comfort working across heterogeneous hardware - GPU kernels (CUDA, ROCm/HIP, Triton), and ideally exposure to non-GPU execution models (systolic arrays, dataflow, wafer-scale, in-memory/analog compute).
  • Hands-on experience building with coding agents / LLMs - prompting, tool-use loops, evaluating and constraining model output, and designing systems where the model writes and optimizes code while tests and benchmarks catch regressions.
  • Fluency in Python and at least one systems language (Rust, C, or C++).
  • Comfort reading and implementing ideas directly from research papers, not just from existing open-source code.

Nice to have

  • Contributed to vLLM, TVM, LLVM, MLIR, or a hardware vendor's compiler/runtime stack.
  • Experience with distributed inference or training internals (NCCL, Megatron-LM, DeepSpeed) and collective-communication algorithms.
  • Familiarity with model categories beyond LLMs - vision, audio, multimodal, or robotics inference.
  • Experience building or maintaining a benchmark suite or performance regression system at scale.
  • Background reproducing results from ML systems papers (KernelBench-style benchmarks, evolutionary/search-based kernel optimization).

Who You Are

You want to build the thing that makes every chip's inference performance a solved problem instead of a standing engineering project. You think in terms of systems that scale across hardware and models rather than one-off implementations, you're energized by AI agents doing the generation and optimization work under a tight benchmark harness, and you want to help set the global standard for how inference libraries get built and maintained.

Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.

  • Seniority level Mid-Senior level
  • Employment type Full-time
  • Job function Engineering and Information Technology
  • Industries Software Development

Referrals increase your chances of interviewing at Infinity Artificial Intelligence Institute by 2x

See who you know

Get notified about new Member of Technical Staff jobs in San Francisco Bay Area.

Sign in to create job alert

Similar jobs

  • Member of Technical Staff, Inference

Member of Technical Staff, Inference

Inferact

San Francisco, CA $200,000 - $400,000 1 month ago

  • Member of Technical Staff

Member of Technical Staff

Morph

San Francisco, CA $175,000 - $350,000 6 days ago

  • Member of Technical Staff - RL Inference

Member of Technical Staff - RL Inference

SpaceXAI

Palo Alto, CA 1 week ago

  • Member of Technical Staff, Inference

Member of Technical Staff, Inference

Reactor

San Francisco, CA 3 months ago

  • Member of Technical Staff (AI Inference Engineer)

Member of Technical Staff (AI Inference Engineer)

Perplexity

Palo Alto, CA $220,000 - $485,000 1 week ago

  • Member of Technical Staff, Kernels

Member of Technical Staff, Kernels

Inception

San Francisco Bay Area $200,000 - $350,000 2 weeks ago

  • Member of Technical Staff - Research Engineer

Member of Technical Staff - Research Engineer

Black Forest Labs

San Francisco, CA 1 month ago

  • Member of Technical Staff, TPU Performance Engineering

Member of Technical Staff, TPU Performance Engineering

Inferact

San Francisco, CA

$200,000.00

$400,000.00

1 month ago

  • Member of Technical Staff, AI Platform & Architecture (Infrastructure)

Member of Technical Staff, AI Platform & Architecture (Infrastructure)

Postman

San Francisco, CA 4 days ago

  • Member of Technical Staff - AI-Optimized Inference

Member of Technical Staff - AI-Optimized Inference

Touring Capital

San Francisco Bay Area 11 hours ago

  • Member of Technical Staff - Inference

Member of Technical Staff - Inference

Prime Intellect

San Francisco, CA 2 months ago

  • Member of Technical Staff — Inference Palo Alto, CA

Member of Technical Staff — Inference Palo Alto, CA

RadixArk

Palo Alto, CA 1 month ago

  • Member of Technical Staff - Model Serving / API Backend Engineer

Member of Technical Staff - Model Serving / API Backend Engineer

Black Forest Labs

San Francisco, CA 2 weeks ago

  • Member of Technical Staff, Training Infra

Member of Technical Staff, Training Infra

Inception

San Francisco Bay Area

$200,000.00

$350,000.00

1 week ago

  • Member of Technical Staff, Model Efficiency

Member of Technical Staff, Model Efficiency

Cohere

San Francisco, CA 2 weeks ago

  • Member of Technical Staff, Inference & RL Systems

Member of Technical Staff, Inference & RL Systems

Magic

San Francisco, CA

$225,000.00

$550,000.00

1 week ago

  • AI Infrastructure Engineer

AI Infrastructure Engineer

Intel

Santa Clara, CA 1 day ago

  • Member of Technical Staff - AI-Optimized Inference

Member of Technical Staff - AI-Optimized Inference

Infinity Artificial Intelligence Institute

San Francisco Bay Area 22 hours ago

  • Staff Software Engineer - AI Platform

Staff Software Engineer - AI Platform

LinkedIn

Mountain View, CA 1 week ago

  • Member of Technical Staff — Developer Technology Palo Alto, CA

Member of Technical Staff — Developer Technology Palo Alto, CA

RadixArk

Palo Alto, CA 2 weeks ago

  • Staff ML Compiler Engineer

Staff ML Compiler Engineer

General Motors

Sunnyvale, CA 3 days ago

  • Staff ML Compiler Engineer

Staff ML Compiler Engineer

General Motors

San Francisco, CA 4 days ago

  • Member of Technical Staff - ML Systems & Inference

Member of Technical Staff - ML Systems & Inference

Gimlet Labs

San Francisco, CA

$150,000.00

$350,000.00

3 months ago

  • Member of Technical Staff - Mid-Training Infra

Member of Technical Staff - Mid-Training Infra

Reflection

San Francisco, CA 6 days ago

  • Member of Technical Staff, Training Infra Engineer

Member of Technical Staff, Training Infra Engineer

Cohere

San Francisco, CA 1 week ago

  • Member of Technical Staff, Specialized Focus

Member of Technical Staff, Specialized Focus

Eventual

San Francisco, CA 15 hours ago

  • Systems Generalist, GPT Infrastructure

Systems Generalist, GPT Infrastructure

OpenAI

San Francisco, CA

$293,000.00

$445,000.00

3 hours ago

People also viewed

  • Member of Technical Staff - Applied AI Research

Member of Technical Staff - Applied AI Research

Gimlet Labs

San Francisco, CA $150,000 - $350,000 3 months ago

  • Senior Software Engineer, AI Inference Systems

Senior Software Engineer, AI Inference Systems

NVIDIA

Santa Clara, CA 2 weeks ago

  • Member of Technical Staff, ML Performance

Member of Technical Staff, ML Performance

Odyssey

Palo Alto, CA 4 months ago

  • Member of Technical Staff - Platform

Member of Technical Staff - Platform

Vals AI

San Francisco, CA 1 month ago

  • Software Engineer, Model Inference

Software Engineer, Model Inference

OpenAI

San Francisco, CA $295,000 - $555,000 2 weeks ago

  • Staff AI Inference and Acceleration Engineer

Staff AI Inference and Acceleration Engineer

Figure

San Jose, CA 1 month ago

  • Staff Engineer, Inference Optimizations

Staff Engineer, Inference Optimizations

DigitalOcean

San Francisco, CA 2 weeks ago

  • Member of Technical Staff, Backend

Member of Technical Staff, Backend

ComfyUI

San Francisco, CA 7 months ago

  • Member of Technical Staff

Member of Technical Staff

Simplify

San Francisco, CA 3 days ago

  • Member of Technical Staff - Edge Inference Engineer

Member of Technical Staff - Edge Inference Engineer

Liquid AI

San Francisco, CA 1 week ago

Similar Searches

  • Senior Member of Technical Staff jobs

8,067 open jobs

  • Member Technical jobs

62,993 open jobs

  • Senior Wireless Engineer jobs

35,731 open jobs

  • Physical Design Engineer jobs

6,753 open jobs

  • Staff Test Engineer jobs

2,677 open jobs

  • Principal Firmware Engineer jobs

1,780 open jobs

  • Line Technician jobs

109,890 open jobs

  • Lead Infrastructure Engineer jobs

16,227 open jobs

  • Support Team Manager jobs

83,600 open jobs

  • Vice President Software jobs

49,146 open jobs

  • Senior Lead Software Engineer jobs

49,381 open jobs

  • Principal Researcher jobs

4,530 open jobs

  • Switch Engineer jobs

9,391 open jobs

  • Staff Software Engineer jobs

64,945 open jobs

  • Lead Quality Engineer jobs

10,548 open jobs

  • Control Coordinator jobs

39,876 open jobs

  • Market Maker jobs

1,432 open jobs

  • Yield Engineer jobs

9,445 open jobs

  • Computer Scientist jobs

49,477 open jobs

  • Lead Test Engineer jobs

13,921 open jobs

  • House Supervisor jobs

29,485 open jobs

  • Core Engineer jobs

33,936 open jobs

  • Cable Technician jobs

11,124 open jobs

  • Principal Software Engineer jobs

73,845 open jobs

  • Logic Design Engineer jobs

1,858 open jobs

Explore top content on LinkedIn

Find curated posts and insights for relevant topics all in one place.

View top content