Member of Technical Staff - Universal Inference
IT
San Francisco, CA, USA
USD 200k-400k / year
Posted on Aug 8, 2026
Infinity Artificial Intelligence Institute San Francisco Bay Area
Member of Technical Staff - Universal Inference
Infinity Artificial Intelligence Institute San Francisco Bay Area
1 day ago 33 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
Company: Infinity
Every accelerator that comes online needs an inference stack, and today that stack is hand-built per chip, per model, per optimization - a permanent, growing backlog of engineering work that scales linearly with the number of chips and models in the world.
We're building Infy: a global inference library, generated and continuously maintained by AI, that targets every major chip - NVIDIA, AMD, Trainium, TPU, Maia, Cerebras, Groq, Tenstorrent, and dozens more - across every model category, from LLMs to vision, audio, and robotics. Think vLLM, but for every accelerator on the market, kept automatically up to date as chips, models, and optimization techniques change.
What You'll Work On
Depending on your strengths, you'll own one or more layers of the system:
We care more about depth and range than a specific checklist, but strong candidates will have most of:
You want to build the thing that makes every chip's inference performance a solved problem instead of a standing engineering project. You think in terms of systems that scale across hardware and models rather than one-off implementations, you're energized by AI agents doing the generation and optimization work under a tight benchmark harness, and you want to help set the global standard for how inference libraries get built and maintained.
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.
See who you know
Get notified about new Member of Technical Staff jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
Inferact
San Francisco, CA $200,000 - $400,000 1 month ago
Morph
San Francisco, CA $175,000 - $350,000 6 days ago
SpaceXAI
Palo Alto, CA 1 week ago
Reactor
San Francisco, CA 3 months ago
Perplexity
Palo Alto, CA $220,000 - $485,000 1 week ago
Inception
San Francisco Bay Area $200,000 - $350,000 2 weeks ago
Black Forest Labs
San Francisco, CA 1 month ago
Inferact
San Francisco, CA
$200,000.00
$400,000.00
1 month ago
Postman
San Francisco, CA 4 days ago
Touring Capital
San Francisco Bay Area 11 hours ago
Prime Intellect
San Francisco, CA 2 months ago
RadixArk
Palo Alto, CA 1 month ago
Black Forest Labs
San Francisco, CA 2 weeks ago
Inception
San Francisco Bay Area
$200,000.00
$350,000.00
1 week ago
Cohere
San Francisco, CA 2 weeks ago
Magic
San Francisco, CA
$225,000.00
$550,000.00
1 week ago
Intel
Santa Clara, CA 1 day ago
Infinity Artificial Intelligence Institute
San Francisco Bay Area 22 hours ago
LinkedIn
Mountain View, CA 1 week ago
RadixArk
Palo Alto, CA 2 weeks ago
General Motors
Sunnyvale, CA 3 days ago
General Motors
San Francisco, CA 4 days ago
Gimlet Labs
San Francisco, CA
$150,000.00
$350,000.00
3 months ago
Reflection
San Francisco, CA 6 days ago
Cohere
San Francisco, CA 1 week ago
Eventual
San Francisco, CA 15 hours ago
OpenAI
San Francisco, CA
$293,000.00
$445,000.00
3 hours ago
People also viewed
Gimlet Labs
San Francisco, CA $150,000 - $350,000 3 months ago
NVIDIA
Santa Clara, CA 2 weeks ago
Odyssey
Palo Alto, CA 4 months ago
Vals AI
San Francisco, CA 1 month ago
OpenAI
San Francisco, CA $295,000 - $555,000 2 weeks ago
Figure
San Jose, CA 1 month ago
DigitalOcean
San Francisco, CA 2 weeks ago
ComfyUI
San Francisco, CA 7 months ago
Simplify
San Francisco, CA 3 days ago
Liquid AI
San Francisco, CA 1 week ago
Similar Searches
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content
Member of Technical Staff - Universal Inference
Infinity Artificial Intelligence Institute San Francisco Bay Area
1 day ago 33 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
- Report this job
Company: Infinity
- Team: Systems / AI Infrastructure
- Location: San Francisco (on-site)
- Type: Full-time
Every accelerator that comes online needs an inference stack, and today that stack is hand-built per chip, per model, per optimization - a permanent, growing backlog of engineering work that scales linearly with the number of chips and models in the world.
We're building Infy: a global inference library, generated and continuously maintained by AI, that targets every major chip - NVIDIA, AMD, Trainium, TPU, Maia, Cerebras, Groq, Tenstorrent, and dozens more - across every model category, from LLMs to vision, audio, and robotics. Think vLLM, but for every accelerator on the market, kept automatically up to date as chips, models, and optimization techniques change.
What You'll Work On
Depending on your strengths, you'll own one or more layers of the system:
- Optimization agents - the strategy → analyze → generate → curate loop that turns a chip + model pair into an optimized kernel or serving path, and the orchestrator that runs this across many chips and models in parallel.
- Kernel optimization loop - the hierarchical agent system that iterates on individual kernels (matmul, attention variants, normalization, RoPE, MoE, collectives) against reference implementations and measured-peak performance gates, drawing on and growing a registry of thousands of known kernels.
- Library generator - the system that takes a validated set of optimized components and assembles them into a real, installable inference library per chip, with a standards-compatible serving interface.
- Hardware probe - agents that build a structured representation of a chip's memory hierarchy, compute regions, and inter-component communication characteristics, so optimization strategy can branch correctly on the chip's actual architecture.
- Coverage tracking - the enablement and optimization tables that track, per chip and per model category, what's implemented and how fast it runs, and that drive prioritization of what to build next.
- SOTA paper replication - agents that read inference-optimization papers and automatically implement and validate the techniques they describe, so the library's optimizations keep pace with published research rather than lagging behind it.
- Continuous tracking - the pipeline that watches for new papers, new hardware, and new models, and feeds that into what Infy builds and rebuilds next.
- Benchmark harness - the metrics layer (TTFT, TPOT, ITL, E2EL, goodput, and model-category-specific equivalents like image-generation latency) that every optimization is measured against.
- Inference service - the gateway, router, and billing layers that turn the library into a running, AI-optimized inference service people can actually call.
We care more about depth and range than a specific checklist, but strong candidates will have most of:
- Real experience with ML inference internals - vLLM or similar serving stacks, attention kernel variants, KV cache management, quantization, continuous batching.
- Comfort working across heterogeneous hardware - GPU kernels (CUDA, ROCm/HIP, Triton), and ideally exposure to non-GPU execution models (systolic arrays, dataflow, wafer-scale, in-memory/analog compute).
- Hands-on experience building with coding agents / LLMs - prompting, tool-use loops, evaluating and constraining model output, and designing systems where the model writes and optimizes code while tests and benchmarks catch regressions.
- Fluency in Python and at least one systems language (Rust, C, or C++).
- Comfort reading and implementing ideas directly from research papers, not just from existing open-source code.
- Contributed to vLLM, TVM, LLVM, MLIR, or a hardware vendor's compiler/runtime stack.
- Experience with distributed inference or training internals (NCCL, Megatron-LM, DeepSpeed) and collective-communication algorithms.
- Familiarity with model categories beyond LLMs - vision, audio, multimodal, or robotics inference.
- Experience building or maintaining a benchmark suite or performance regression system at scale.
- Background reproducing results from ML systems papers (KernelBench-style benchmarks, evolutionary/search-based kernel optimization).
You want to build the thing that makes every chip's inference performance a solved problem instead of a standing engineering project. You think in terms of systems that scale across hardware and models rather than one-off implementations, you're energized by AI agents doing the generation and optimization work under a tight benchmark harness, and you want to help set the global standard for how inference libraries get built and maintained.
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.
- Seniority level Mid-Senior level
- Employment type Full-time
- Job function Engineering and Information Technology
- Industries Software Development
See who you know
Get notified about new Member of Technical Staff jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
- Member of Technical Staff, Inference
Inferact
San Francisco, CA $200,000 - $400,000 1 month ago
- Member of Technical Staff
Morph
San Francisco, CA $175,000 - $350,000 6 days ago
- Member of Technical Staff - RL Inference
SpaceXAI
Palo Alto, CA 1 week ago
- Member of Technical Staff, Inference
Reactor
San Francisco, CA 3 months ago
- Member of Technical Staff (AI Inference Engineer)
Perplexity
Palo Alto, CA $220,000 - $485,000 1 week ago
- Member of Technical Staff, Kernels
Inception
San Francisco Bay Area $200,000 - $350,000 2 weeks ago
- Member of Technical Staff - Research Engineer
Black Forest Labs
San Francisco, CA 1 month ago
- Member of Technical Staff, TPU Performance Engineering
Inferact
San Francisco, CA
$200,000.00
$400,000.00
1 month ago
- Member of Technical Staff, AI Platform & Architecture (Infrastructure)
Postman
San Francisco, CA 4 days ago
- Member of Technical Staff - AI-Optimized Inference
Touring Capital
San Francisco Bay Area 11 hours ago
- Member of Technical Staff - Inference
Prime Intellect
San Francisco, CA 2 months ago
- Member of Technical Staff — Inference Palo Alto, CA
RadixArk
Palo Alto, CA 1 month ago
- Member of Technical Staff - Model Serving / API Backend Engineer
Black Forest Labs
San Francisco, CA 2 weeks ago
- Member of Technical Staff, Training Infra
Inception
San Francisco Bay Area
$200,000.00
$350,000.00
1 week ago
- Member of Technical Staff, Model Efficiency
Cohere
San Francisco, CA 2 weeks ago
- Member of Technical Staff, Inference & RL Systems
Magic
San Francisco, CA
$225,000.00
$550,000.00
1 week ago
- AI Infrastructure Engineer
Intel
Santa Clara, CA 1 day ago
- Member of Technical Staff - AI-Optimized Inference
Infinity Artificial Intelligence Institute
San Francisco Bay Area 22 hours ago
- Staff Software Engineer - AI Platform
Mountain View, CA 1 week ago
- Member of Technical Staff — Developer Technology Palo Alto, CA
RadixArk
Palo Alto, CA 2 weeks ago
- Staff ML Compiler Engineer
General Motors
Sunnyvale, CA 3 days ago
- Staff ML Compiler Engineer
General Motors
San Francisco, CA 4 days ago
- Member of Technical Staff - ML Systems & Inference
Gimlet Labs
San Francisco, CA
$150,000.00
$350,000.00
3 months ago
- Member of Technical Staff - Mid-Training Infra
Reflection
San Francisco, CA 6 days ago
- Member of Technical Staff, Training Infra Engineer
Cohere
San Francisco, CA 1 week ago
- Member of Technical Staff, Specialized Focus
Eventual
San Francisco, CA 15 hours ago
- Systems Generalist, GPT Infrastructure
OpenAI
San Francisco, CA
$293,000.00
$445,000.00
3 hours ago
People also viewed
- Member of Technical Staff - Applied AI Research
Gimlet Labs
San Francisco, CA $150,000 - $350,000 3 months ago
- Senior Software Engineer, AI Inference Systems
NVIDIA
Santa Clara, CA 2 weeks ago
- Member of Technical Staff, ML Performance
Odyssey
Palo Alto, CA 4 months ago
- Member of Technical Staff - Platform
Vals AI
San Francisco, CA 1 month ago
- Software Engineer, Model Inference
OpenAI
San Francisco, CA $295,000 - $555,000 2 weeks ago
- Staff AI Inference and Acceleration Engineer
Figure
San Jose, CA 1 month ago
- Staff Engineer, Inference Optimizations
DigitalOcean
San Francisco, CA 2 weeks ago
- Member of Technical Staff, Backend
ComfyUI
San Francisco, CA 7 months ago
- Member of Technical Staff
Simplify
San Francisco, CA 3 days ago
- Member of Technical Staff - Edge Inference Engineer
Liquid AI
San Francisco, CA 1 week ago
Similar Searches
- Senior Member of Technical Staff jobs
- Member Technical jobs
- Senior Wireless Engineer jobs
- Physical Design Engineer jobs
- Staff Test Engineer jobs
- Principal Firmware Engineer jobs
- Line Technician jobs
- Lead Infrastructure Engineer jobs
- Support Team Manager jobs
- Vice President Software jobs
- Senior Lead Software Engineer jobs
- Principal Researcher jobs
- Switch Engineer jobs
- Staff Software Engineer jobs
- Lead Quality Engineer jobs
- Control Coordinator jobs
- Market Maker jobs
- Yield Engineer jobs
- Computer Scientist jobs
- Lead Test Engineer jobs
- House Supervisor jobs
- Core Engineer jobs
- Cable Technician jobs
- Principal Software Engineer jobs
- Logic Design Engineer jobs
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content