Member of Technical Staff - Debugger Generation
IT
San Francisco, CA, USA
Posted on Aug 8, 2026
Infinity Artificial Intelligence Institute San Francisco Bay Area
Member of Technical Staff - Debugger Generation
Infinity Artificial Intelligence Institute San Francisco Bay Area
1 day ago Be among the first 25 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
Company: Infinity
Debugging an accelerator usually means chasing something that is already gone. A run hangs, a race condition fires, or a result refuses to reproduce, and the state that would explain it evaporated the moment execution moved on. The standard recourse is to instrument more, run again, and narrow in one increasingly detailed dump at a time; on non-deterministic failures, that loop may never converge. Numerical bugs are worse to localize by hand, because the corruption usually sits far upstream of where it finally surfaces.
A real debugger dissolves that problem by refusing to let the state disappear. It freezes the entire program, including where every value physically lives, and lets you move through execution in both directions. This role builds the agent that generates a debugger like that for a new chip, and its core is continuous checkpointing with restore-to-checkpoint: snapshot the run as it proceeds, then jump back to the instant before the hang, or before the numbers began to degrade, and inspect exactly what was and was not in memory, without rewriting code to reproduce the moment.
The same frozen-state representation turns out to be the right substrate for making code fast, not only for finding out why it broke. Restore to any point, change one part of an inference pass, measure the effect from there, and undo it by restoring, with no full rebuild and no full rerun. That collapse of the rewrite-and-rerun cycle into checkpoint-and-restore is where a 10 to 20x speedup in experimentation comes from. Because numerical bugs are almost always deterministic given their inputs, the same loop closes cleanly: restore, scan for NaN and inf, diff intermediate tensors against a known-good pass, ablate a precision change, restore, re-run. And because implementation timelines are increasingly set by how fast agents can attempt and discard hypotheses, the largest lever may be that this is a far better interface for an agent than writing throwaway scripts, a place to read every value and its location and invalidate its own wrong assumptions in seconds rather than runs.
The specification is set by the best debuggers the industry already has, the whole thing has to run on the chip itself with no simulator standing in, and it has to generalize across hardware, AI accelerators first but mobile SoCs like Snapdragon too. The near-term target is concrete, and unglamorous in the right way: today’s d-Matrix Corsair debugger permits exactly one breakpoint and no checkpointing at all. We’re co-creating its successor, with real breakpoints, continuous checkpointing, and restore to intermediate state, so that AMPs can run inside the debugger with Corsair state rewound to any point in the computation. That is the same chip Ignition, our bringup agent, took from first hardware access to tensor-parallel matmuls across all 32 compute units in 10 hours and to three frontier models end-to-end in 10 days, now live as the Infinity d-Matrix Cloud.
What you’ll work on
You’ll build the debugger and the agent that generates it. Depending on your strengths:
We weight range and depth over any particular résumé. Strong candidates will have most of the following:
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We’ve signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We’re headquartered in San Francisco.
See who you know
Get notified about new Member of Technical Staff jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
Touring Capital
San Francisco Bay Area 11 hours ago
Qualcomm
Santa Clara, CA 3 days ago
Waymo
Mountain View, CA 1 day ago
General Motors
Sunnyvale, CA 3 days ago
General Motors
San Francisco, CA 4 days ago
Google
San Jose, CA 1 week ago
Samsung Semiconductor
San Jose, CA 1 day ago
Gimlet Labs
San Francisco, CA
$180,000.00
$400,000.00
3 months ago
Intel
Santa Clara, CA 1 week ago
Google
San Jose, CA 3 days ago
NVIDIA
Santa Clara, CA 16 hours ago
Sony Interactive Entertainment
San Mateo, CA 1 week ago
Theorem
San Francisco, CA
$300,000.00
$500,000.00
9 months ago
Lemurian Labs
Santa Clara, CA 1 week ago
OpenAI
San Francisco, CA
$225,000.00
$445,000.00
2 weeks ago
Morph
San Francisco, CA
$175,000.00
$350,000.00
6 days ago
Intel
Santa Clara, CA 1 week ago
Perplexity
San Francisco, CA
$250,000.00
$405,000.00
2 weeks ago
AMD
San Jose, CA
$204,000.00
$306,000.00
1 day ago
hillclimb
San Francisco, CA
$200,000.00
$300,000.00
1 month ago
Anthropic
San Francisco, CA 2 weeks ago
dynamism
San Francisco, CA
$400,000.00
$500,000.00
4 days ago
Samsung Semiconductor
San Jose, CA 4 days ago
Google DeepMind
Mountain View, CA 1 day ago
Theorem
San Francisco, CA
$150,000.00
$250,000.00
9 months ago
OpenAI
San Francisco, CA
$266,000.00
$445,000.00
1 week ago
Perplexity
San Francisco, CA
$220,000.00
$405,000.00
1 week ago
People also viewed
dynamism
San Francisco, CA $400,000 - $500,000 4 days ago
Waymo
Mountain View, CA 2 weeks ago
NVIDIA
Santa Clara, CA 5 days ago
ByteDance
San Jose, CA 1 day ago
SpaceXAI
Palo Alto, CA 2 weeks ago
Vela
San Francisco, CA $120,000 - $300,000 2 weeks ago
Anthropic
San Francisco, CA 2 weeks ago
SpaceXAI
Palo Alto, CA 5 hours ago
Magic
San Francisco, CA $225,000 - $550,000 1 week ago
AMD
Santa Clara, CA $204,000 - $306,000 4 days ago
Similar Searches
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content
Member of Technical Staff - Debugger Generation
Infinity Artificial Intelligence Institute San Francisco Bay Area
1 day ago Be among the first 25 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
- Report this job
Company: Infinity
- Team: Systems / AI Infrastructure
- Location: San Francisco (on-site)
- Type: Full-time
Debugging an accelerator usually means chasing something that is already gone. A run hangs, a race condition fires, or a result refuses to reproduce, and the state that would explain it evaporated the moment execution moved on. The standard recourse is to instrument more, run again, and narrow in one increasingly detailed dump at a time; on non-deterministic failures, that loop may never converge. Numerical bugs are worse to localize by hand, because the corruption usually sits far upstream of where it finally surfaces.
A real debugger dissolves that problem by refusing to let the state disappear. It freezes the entire program, including where every value physically lives, and lets you move through execution in both directions. This role builds the agent that generates a debugger like that for a new chip, and its core is continuous checkpointing with restore-to-checkpoint: snapshot the run as it proceeds, then jump back to the instant before the hang, or before the numbers began to degrade, and inspect exactly what was and was not in memory, without rewriting code to reproduce the moment.
The same frozen-state representation turns out to be the right substrate for making code fast, not only for finding out why it broke. Restore to any point, change one part of an inference pass, measure the effect from there, and undo it by restoring, with no full rebuild and no full rerun. That collapse of the rewrite-and-rerun cycle into checkpoint-and-restore is where a 10 to 20x speedup in experimentation comes from. Because numerical bugs are almost always deterministic given their inputs, the same loop closes cleanly: restore, scan for NaN and inf, diff intermediate tensors against a known-good pass, ablate a precision change, restore, re-run. And because implementation timelines are increasingly set by how fast agents can attempt and discard hypotheses, the largest lever may be that this is a far better interface for an agent than writing throwaway scripts, a place to read every value and its location and invalidate its own wrong assumptions in seconds rather than runs.
The specification is set by the best debuggers the industry already has, the whole thing has to run on the chip itself with no simulator standing in, and it has to generalize across hardware, AI accelerators first but mobile SoCs like Snapdragon too. The near-term target is concrete, and unglamorous in the right way: today’s d-Matrix Corsair debugger permits exactly one breakpoint and no checkpointing at all. We’re co-creating its successor, with real breakpoints, continuous checkpointing, and restore to intermediate state, so that AMPs can run inside the debugger with Corsair state rewound to any point in the computation. That is the same chip Ignition, our bringup agent, took from first hardware access to tensor-parallel matmuls across all 32 compute units in 10 hours and to three frontier models end-to-end in 10 days, now live as the Infinity d-Matrix Cloud.
What you’ll work on
You’ll build the debugger and the agent that generates it. Depending on your strengths:
- Continuous checkpointing and restore. Capture full program state at varying granularity and roll back to any saved point, and build the reverse-execution machinery beneath it: invert computations where they are invertible, and replay forward from the nearest checkpoint where they are not. Time reversal, made practical.
- Breakpoints and state inspection. Halt a computation and surface every variable, its value, and its physical location, with metadata attached at each memory write that ties the value back to its node in the inference computation graph, so “what is this, and where did it come from” is answerable rather than archaeological.
- In-debugger optimization. Make the state representation good enough that a change to an inference pass can be applied, measured, and ablated in place, so experimentation runs at checkpoint speed instead of rebuild speed.
- Parallel debugging. Capture and faithfully reproduce how parallel execution units interact, gangs for instance, because that interaction is precisely where most race conditions live, and where single-threaded debuggers go blind.
- Differential debugging. Diff a suspect inference pass against a known-good reference to localize where it diverged, numerically or structurally, instead of bisecting by hand.
- Breadth across silicon. Bring the same capabilities to very different hardware, from AI accelerators to mobile parts like Snapdragon and its Hexagon DSP, on real silicon with nothing simulated in the loop.
- The Corsair build. The first concrete target: adding real breakpoints and continuous checkpointing to the existing d-Matrix debugger, and running AMPs inside it with state restored to intermediate points in the computation.
- The agent-facing interface. Expose all of it so an agent can pause, read every value and its location, and disprove its own hypotheses through the debugger rather than by writing and running throwaway scripts, which more than anything else sets how fast a hard bug gets solved.
We weight range and depth over any particular résumé. Strong candidates will have most of the following:
- Genuine debugging depth. You’ve built a debugger, or depended on one hard enough to know exactly where they fall short, ideally having tracked a bug all the way down to the silicon.
- A real model of program state on hardware: memory models, execution ordering, and what it actually costs to capture and restore all of it on a massively parallel machine.
- Experience with time-travel, record-replay, or checkpoint-restore systems, or a strong pull toward building one from first principles.
- An instinct for tools whose primary user is an agent, not a human at a keyboard, and for how that changes the interface.
- Fluency in Python and a systems language: Rust, C, or C++.
- Built or worked on a time-travel or record-replay debugger: rr, gdb reverse debugging, WinDbg Time Travel, or your own.
- Firmware, bare-metal, on-chip debug, or JTAG experience.
- Worked on AI accelerators or mobile SoCs such as Snapdragon and its Hexagon DSP.
- Hunted numerical bugs in ML, the NaN-at-layer-forty kind that only manifests sometimes.
- Familiarity with checkpoint-and-restore systems such as CRIU.
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We’ve signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We’re headquartered in San Francisco.
- Seniority level Mid-Senior level
- Employment type Full-time
- Job function Engineering and Information Technology
- Industries Software Development
See who you know
Get notified about new Member of Technical Staff jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
- Member of Technical Staff - Debugger Generation
Touring Capital
San Francisco Bay Area 11 hours ago
- Staff Software Engineer – Platform Debug
Qualcomm
Santa Clara, CA 3 days ago
- Staff Software Engineer, Compute Reliability
Waymo
Mountain View, CA 1 day ago
- Staff ML Compiler Engineer
General Motors
Sunnyvale, CA 3 days ago
- Staff ML Compiler Engineer
General Motors
San Francisco, CA 4 days ago
- Software Engineer, LLVM and Production Toolchain
San Jose, CA 1 week ago
- Staff Engineer, Compiler
Samsung Semiconductor
San Jose, CA 1 day ago
- Member of Technical Staff - Compilers
Gimlet Labs
San Francisco, CA
$180,000.00
$400,000.00
3 months ago
- Compiler Engineer
Intel
Santa Clara, CA 1 week ago
- Staff Software Engineer, Source Client
San Jose, CA 3 days ago
- GPU Performance Profiling Engineer
NVIDIA
Santa Clara, CA 16 hours ago
- Senior Software Engineer
Sony Interactive Entertainment
San Mateo, CA 1 week ago
- Senior SWE
Theorem
San Francisco, CA
$300,000.00
$500,000.00
9 months ago
- Senior Developer Tools Engineer
Lemurian Labs
Santa Clara, CA 1 week ago
- Hardware Tools Engineer
OpenAI
San Francisco, CA
$225,000.00
$445,000.00
2 weeks ago
- Member of Technical Staff
Morph
San Francisco, CA
$175,000.00
$350,000.00
6 days ago
- Development Tools Software Engineer
Intel
Santa Clara, CA 1 week ago
- Member of Technical Staff (Software Engineer, Acceleration)
Perplexity
San Francisco, CA
$250,000.00
$405,000.00
2 weeks ago
- Staff Software Development Engineer - Rust, Compilers, and GPU Systems
AMD
San Jose, CA
$204,000.00
$306,000.00
1 day ago
- Member of Technical Staff, Research
hillclimb
San Francisco, CA
$200,000.00
$300,000.00
1 month ago
- Staff+ Software Engineer, Experimentation
Anthropic
San Francisco, CA 2 weeks ago
- Member of Technical Staff, Deployments
dynamism
San Francisco, CA
$400,000.00
$500,000.00
4 days ago
- Staff Engineer, Workbench Platform
Samsung Semiconductor
San Jose, CA 4 days ago
- Software Engineer, GenAI Silicon Automation, DeepMind
Google DeepMind
Mountain View, CA 1 day ago
- Systems Engineer
Theorem
San Francisco, CA
$150,000.00
$250,000.00
9 months ago
- Software Engineer, Kernel Performance & AI Tooling
OpenAI
San Francisco, CA
$266,000.00
$445,000.00
1 week ago
- Member of Technical Staff (Software Engineer, Computer)
Perplexity
San Francisco, CA
$220,000.00
$405,000.00
1 week ago
People also viewed
- Member of Technical Staff
dynamism
San Francisco, CA $400,000 - $500,000 4 days ago
- Staff Software Engineer (TLM), Multiverse
Waymo
Mountain View, CA 2 weeks ago
- Senior AI Frameworks Engineer
NVIDIA
Santa Clara, CA 5 days ago
- Tech Lead Software Engineer, Programming Language
ByteDance
San Jose, CA 1 day ago
- Member of Technical Staff - Coding Agents
SpaceXAI
Palo Alto, CA 2 weeks ago
- Member of Technical Staff @ Vela
Vela
San Francisco, CA $120,000 - $300,000 2 weeks ago
- Staff+ Software Engineer, Developer Acceleration
Anthropic
San Francisco, CA 2 weeks ago
- Member of Technical Staff - Observability
SpaceXAI
Palo Alto, CA 5 hours ago
- Member of Technical Staff, Kernels
Magic
San Francisco, CA $225,000 - $550,000 1 week ago
- Staff Software Development Engineer - AI Profiler Tools
AMD
Santa Clara, CA $204,000 - $306,000 4 days ago
Similar Searches
- Senior Member of Technical Staff jobs
- Member Technical jobs
- Senior Wireless Engineer jobs
- Physical Design Engineer jobs
- Staff Test Engineer jobs
- Principal Firmware Engineer jobs
- Line Technician jobs
- Lead Infrastructure Engineer jobs
- Support Team Manager jobs
- Vice President Software jobs
- Senior Lead Software Engineer jobs
- Principal Researcher jobs
- Switch Engineer jobs
- Staff Software Engineer jobs
- Lead Quality Engineer jobs
- Control Coordinator jobs
- Market Maker jobs
- Yield Engineer jobs
- Computer Scientist jobs
- Lead Test Engineer jobs
- House Supervisor jobs
- Core Engineer jobs
- Cable Technician jobs
- Principal Software Engineer jobs
- Logic Design Engineer jobs
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content