Debugger Generation Agent Engineer - Member of Technical Staff
IT
San Francisco, CA, USA
USD 180k-400k / year
Posted on Jul 23, 2026
Infinity Artificial Intelligence Institute San Francisco Bay Area
Debugger Generation Agent Engineer - Member of Technical Staff
Infinity Artificial Intelligence Institute San Francisco Bay Area
6 days ago 38 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
Company: Infinity
The same frozen-state representation turns out to be the right substrate for making code fast, not only for finding out why it broke. Restore to any point, change one part of an inference pass, measure the effect from there, and undo it by restoring, with no full rebuild and no full rerun. That collapse of the rewrite-and-rerun cycle into checkpoint-and-restore is where a 10 to 20x speedup in experimentation comes from. Because numerical bugs are almost always deterministic given their inputs, the same loop closes cleanly: restore, scan for NaN and inf, diff intermediate tensors against a known-good pass, ablate a precision change, restore, re-run. And because implementation timelines are increasingly set by how fast agents can attempt and discard hypotheses, the largest lever may be that this is a far better interface for an agent than writing throwaway scripts, a place to read every value and its location and invalidate its own wrong assumptions in seconds rather than runs.
The specification is set by the best debuggers the industry already has, the whole thing has to run on the chip itself with no simulator standing in, and it has to generalize across hardware, AI accelerators first but mobile SoCs like Snapdragon too. The near-term target is concrete, and unglamorous in the right way: today’s d-Matrix Corsair debugger permits exactly one breakpoint and no checkpointing at all. We’re co-creating its successor, with real breakpoints, continuous checkpointing, and restore to intermediate state, so that AMPs can run inside the debugger with Corsair state rewound to any point in the computation. That is the same chip Ignition, our bringup agent, took from first hardware access to tensor-parallel matmuls across all 32 compute units in 10 hours and to three frontier models end-to-end in 10 days, now live as the Infinity d-Matrix Cloud.
What you’ll work on
You’ll build the debugger and the agent that generates it. Depending on your strengths:
We weight range and depth over any particular résumé. Strong candidates will have most of the following:
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We’ve signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We’re headquartered in San Francisco.
See who you know
Get notified about new Technical Representative jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
Qualcomm
Santa Clara, CA 1 week ago
Waymo
Mountain View, CA 1 week ago
OpenAI
San Francisco, CA $266,000 - $445,000 2 weeks ago
Touring Capital
San Francisco Bay Area 1 day ago
OpenAI
San Francisco, CA $225,000 - $445,000 2 weeks ago
Lemurian Labs
Santa Clara, CA 2 weeks ago
Gimlet Labs
San Francisco, CA
$180,000.00
$400,000.00
2 months ago
AMD
Santa Clara, CA
$204,000.00
$306,000.00
1 week ago
NVIDIA
Santa Clara, CA 1 hour ago
AMD
San Jose, CA
$204,000.00
$306,000.00
1 week ago
Sony Interactive Entertainment
San Mateo, CA 2 weeks ago
General Motors
Sunnyvale, CA 1 week ago
General Motors
San Francisco, CA 1 week ago
Waymo
San Francisco, CA 2 months ago
Roblox
San Mateo, CA 5 days ago
NVIDIA
Santa Clara, CA 2 weeks ago
Morph Systems
Mountain View, CA 1 week ago
Google
Mountain View, CA 2 weeks ago
Harper
San Francisco, CA
$250,000.00
$350,000.00
1 month ago
Meta
Sunnyvale, CA
$154,003.00
$217,000.00
1 week ago
Microsoft
Mountain View, CA 6 days ago
Samsung Semiconductor
San Jose, CA 1 week ago
Snowflake
Menlo Park, CA 1 day ago
Perplexity
San Francisco, CA
$250,000.00
$405,000.00
1 day ago
Nuro
Mountain View, CA 3 months ago
SiFive
Santa Clara, CA 3 months ago
MatX
Mountain View, CA 2 weeks ago
People also viewed
Efficient Computer
San Jose, CA 2 weeks ago
SiFive
Berkeley, CA 3 months ago
OpenAI
San Francisco, CA $225,000 - $455,000 2 weeks ago
OpenAI
San Francisco, CA $293,000 - $445,000 5 days ago
Waymo
Mountain View, CA 2 months ago
NVIDIA
Santa Clara, CA 4 days ago
NVIDIA
Santa Clara, CA 2 weeks ago
Waymo
San Francisco, CA 2 months ago
NVIDIA
Santa Clara, CA 1 week ago
Waymo
Mountain View, CA 2 months ago
Similar Searches
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content
Debugger Generation Agent Engineer - Member of Technical Staff
Infinity Artificial Intelligence Institute San Francisco Bay Area
6 days ago 38 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
- Report this job
Company: Infinity
- Team: Systems / AI Infrastructure
- Location: San Francisco (on-site)
- Type: Full-time The Mission Debugging an accelerator usually means chasing something that is already gone. A run hangs, a race condition fires, or a result refuses to reproduce, and the state that would explain it evaporated the moment execution moved on. The standard recourse is to instrument more, run again, and narrow in one increasingly detailed dump at a time; on non-deterministic failures, that loop may never converge. Numerical bugs are worse to localize by hand, because the corruption usually sits far upstream of where it finally surfaces.
The same frozen-state representation turns out to be the right substrate for making code fast, not only for finding out why it broke. Restore to any point, change one part of an inference pass, measure the effect from there, and undo it by restoring, with no full rebuild and no full rerun. That collapse of the rewrite-and-rerun cycle into checkpoint-and-restore is where a 10 to 20x speedup in experimentation comes from. Because numerical bugs are almost always deterministic given their inputs, the same loop closes cleanly: restore, scan for NaN and inf, diff intermediate tensors against a known-good pass, ablate a precision change, restore, re-run. And because implementation timelines are increasingly set by how fast agents can attempt and discard hypotheses, the largest lever may be that this is a far better interface for an agent than writing throwaway scripts, a place to read every value and its location and invalidate its own wrong assumptions in seconds rather than runs.
The specification is set by the best debuggers the industry already has, the whole thing has to run on the chip itself with no simulator standing in, and it has to generalize across hardware, AI accelerators first but mobile SoCs like Snapdragon too. The near-term target is concrete, and unglamorous in the right way: today’s d-Matrix Corsair debugger permits exactly one breakpoint and no checkpointing at all. We’re co-creating its successor, with real breakpoints, continuous checkpointing, and restore to intermediate state, so that AMPs can run inside the debugger with Corsair state rewound to any point in the computation. That is the same chip Ignition, our bringup agent, took from first hardware access to tensor-parallel matmuls across all 32 compute units in 10 hours and to three frontier models end-to-end in 10 days, now live as the Infinity d-Matrix Cloud.
What you’ll work on
You’ll build the debugger and the agent that generates it. Depending on your strengths:
- Continuous checkpointing and restore. Capture full program state at varying granularity and roll back to any saved point, and build the reverse-execution machinery beneath it: invert computations where they are invertible, and replay forward from the nearest checkpoint where they are not. Time reversal, made practical.
- Breakpoints and state inspection. Halt a computation and surface every variable, its value, and its physical location, with metadata attached at each memory write that ties the value back to its node in the inference computation graph, so “what is this, and where did it come from” is answerable rather than archaeological.
- In-debugger optimization. Make the state representation good enough that a change to an inference pass can be applied, measured, and ablated in place, so experimentation runs at checkpoint speed instead of rebuild speed.
- Parallel debugging. Capture and faithfully reproduce how parallel execution units interact, gangs for instance, because that interaction is precisely where most race conditions live, and where single-threaded debuggers go blind.
- Differential debugging. Diff a suspect inference pass against a known-good reference to localize where it diverged, numerically or structurally, instead of bisecting by hand.
- Breadth across silicon. Bring the same capabilities to very different hardware, from AI accelerators to mobile parts like Snapdragon and its Hexagon DSP, on real silicon with nothing simulated in the loop.
- The Corsair build. The first concrete target: adding real breakpoints and continuous checkpointing to the existing d-Matrix debugger, and running AMPs inside it with state restored to intermediate points in the computation.
- The agent-facing interface. Expose all of it so an agent can pause, read every value and its location, and disprove its own hypotheses through the debugger rather than by writing and running throwaway scripts, which more than anything else sets how fast a hard bug gets solved.
We weight range and depth over any particular résumé. Strong candidates will have most of the following:
- Genuine debugging depth. You’ve built a debugger, or depended on one hard enough to know exactly where they fall short, ideally having tracked a bug all the way down to the silicon.
- A real model of program state on hardware: memory models, execution ordering, and what it actually costs to capture and restore all of it on a massively parallel machine.
- Experience with time-travel, record-replay, or checkpoint-restore systems, or a strong pull toward building one from first principles.
- An instinct for tools whose primary user is an agent, not a human at a keyboard, and for how that changes the interface.
- Fluency in Python and a systems language: Rust, C, or C++.
- Built or worked on a time-travel or record-replay debugger: rr, gdb reverse debugging, WinDbg Time Travel, or your own.
- Firmware, bare-metal, on-chip debug, or JTAG experience.
- Worked on AI accelerators or mobile SoCs such as Snapdragon and its Hexagon DSP.
- Hunted numerical bugs in ML, the NaN-at-layer-forty kind that only manifests sometimes.
- Familiarity with checkpoint-and-restore systems such as CRIU.
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We’ve signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We’re headquartered in San Francisco.
- Seniority level Entry level
- Employment type Full-time
- Job function Sales and Business Development
- Industries Software Development
See who you know
Get notified about new Technical Representative jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
- Staff Software Engineer – Platform Debug
Qualcomm
Santa Clara, CA 1 week ago
- Staff Software Engineer, Compute Reliability
Waymo
Mountain View, CA 1 week ago
- Software Engineer, Kernel Performance & AI Tooling
OpenAI
San Francisco, CA $266,000 - $445,000 2 weeks ago
- Debugger Generation Agent Engineer - Member of Technical Staff
Touring Capital
San Francisco Bay Area 1 day ago
- Hardware Tools Engineer
OpenAI
San Francisco, CA $225,000 - $445,000 2 weeks ago
- Senior Developer Tools Engineer
Lemurian Labs
Santa Clara, CA 2 weeks ago
- Member of Technical Staff - Compilers
Gimlet Labs
San Francisco, CA
$180,000.00
$400,000.00
2 months ago
- Staff Software Development Engineer - AI Profiler Tools
AMD
Santa Clara, CA
$204,000.00
$306,000.00
1 week ago
- Software Engineer, GPU Performance Tools
NVIDIA
Santa Clara, CA 1 hour ago
- Senior Staff Software Development Engineer
AMD
San Jose, CA
$204,000.00
$306,000.00
1 week ago
- Senior Software Engineer
Sony Interactive Entertainment
San Mateo, CA 2 weeks ago
- Staff ML Compiler Engineer
General Motors
Sunnyvale, CA 1 week ago
- Staff ML Compiler Engineer
General Motors
San Francisco, CA 1 week ago
- Staff Software Engineer, Multiverse
Waymo
San Francisco, CA 2 months ago
- Principal Software Engineer, Crash Reporting
Roblox
San Mateo, CA 5 days ago
- Senior Failure Analysis Engineer
NVIDIA
Santa Clara, CA 2 weeks ago
- Software Engineer
Morph Systems
Mountain View, CA 1 week ago
- Staff Software Engineer, Machine Learning Compiler, Google Research
Mountain View, CA 2 weeks ago
- Staff Engineer, Engineering Productivity & AI Quality
Harper
San Francisco, CA
$250,000.00
$350,000.00
1 month ago
- Software Engineer, Systems ML - Compilers / Backend
Meta
Sunnyvale, CA
$154,003.00
$217,000.00
1 week ago
- Principal AI Accelerator Tools Development Engineer
Microsoft
Mountain View, CA 6 days ago
- Staff Engineer, Compiler
Samsung Semiconductor
San Jose, CA 1 week ago
- Software Engineer
Snowflake
Menlo Park, CA 1 day ago
- Member of Technical Staff (Software Engineer, Acceleration)
Perplexity
San Francisco, CA
$250,000.00
$405,000.00
1 day ago
- Software Engineer, Performance Tooling and Infrastructure
Nuro
Mountain View, CA 3 months ago
- Software Engineer - Platform Technologies
SiFive
Santa Clara, CA 3 months ago
- Software Engineer - Compiler
MatX
Mountain View, CA 2 weeks ago
People also viewed
- Performance Research Engineer (multiple levels)
Efficient Computer
San Jose, CA 2 weeks ago
- Software Engineer - Platform Technologies
SiFive
Berkeley, CA 3 months ago
- Software Engineer, Hardware
OpenAI
San Francisco, CA $225,000 - $455,000 2 weeks ago
- Systems Generalist, GPT Infrastructure
OpenAI
San Francisco, CA $293,000 - $445,000 5 days ago
- Staff Software Engineer, Multiverse
Waymo
Mountain View, CA 2 months ago
- Senior Compiler Engineer - AI
NVIDIA
Santa Clara, CA 4 days ago
- Senior Software Engineer, AI Speed Infrastructure
NVIDIA
Santa Clara, CA 2 weeks ago
- Software Engineer, Multiverse
Waymo
San Francisco, CA 2 months ago
- Senior AI Frameworks Engineer
NVIDIA
Santa Clara, CA 1 week ago
- Software Engineer, Multiverse
Waymo
Mountain View, CA 2 months ago
Similar Searches
- Help Desk Representative jobs
- Technical Service Representative jobs
- Technical Sales Assistant jobs
- Liaison Engineer jobs
- Senior Microbiologist jobs
- Technical Sales Representative jobs
- Senior Technical Sales Representative jobs
- Market Development Specialist jobs
- Senior Sales Administrator jobs
- Agronomist jobs
- Security Representative jobs
- Field Service Representative jobs
- Extension Educator jobs
- Career Services Specialist jobs
- Technical Supervisor jobs
- Quality Control Microbiologist jobs
- Construction Operations Manager jobs
- Licensed Aircraft Engineer jobs
- Technology Innovation Manager jobs
- Technical Services Specialist jobs
- Quality Expert jobs
- Field Representative jobs
- Reservoir Engineer jobs
- Capacity Planner jobs
- Flight Operations Engineer jobs
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content