Debugger Generation Agent Engineer - Member of Technical Staff

Touring Capital
Touring Capital

IT

San Francisco, CA, USA

USD 180k-400k / year

Posted on Jul 23, 2026
Infinity Artificial Intelligence Institute San Francisco Bay Area

Debugger Generation Agent Engineer - Member of Technical Staff

Infinity Artificial Intelligence Institute San Francisco Bay Area

6 days ago 38 applicants

See who Infinity Artificial Intelligence Institute has hired for this role

Save

  • Report this job

Debugger Generation Agent

Company: Infinity

  • Team: Systems / AI Infrastructure
  • Location: San Francisco (on-site)
  • Type: Full-time The Mission Debugging an accelerator usually means chasing something that is already gone. A run hangs, a race condition fires, or a result refuses to reproduce, and the state that would explain it evaporated the moment execution moved on. The standard recourse is to instrument more, run again, and narrow in one increasingly detailed dump at a time; on non-deterministic failures, that loop may never converge. Numerical bugs are worse to localize by hand, because the corruption usually sits far upstream of where it finally surfaces.

A real debugger dissolves that problem by refusing to let the state disappear. It freezes the entire program, including where every value physically lives, and lets you move through execution in both directions. This role builds the agent that generates a debugger like that for a new chip, and its core is continuous checkpointing with restore-to-checkpoint: snapshot the run as it proceeds, then jump back to the instant before the hang, or before the numbers began to degrade, and inspect exactly what was and was not in memory, without rewriting code to reproduce the moment.

The same frozen-state representation turns out to be the right substrate for making code fast, not only for finding out why it broke. Restore to any point, change one part of an inference pass, measure the effect from there, and undo it by restoring, with no full rebuild and no full rerun. That collapse of the rewrite-and-rerun cycle into checkpoint-and-restore is where a 10 to 20x speedup in experimentation comes from. Because numerical bugs are almost always deterministic given their inputs, the same loop closes cleanly: restore, scan for NaN and inf, diff intermediate tensors against a known-good pass, ablate a precision change, restore, re-run. And because implementation timelines are increasingly set by how fast agents can attempt and discard hypotheses, the largest lever may be that this is a far better interface for an agent than writing throwaway scripts, a place to read every value and its location and invalidate its own wrong assumptions in seconds rather than runs.

The specification is set by the best debuggers the industry already has, the whole thing has to run on the chip itself with no simulator standing in, and it has to generalize across hardware, AI accelerators first but mobile SoCs like Snapdragon too. The near-term target is concrete, and unglamorous in the right way: today’s d-Matrix Corsair debugger permits exactly one breakpoint and no checkpointing at all. We’re co-creating its successor, with real breakpoints, continuous checkpointing, and restore to intermediate state, so that AMPs can run inside the debugger with Corsair state rewound to any point in the computation. That is the same chip Ignition, our bringup agent, took from first hardware access to tensor-parallel matmuls across all 32 compute units in 10 hours and to three frontier models end-to-end in 10 days, now live as the Infinity d-Matrix Cloud.

What you’ll work on

You’ll build the debugger and the agent that generates it. Depending on your strengths:

  • Continuous checkpointing and restore. Capture full program state at varying granularity and roll back to any saved point, and build the reverse-execution machinery beneath it: invert computations where they are invertible, and replay forward from the nearest checkpoint where they are not. Time reversal, made practical.
  • Breakpoints and state inspection. Halt a computation and surface every variable, its value, and its physical location, with metadata attached at each memory write that ties the value back to its node in the inference computation graph, so “what is this, and where did it come from” is answerable rather than archaeological.
  • In-debugger optimization. Make the state representation good enough that a change to an inference pass can be applied, measured, and ablated in place, so experimentation runs at checkpoint speed instead of rebuild speed.
  • Parallel debugging. Capture and faithfully reproduce how parallel execution units interact, gangs for instance, because that interaction is precisely where most race conditions live, and where single-threaded debuggers go blind.
  • Differential debugging. Diff a suspect inference pass against a known-good reference to localize where it diverged, numerically or structurally, instead of bisecting by hand.
  • Breadth across silicon. Bring the same capabilities to very different hardware, from AI accelerators to mobile parts like Snapdragon and its Hexagon DSP, on real silicon with nothing simulated in the loop.
  • The Corsair build. The first concrete target: adding real breakpoints and continuous checkpointing to the existing d-Matrix debugger, and running AMPs inside it with state restored to intermediate points in the computation.
  • The agent-facing interface. Expose all of it so an agent can pause, read every value and its location, and disprove its own hypotheses through the debugger rather than by writing and running throwaway scripts, which more than anything else sets how fast a hard bug gets solved.

What we’re looking for

We weight range and depth over any particular résumé. Strong candidates will have most of the following:

  • Genuine debugging depth. You’ve built a debugger, or depended on one hard enough to know exactly where they fall short, ideally having tracked a bug all the way down to the silicon.
  • A real model of program state on hardware: memory models, execution ordering, and what it actually costs to capture and restore all of it on a massively parallel machine.
  • Experience with time-travel, record-replay, or checkpoint-restore systems, or a strong pull toward building one from first principles.
  • An instinct for tools whose primary user is an agent, not a human at a keyboard, and for how that changes the interface.
  • Fluency in Python and a systems language: Rust, C, or C++.

Nice to have

  • Built or worked on a time-travel or record-replay debugger: rr, gdb reverse debugging, WinDbg Time Travel, or your own.
  • Firmware, bare-metal, on-chip debug, or JTAG experience.
  • Worked on AI accelerators or mobile SoCs such as Snapdragon and its Hexagon DSP.
  • Hunted numerical bugs in ML, the NaN-at-layer-forty kind that only manifests sometimes.
  • Familiarity with checkpoint-and-restore systems such as CRIU.

Who we are

Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We’ve signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We’re headquartered in San Francisco.

  • Seniority level Entry level
  • Employment type Full-time
  • Job function Sales and Business Development
  • Industries Software Development

Referrals increase your chances of interviewing at Infinity Artificial Intelligence Institute by 2x

See who you know

Get notified about new Technical Representative jobs in San Francisco Bay Area.

Sign in to create job alert

Similar jobs

  • Staff Software Engineer – Platform Debug

Staff Software Engineer – Platform Debug

Qualcomm

Santa Clara, CA 1 week ago

  • Staff Software Engineer, Compute Reliability

Staff Software Engineer, Compute Reliability

Waymo

Mountain View, CA 1 week ago

  • Software Engineer, Kernel Performance & AI Tooling

Software Engineer, Kernel Performance & AI Tooling

OpenAI

San Francisco, CA $266,000 - $445,000 2 weeks ago

  • Debugger Generation Agent Engineer - Member of Technical Staff

Debugger Generation Agent Engineer - Member of Technical Staff

Touring Capital

San Francisco Bay Area 1 day ago

  • Hardware Tools Engineer

Hardware Tools Engineer

OpenAI

San Francisco, CA $225,000 - $445,000 2 weeks ago

  • Senior Developer Tools Engineer

Senior Developer Tools Engineer

Lemurian Labs

Santa Clara, CA 2 weeks ago

  • Member of Technical Staff - Compilers

Member of Technical Staff - Compilers

Gimlet Labs

San Francisco, CA

$180,000.00

$400,000.00

2 months ago

  • Staff Software Development Engineer - AI Profiler Tools

Staff Software Development Engineer - AI Profiler Tools

AMD

Santa Clara, CA

$204,000.00

$306,000.00

1 week ago

  • Software Engineer, GPU Performance Tools

Software Engineer, GPU Performance Tools

NVIDIA

Santa Clara, CA 1 hour ago

  • Senior Staff Software Development Engineer

Senior Staff Software Development Engineer

AMD

San Jose, CA

$204,000.00

$306,000.00

1 week ago

  • Senior Software Engineer

Senior Software Engineer

Sony Interactive Entertainment

San Mateo, CA 2 weeks ago

  • Staff ML Compiler Engineer

Staff ML Compiler Engineer

General Motors

Sunnyvale, CA 1 week ago

  • Staff ML Compiler Engineer

Staff ML Compiler Engineer

General Motors

San Francisco, CA 1 week ago

  • Staff Software Engineer, Multiverse

Staff Software Engineer, Multiverse

Waymo

San Francisco, CA 2 months ago

  • Principal Software Engineer, Crash Reporting

Principal Software Engineer, Crash Reporting

Roblox

San Mateo, CA 5 days ago

  • Senior Failure Analysis Engineer

Senior Failure Analysis Engineer

NVIDIA

Santa Clara, CA 2 weeks ago

  • Software Engineer

Software Engineer

Morph Systems

Mountain View, CA 1 week ago

  • Staff Software Engineer, Machine Learning Compiler, Google Research

Staff Software Engineer, Machine Learning Compiler, Google Research

Google

Mountain View, CA 2 weeks ago

  • Staff Engineer, Engineering Productivity & AI Quality

Staff Engineer, Engineering Productivity & AI Quality

Harper

San Francisco, CA

$250,000.00

$350,000.00

1 month ago

  • Software Engineer, Systems ML - Compilers / Backend

Software Engineer, Systems ML - Compilers / Backend

Meta

Sunnyvale, CA

$154,003.00

$217,000.00

1 week ago

  • Principal AI Accelerator Tools Development Engineer

Principal AI Accelerator Tools Development Engineer

Microsoft

Mountain View, CA 6 days ago

  • Staff Engineer, Compiler

Staff Engineer, Compiler

Samsung Semiconductor

San Jose, CA 1 week ago

  • Software Engineer

Software Engineer

Snowflake

Menlo Park, CA 1 day ago

  • Member of Technical Staff (Software Engineer, Acceleration)

Member of Technical Staff (Software Engineer, Acceleration)

Perplexity

San Francisco, CA

$250,000.00

$405,000.00

1 day ago

  • Software Engineer, Performance Tooling and Infrastructure

Software Engineer, Performance Tooling and Infrastructure

Nuro

Mountain View, CA 3 months ago

  • Software Engineer - Platform Technologies

Software Engineer - Platform Technologies

SiFive

Santa Clara, CA 3 months ago

  • Software Engineer - Compiler

Software Engineer - Compiler

MatX

Mountain View, CA 2 weeks ago

People also viewed

  • Performance Research Engineer (multiple levels)

Performance Research Engineer (multiple levels)

Efficient Computer

San Jose, CA 2 weeks ago

  • Software Engineer - Platform Technologies

Software Engineer - Platform Technologies

SiFive

Berkeley, CA 3 months ago

  • Software Engineer, Hardware

Software Engineer, Hardware

OpenAI

San Francisco, CA $225,000 - $455,000 2 weeks ago

  • Systems Generalist, GPT Infrastructure

Systems Generalist, GPT Infrastructure

OpenAI

San Francisco, CA $293,000 - $445,000 5 days ago

  • Staff Software Engineer, Multiverse

Staff Software Engineer, Multiverse

Waymo

Mountain View, CA 2 months ago

  • Senior Compiler Engineer - AI

Senior Compiler Engineer - AI

NVIDIA

Santa Clara, CA 4 days ago

  • Senior Software Engineer, AI Speed Infrastructure

Senior Software Engineer, AI Speed Infrastructure

NVIDIA

Santa Clara, CA 2 weeks ago

  • Software Engineer, Multiverse

Software Engineer, Multiverse

Waymo

San Francisco, CA 2 months ago

  • Senior AI Frameworks Engineer

Senior AI Frameworks Engineer

NVIDIA

Santa Clara, CA 1 week ago

  • Software Engineer, Multiverse

Software Engineer, Multiverse

Waymo

Mountain View, CA 2 months ago

Similar Searches

  • Help Desk Representative jobs

5,749 open jobs

  • Technical Service Representative jobs

16 open jobs

  • Technical Sales Assistant jobs

4,061 open jobs

  • Liaison Engineer jobs

31,800 open jobs

  • Senior Microbiologist jobs

1,762 open jobs

  • Technical Sales Representative jobs

47,541 open jobs

  • Senior Technical Sales Representative jobs

13,967 open jobs

  • Market Development Specialist jobs

11,317 open jobs

  • Senior Sales Administrator jobs

72,647 open jobs

  • Agronomist jobs

1,084 open jobs

  • Security Representative jobs

57,819 open jobs

  • Field Service Representative jobs

15,986 open jobs

  • Extension Educator jobs

37,149 open jobs

  • Career Services Specialist jobs

40,323 open jobs

  • Technical Supervisor jobs

41,786 open jobs

  • Quality Control Microbiologist jobs

1,624 open jobs

  • Construction Operations Manager jobs

3,372 open jobs

  • Licensed Aircraft Engineer jobs

2,741 open jobs

  • Technology Innovation Manager jobs

12,177 open jobs

  • Technical Services Specialist jobs

10,475 open jobs

  • Quality Expert jobs

44,880 open jobs

  • Field Representative jobs

97,863 open jobs

  • Reservoir Engineer jobs

1,754 open jobs

  • Capacity Planner jobs

14,239 open jobs

  • Flight Operations Engineer jobs

1,952 open jobs

Explore top content on LinkedIn

Find curated posts and insights for relevant topics all in one place.

View top content