Member of Technical Staff - Debugger Generation

Touring Capital
Touring Capital

IT

San Francisco, CA, USA

Posted on Aug 8, 2026
Infinity Artificial Intelligence Institute San Francisco Bay Area

Member of Technical Staff - Debugger Generation

Infinity Artificial Intelligence Institute San Francisco Bay Area

1 day ago Be among the first 25 applicants

See who Infinity Artificial Intelligence Institute has hired for this role

Save

  • Report this job

Member of Technical Staff - Debugger Generation

Company: Infinity

  • Team: Systems / AI Infrastructure
  • Location: San Francisco (on-site)
  • Type: Full-time

The Mission

Debugging an accelerator usually means chasing something that is already gone. A run hangs, a race condition fires, or a result refuses to reproduce, and the state that would explain it evaporated the moment execution moved on. The standard recourse is to instrument more, run again, and narrow in one increasingly detailed dump at a time; on non-deterministic failures, that loop may never converge. Numerical bugs are worse to localize by hand, because the corruption usually sits far upstream of where it finally surfaces.

A real debugger dissolves that problem by refusing to let the state disappear. It freezes the entire program, including where every value physically lives, and lets you move through execution in both directions. This role builds the agent that generates a debugger like that for a new chip, and its core is continuous checkpointing with restore-to-checkpoint: snapshot the run as it proceeds, then jump back to the instant before the hang, or before the numbers began to degrade, and inspect exactly what was and was not in memory, without rewriting code to reproduce the moment.

The same frozen-state representation turns out to be the right substrate for making code fast, not only for finding out why it broke. Restore to any point, change one part of an inference pass, measure the effect from there, and undo it by restoring, with no full rebuild and no full rerun. That collapse of the rewrite-and-rerun cycle into checkpoint-and-restore is where a 10 to 20x speedup in experimentation comes from. Because numerical bugs are almost always deterministic given their inputs, the same loop closes cleanly: restore, scan for NaN and inf, diff intermediate tensors against a known-good pass, ablate a precision change, restore, re-run. And because implementation timelines are increasingly set by how fast agents can attempt and discard hypotheses, the largest lever may be that this is a far better interface for an agent than writing throwaway scripts, a place to read every value and its location and invalidate its own wrong assumptions in seconds rather than runs.

The specification is set by the best debuggers the industry already has, the whole thing has to run on the chip itself with no simulator standing in, and it has to generalize across hardware, AI accelerators first but mobile SoCs like Snapdragon too. The near-term target is concrete, and unglamorous in the right way: today’s d-Matrix Corsair debugger permits exactly one breakpoint and no checkpointing at all. We’re co-creating its successor, with real breakpoints, continuous checkpointing, and restore to intermediate state, so that AMPs can run inside the debugger with Corsair state rewound to any point in the computation. That is the same chip Ignition, our bringup agent, took from first hardware access to tensor-parallel matmuls across all 32 compute units in 10 hours and to three frontier models end-to-end in 10 days, now live as the Infinity d-Matrix Cloud.

What you’ll work on

You’ll build the debugger and the agent that generates it. Depending on your strengths:

  • Continuous checkpointing and restore. Capture full program state at varying granularity and roll back to any saved point, and build the reverse-execution machinery beneath it: invert computations where they are invertible, and replay forward from the nearest checkpoint where they are not. Time reversal, made practical.
  • Breakpoints and state inspection. Halt a computation and surface every variable, its value, and its physical location, with metadata attached at each memory write that ties the value back to its node in the inference computation graph, so “what is this, and where did it come from” is answerable rather than archaeological.
  • In-debugger optimization. Make the state representation good enough that a change to an inference pass can be applied, measured, and ablated in place, so experimentation runs at checkpoint speed instead of rebuild speed.
  • Parallel debugging. Capture and faithfully reproduce how parallel execution units interact, gangs for instance, because that interaction is precisely where most race conditions live, and where single-threaded debuggers go blind.
  • Differential debugging. Diff a suspect inference pass against a known-good reference to localize where it diverged, numerically or structurally, instead of bisecting by hand.
  • Breadth across silicon. Bring the same capabilities to very different hardware, from AI accelerators to mobile parts like Snapdragon and its Hexagon DSP, on real silicon with nothing simulated in the loop.
  • The Corsair build. The first concrete target: adding real breakpoints and continuous checkpointing to the existing d-Matrix debugger, and running AMPs inside it with state restored to intermediate points in the computation.
  • The agent-facing interface. Expose all of it so an agent can pause, read every value and its location, and disprove its own hypotheses through the debugger rather than by writing and running throwaway scripts, which more than anything else sets how fast a hard bug gets solved.

What we’re looking for

We weight range and depth over any particular résumé. Strong candidates will have most of the following:

  • Genuine debugging depth. You’ve built a debugger, or depended on one hard enough to know exactly where they fall short, ideally having tracked a bug all the way down to the silicon.
  • A real model of program state on hardware: memory models, execution ordering, and what it actually costs to capture and restore all of it on a massively parallel machine.
  • Experience with time-travel, record-replay, or checkpoint-restore systems, or a strong pull toward building one from first principles.
  • An instinct for tools whose primary user is an agent, not a human at a keyboard, and for how that changes the interface.
  • Fluency in Python and a systems language: Rust, C, or C++.

Nice to have

  • Built or worked on a time-travel or record-replay debugger: rr, gdb reverse debugging, WinDbg Time Travel, or your own.
  • Firmware, bare-metal, on-chip debug, or JTAG experience.
  • Worked on AI accelerators or mobile SoCs such as Snapdragon and its Hexagon DSP.
  • Hunted numerical bugs in ML, the NaN-at-layer-forty kind that only manifests sometimes.
  • Familiarity with checkpoint-and-restore systems such as CRIU.

Who we are

Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We’ve signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We’re headquartered in San Francisco.

  • Seniority level Mid-Senior level
  • Employment type Full-time
  • Job function Engineering and Information Technology
  • Industries Software Development

Referrals increase your chances of interviewing at Infinity Artificial Intelligence Institute by 2x

See who you know

Get notified about new Member of Technical Staff jobs in San Francisco Bay Area.

Sign in to create job alert

Similar jobs

  • Member of Technical Staff - Debugger Generation

Member of Technical Staff - Debugger Generation

Touring Capital

San Francisco Bay Area 11 hours ago

  • Staff Software Engineer – Platform Debug

Staff Software Engineer – Platform Debug

Qualcomm

Santa Clara, CA 3 days ago

  • Staff Software Engineer, Compute Reliability

Staff Software Engineer, Compute Reliability

Waymo

Mountain View, CA 1 day ago

  • Staff ML Compiler Engineer

Staff ML Compiler Engineer

General Motors

Sunnyvale, CA 3 days ago

  • Staff ML Compiler Engineer

Staff ML Compiler Engineer

General Motors

San Francisco, CA 4 days ago

  • Software Engineer, LLVM and Production Toolchain

Software Engineer, LLVM and Production Toolchain

Google

San Jose, CA 1 week ago

  • Staff Engineer, Compiler

Staff Engineer, Compiler

Samsung Semiconductor

San Jose, CA 1 day ago

  • Member of Technical Staff - Compilers

Member of Technical Staff - Compilers

Gimlet Labs

San Francisco, CA

$180,000.00

$400,000.00

3 months ago

  • Compiler Engineer

Compiler Engineer

Intel

Santa Clara, CA 1 week ago

  • Staff Software Engineer, Source Client

Staff Software Engineer, Source Client

Google

San Jose, CA 3 days ago

  • GPU Performance Profiling Engineer

GPU Performance Profiling Engineer

NVIDIA

Santa Clara, CA 16 hours ago

  • Senior Software Engineer

Senior Software Engineer

Sony Interactive Entertainment

San Mateo, CA 1 week ago

  • Senior SWE

Senior SWE

Theorem

San Francisco, CA

$300,000.00

$500,000.00

9 months ago

  • Senior Developer Tools Engineer

Senior Developer Tools Engineer

Lemurian Labs

Santa Clara, CA 1 week ago

  • Hardware Tools Engineer

Hardware Tools Engineer

OpenAI

San Francisco, CA

$225,000.00

$445,000.00

2 weeks ago

  • Member of Technical Staff

Member of Technical Staff

Morph

San Francisco, CA

$175,000.00

$350,000.00

6 days ago

  • Development Tools Software Engineer

Development Tools Software Engineer

Intel

Santa Clara, CA 1 week ago

  • Member of Technical Staff (Software Engineer, Acceleration)

Member of Technical Staff (Software Engineer, Acceleration)

Perplexity

San Francisco, CA

$250,000.00

$405,000.00

2 weeks ago

  • Staff Software Development Engineer - Rust, Compilers, and GPU Systems

Staff Software Development Engineer - Rust, Compilers, and GPU Systems

AMD

San Jose, CA

$204,000.00

$306,000.00

1 day ago

  • Member of Technical Staff, Research

Member of Technical Staff, Research

hillclimb

San Francisco, CA

$200,000.00

$300,000.00

1 month ago

  • Staff+ Software Engineer, Experimentation

Staff+ Software Engineer, Experimentation

Anthropic

San Francisco, CA 2 weeks ago

  • Member of Technical Staff, Deployments

Member of Technical Staff, Deployments

dynamism

San Francisco, CA

$400,000.00

$500,000.00

4 days ago

  • Staff Engineer, Workbench Platform

Staff Engineer, Workbench Platform

Samsung Semiconductor

San Jose, CA 4 days ago

  • Software Engineer, GenAI Silicon Automation, DeepMind

Software Engineer, GenAI Silicon Automation, DeepMind

Google DeepMind

Mountain View, CA 1 day ago

  • Systems Engineer

Systems Engineer

Theorem

San Francisco, CA

$150,000.00

$250,000.00

9 months ago

  • Software Engineer, Kernel Performance & AI Tooling

Software Engineer, Kernel Performance & AI Tooling

OpenAI

San Francisco, CA

$266,000.00

$445,000.00

1 week ago

  • Member of Technical Staff (Software Engineer, Computer)

Member of Technical Staff (Software Engineer, Computer)

Perplexity

San Francisco, CA

$220,000.00

$405,000.00

1 week ago

People also viewed

  • Member of Technical Staff

Member of Technical Staff

dynamism

San Francisco, CA $400,000 - $500,000 4 days ago

  • Staff Software Engineer (TLM), Multiverse

Staff Software Engineer (TLM), Multiverse

Waymo

Mountain View, CA 2 weeks ago

  • Senior AI Frameworks Engineer

Senior AI Frameworks Engineer

NVIDIA

Santa Clara, CA 5 days ago

  • Tech Lead Software Engineer, Programming Language

Tech Lead Software Engineer, Programming Language

ByteDance

San Jose, CA 1 day ago

  • Member of Technical Staff - Coding Agents

Member of Technical Staff - Coding Agents

SpaceXAI

Palo Alto, CA 2 weeks ago

  • Member of Technical Staff @ Vela

Member of Technical Staff @ Vela

Vela

San Francisco, CA $120,000 - $300,000 2 weeks ago

  • Staff+ Software Engineer, Developer Acceleration

Staff+ Software Engineer, Developer Acceleration

Anthropic

San Francisco, CA 2 weeks ago

  • Member of Technical Staff - Observability

Member of Technical Staff - Observability

SpaceXAI

Palo Alto, CA 5 hours ago

  • Member of Technical Staff, Kernels

Member of Technical Staff, Kernels

Magic

San Francisco, CA $225,000 - $550,000 1 week ago

  • Staff Software Development Engineer - AI Profiler Tools

Staff Software Development Engineer - AI Profiler Tools

AMD

Santa Clara, CA $204,000 - $306,000 4 days ago

Similar Searches

  • Senior Member of Technical Staff jobs

8,067 open jobs

  • Member Technical jobs

62,993 open jobs

  • Senior Wireless Engineer jobs

35,731 open jobs

  • Physical Design Engineer jobs

6,753 open jobs

  • Staff Test Engineer jobs

2,677 open jobs

  • Principal Firmware Engineer jobs

1,780 open jobs

  • Line Technician jobs

109,890 open jobs

  • Lead Infrastructure Engineer jobs

16,227 open jobs

  • Support Team Manager jobs

83,600 open jobs

  • Vice President Software jobs

49,146 open jobs

  • Senior Lead Software Engineer jobs

49,381 open jobs

  • Principal Researcher jobs

4,530 open jobs

  • Switch Engineer jobs

9,391 open jobs

  • Staff Software Engineer jobs

64,945 open jobs

  • Lead Quality Engineer jobs

10,548 open jobs

  • Control Coordinator jobs

39,876 open jobs

  • Market Maker jobs

1,432 open jobs

  • Yield Engineer jobs

9,445 open jobs

  • Computer Scientist jobs

49,477 open jobs

  • Lead Test Engineer jobs

13,921 open jobs

  • House Supervisor jobs

29,485 open jobs

  • Core Engineer jobs

33,936 open jobs

  • Cable Technician jobs

11,124 open jobs

  • Principal Software Engineer jobs

73,845 open jobs

  • Logic Design Engineer jobs

1,858 open jobs

Explore top content on LinkedIn

Find curated posts and insights for relevant topics all in one place.

View top content