$1–1

  • On-site
  • Full-time
  • San Francisco
  • Engineering
  • 3mo ago

Job Description

About the role

Dedalus Labs builds persistent computers for AI agents. Our flagship product, Dedalus Machines, gives agents an isolated environment where they can run software, keep files and state, and work over time.

We're looking for a systems engineer who's bothered by a restore that usually works, a tail-latency spike nobody can explain, or an abstraction that makes a failure harder to understand. You want to follow the problem into the runtime, operating system, storage, or network until you can explain what happened and reproduce it.

You'll own substantial parts of the compute runtime. That means defining the behavior, implementing it, testing the failure paths, and carrying the change through rollout and operation. Other engineers should be able to use and maintain what you build.

What you'll work on

Make persistent execution correct. Build and improve snapshot, restore, sleep, resume, and migration mechanisms. Work through what happens when a process dies, a request is canceled, or only part of an operation reaches storage. State what must remain true, then build a test that would catch it becoming false.

Make the runtime faster. Trace startup and execution latency across application code, the hypervisor, kernel, filesystem, and network. Separate time spent computing from time spent waiting. Compare changes on representative workloads, including contention and cold state.

Make isolation and resource controls understandable. Work on virtualization, device interfaces, networking, memory, and scheduling. Define who owns each resource, when it is released, and how failure is reported. Build the diagnostics the next engineer will need.

What we look for

You have built substantial software in Rust, Go, C, C++, or a similar language and can explain its memory use, concurrency, and failure behavior. Rust is preferred but not required. We want depth in at least one of virtualization, operating systems, storage, networking, runtimes, or performance analysis, together with the ability to learn the layers around it.

Show us a difficult implementation or investigation you personally carried through. A kernel or runtime fix, storage-corruption reproducer, isolation mechanism, or carefully measured performance improvement can all be strong evidence. Production, open-source, research, and independent work are relevant. You do not need to have used every component in our stack.

You choose a small design that satisfies the requirement. You can explain an approach you abandoned, a component you reused, or code you deleted after understanding the problem better. You know which result you measured and which guarantee still needs testing.

You work independently and make your reasoning available to the team. When the first explanation fails, you read the implementation, reduce the problem, collect better evidence, or ask a precise question. You change your approach when the evidence calls for it and keep responsibility for the outcome.

You can explain a design, engage with a substantive objection, and help another engineer become effective in the system. You use AI coding tools where they help and take responsibility for understanding, testing, and maintaining the resulting code.

This role may be a poor fit if

  • You want to stay within an application framework and hand off problems once they reach the operating system or runtime.

  • You need every implementation step specified before beginning an investigation.

  • You lose interest once the happy path works, before testing recovery, documenting the behavior, or supporting the rollout.

Logistics

  • Full-time and in person in San Francisco.

  • Relocation support is available.

  • We sponsor visas for exceptional talents.

  • Competitive salary, bonus, and meaningful equity.

  • Meals and office benefits are included.

Show us your work

Use your application to show us one thing you built and one difficult problem you investigated. They can come from the same project. Choose work you already have. We are not asking you to build a new project for this application.

  • Work: Show the code, demo, design, or result. Explain what you personally owned, what you reused, and one decision that made the result better or simpler.

  • Investigation: Describe the symptom, your first explanation, the evidence that changed your thinking, and how you established the result. Include the workload or users involved and any important limitations.

  • Writing: Share one piece of your best writing as a separate, required sample. Simple and succinct is welcome. Choose work that shows your thinking and your voice in a professional or similarly thoughtful context: a strong README, design note, blog post, postmortem, proposal, or essay. A serious, self-contained Reddit post is fine. Tweets, Twitter/X threads, and short social posts do not count. Link it or paste a redacted excerpt. Tell us the intended audience, what you wrote, why you chose it, and any coauthors or AI assistance.

Work from employment, open source, research, coursework, or independent projects is welcome. Public source code is optional. A redacted excerpt or a concrete account of private work is fine. Keep confidential information out of your application.

Tell us which problem in this role you want to own and why. If your strongest evidence is easy to miss on a resume, point us to it. An existing demo or short screen recording is welcome, and video is optional.