Distributed Systems Engineer
$1–1
- On-site
- Full-time
- San Francisco
- Engineering
- 3mo ago
Job Description
About the role
Dedalus Labs builds persistent computers for AI agents. Our flagship product, Dedalus Machines, gives agents an isolated environment where they can run software, keep files and state, and work over time.
We're looking for an engineer who wants to design systems that continue working after individual machines fail. You care about what happens when a request times out after committing, two workers believe they own the same task, or a replica returns after missing writes. You want to make those cases understandable and testable.
You'll own distributed services from the initial contract through implementation, rollout, and operation. We need someone who can turn an incomplete goal into a small design, explain the tradeoffs, and finish the work with the engineers who depend on it.
What you'll work on
Keep persistent state correct. Build storage, replication, snapshot, and recovery mechanisms. Define when a write is acknowledged, which component can change state, and how a recovering node discovers what it missed. Test process loss, incomplete writes, duplicate messages, and delayed responses.
Coordinate compute under changing capacity. Build scheduling and orchestration for concurrent, long-running workloads. Reason about resource reservations, competing requests, stale information, and cancellation. Make the system recover when a worker disappears halfway through an operation.
Make failures diagnosable. Follow a symptom across services, storage, and networking. Build reproducible failure tests and useful logs, metrics, and traces. Use that evidence to simplify recovery and remove recurring manual intervention.
Improve performance under real workloads. Measure contention, queueing, replication, and data movement. Explain what limits throughput or tail latency before changing the design. Verify that the improvement preserves the system's correctness requirements.
What we look for
You have built a substantial distributed system in Go, Rust, C++, or a similar language. You can explain the consistency guarantees, ordering assumptions, failure modes, and operational tradeoffs of work you personally owned. Depth in storage, replication, consensus, scheduling, or distributed networking is especially relevant.
You can make progress in an unfamiliar codebase by reading source and specifications, constructing a reproduction, and choosing an experiment that distinguishes competing explanations. When you learn that an assumption was wrong, you revise the design and the tests.
You have carried an implementation through to a useful result and continued improving it after other people used it. Production services, open-source systems, research implementations, and technically ambitious independent projects can all demonstrate this. We evaluate your contribution and understanding, not an employer name or publication venue.
You work independently and make your reasoning available to the team. When the first explanation fails, you read the implementation, reduce the problem, collect better evidence, or ask a precise question. You change your approach when the evidence calls for it and keep responsibility for the outcome.
You can explain a design, engage with a substantive objection, and help another engineer become effective in the system. You use AI coding tools where they help and take responsibility for understanding, testing, and maintaining the resulting code.
This role may be a poor fit if
Your interest ends at the architecture diagram and you want someone else to implement, test, and operate the system.
You treat retries as sufficient recovery without reasoning about duplicate effects or uncertain state.
You prefer adding services and coordination layers before checking whether a smaller design meets the requirement.
Logistics
In person in San Francisco.
We sponsor visas.
Relocation support available.
Competitive salary and meaningful equity.
Meals and office benefits included.
Show us your work
Use your application to show us one thing you built and one difficult problem you investigated. They can come from the same project. Choose work you already have. We are not asking you to build a new project for this application.
Work: Show the code, demo, design, or result. Explain what you personally owned, what you reused, and one decision that made the result better or simpler.
Investigation: Describe the symptom, your first explanation, the evidence that changed your thinking, and how you established the result. Include the workload or users involved and any important limitations.
Writing: Share one piece of your best writing as a separate, required sample. Simple and succinct is welcome. Choose work that shows your thinking and your voice in a professional or similarly thoughtful context: a strong README, design note, blog post, postmortem, proposal, or essay. A serious, self-contained Reddit post is fine. Tweets, Twitter/X threads, and short social posts do not count. Link it or paste a redacted excerpt. Tell us the intended audience, what you wrote, why you chose it, and any coauthors or AI assistance.
Work from employment, open source, research, coursework, or independent projects is welcome. Public source code is optional. A redacted excerpt or a concrete account of private work is fine. Keep confidential information out of your application.
Tell us which problem in this role you want to own and why. If your strongest evidence is easy to miss on a resume, point us to it. An existing demo or short screen recording is welcome, and video is optional.