Placem
Mercor

Member of Technical Staff, Deeptune Environments

Mercor

New York CityFull timeOn-siteOpenPosted 13 days ago
Apply on Mercor

About Mercor

Mercor's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. Mercor's APEX benchmark family measures AI's real-world impact on professional work. Mercor Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.

 

Mercor is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. Mercor is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.

About Deeptune

Deeptune is an RL environments lab within Mercor. We build training gyms for AI agents — high-fidelity simulations where models learn to solve economically valuable problems through reinforcement learning. We work with the leading AI labs to train the next generation of agentic models, and our environments have already contributed to recent breakthroughs in computer use, code generation, and multi-step task completion.

About the Role

This is a backend and infrastructure role, not a research or model-training role.

You'll build the systems our environments run on: sandboxed execution, orchestration for thousands of concurrent rollouts, agent and task harnesses, tool APIs, the grading layer that decides whether an agent actually succeeded, and the pipelines that turn raw human demonstrations into training-ready environments.

You'll do this alongside researchers who own the RL side. Your job is to make their ideas real: fast, correct, and reliable at scale. That means you should know enough about post-training, evals, and reward modeling to push back on a research spec productively. It does not mean you'll be training models.

The work is high-ownership and lightly specified. You'll set direction, drive outcomes, and stay hands-on. You'll also work directly with AI labs and enterprise partners, which means shipping against real external deadlines rather than internal ones.

What You'll Do

  • Build the systems that build environments end to end: the simulated app or system, the agent-facing tool surface, the task definitions, and the verifiers that score them.

  • Design and operate the infrastructure that runs environments at scale: containers, sandboxing, orchestration, queues, observability

  • Turn messy human data into clean, reproducible training environments.

  • A non-deterministic or slow environment is worse than no environment

  • Own the interface with labs and researchers: translate a research goal into a system that exists next week.

What We're Looking For

We care more about what you've built than how long you've been building it.

  • 2+ years of production backend engineering, including at least 1 year at a startup, ideally as a founding engineer, an early engineer at a fast-growing venture, or a founder yourself

  • Strong generalist with systems depth. We primarily use Python and seek engineers who are fluent in at least one language. Ideally, they should be skilled at applying agents with discernment and able to adapt quickly to different problem requirements.

  • Understanding of ML/LLM concepts. Post-training, evaluation metrics, reward modeling: deep enough to partner with researchers and execute on novel RL environments.

  • Thrives in ambiguity. You scope your own work, make pragmatic calls, and ship without a spec handed to you.

  • Ability to raise the bar around you. Coaching engineers and driving execution, while staying in the code

Bonus, not required: sandboxing or virtualization, browser and computer-use automation, CI/build systems, developer tooling.

You're a Match If

  • Ownership, impact, and building frontier tech are what motivate you

  • Your work is a craft you want to master

  • You thrive in ambiguity and like hard problems

  • You appreciate diverse perspectives and uncommon ideas

  • You're excited to build in person, 5 days a week, 10 am–8 pm ET, from our office at One World Trade