Staff ML Engineer, Agent Training & Environments

Labelbox
San Francisco, California, United States
Full Time
Hybrid
Senior
USD 250,000 – 280,000
Technology Information Technology Telecommunications

Posted 3 days ago

Apply to this job and hundreds more automatically

Get Started
Excellent 4.7

Job summary

The role involves building and running the systems for training agents in reinforcement learning environments, including experiment pipelines, verifiers, and evaluation infrastructure, focusing on reliability, scaling, and efficiency.

Responsibilities

You will build and maintain the infrastructure, environments, and evaluation systems used to train and test frontier AI agents. This involves developing verifiers, managing fine-tuning pipelines, and scaling training infrastructure to support high-throughput experimentation.

Qualifications

API Design Distributed Systems Evaluation Pipelines Fine-tuning Models ML Infrastructure Open-ended Work Verifiers or Graders Python Reinforcement Learning RL Algorithms (e.g., PPO, GRPO) System Design

The role requires over 3 years of experience shipping reliable production systems and deep expertise in post-training reinforcement learning methods. Candidates must demonstrate strong system design judgment and the ability to work effectively in a fast-paced, startup-style environment.

English

Company

Labelbox

Labelbox

Software Development
Location
San Francisco, California, United States

Apply to this job and hundreds more automatically

Get Started
Excellent 4.7
Pete DeQuarto avatar
Pete DeQuarto

AIApply is a total game-changer. Thanks to Auto-Apply, I have been applying to jobs I would never even have come across otherwise!

Akarshak Tanwar avatar
Akarshak Tanwar

In one week I applied to over 150 jobs and locked in 2 confirmed interviews, with a support team that was quick to respond. The auto-apply feature is fascinating and a real pleasure to use!

JC
Julian Carlin

It really worked. In under 48 hours it had applied to 100 jobs on my behalf without any time or effort from me. I am looking forward to seeing how it keeps improving.