RL Infrastructure Engineer | Frontier AI Research | San Francisco – AIONIA
RL Infrastructure · Frontier AI · San Francisco

RL Infrastructure Engineer

A rare infrastructure role at the core of a frontier RL research operation. Seed-stage, well-funded, and staffed by researchers from top AI labs and PhD programs — building systems that automate task objectives and create scalable learnable curriculums across thousands of GPUs.

$300K – $500K + Equity San Francisco · On-Site 5d/wk Seed Stage · Well-Funded Team of 6–8 H1B Transfer · OPT · O-1 Supported 1 Opening · Urgent Hire
Apply via AIONIA

Frontier RL at the Earliest Stage

This is an early-stage AI research company focused on reinforcement learning in open-ended settings — building systems that automate task objectives and create scalable learnable curriculums. The team of 6–8 includes researchers from top frontier AI labs and PhD programs.

Seed-stage and well-funded, the company operates on weekly hypothesis-driven sprints. Each engineer develops a hypothesis and demos progress by end of week — a culture that rewards curiosity, speed, and rigorous thinking in equal measure.

"Each engineer develops a hypothesis and demos progress by end of week — working directly alongside researchers at the frontier."
Confidential Search

The client organization is confidential. All conversations are handled with full discretion. Apply or reach out directly — even if you're not actively looking.

Own the Systems Layer

This is an infrastructure engineering role at the core of a frontier RL research operation. You'll build the systems layer that enables researchers and applied ML engineers to run, debug, and reproduce large-scale RL experiments — covering distributed rollouts, training orchestration, inference, evals, data pipelines, observability, and reliability.

You'll own infrastructure projects end to end, from architecture through deployment and long-term maintenance — translating messy experimental workflows into durable, scalable infrastructure alongside some of the sharpest RL minds in the field.

1 hire — urgent, target close within 1–3 months.

What You'll Build

  • Infrastructure for distributed RL training and inference across thousands of GPUs
  • Reliability, debuggability, and throughput improvements for large-scale RL experiments
  • Interfaces for researchers and applied ML engineers to launch, inspect, compare, and reproduce experiments
  • Elimination of bottlenecks in training, rollout generation, eval execution, data movement, and cluster utilization
  • Engineering standards for RL infrastructure: testing, observability, versioning, and reproducibility

Technical Domain

Python PyTorch vLLM SGLang Kubernetes GPU Clusters veRL FSDP DeepSpeed SkyRL Slime Distributed Training

Who You Are

Must-Have — Non-Negotiable
  • 2+ years building infrastructure for LLM or RL systems
  • Experience at a high-engineering-bar organization — top AI startups, frontier labs, or Big Tech RL research teams
  • Hands-on experience with GPU clusters, distributed training, model serving, or high-throughput inference systems
  • Familiarity with vLLM, SGLang, and modern LLM-RL training frameworks
  • Degree in CS, EECS, Mathematics, or a related field
  • Very high level of curiosity and hypothesis-driven thinking
Nice to Have
  • Experience working closely with ML researchers building infrastructure for messy experimental workflows
  • Evidence of strong independent technical work — open-source projects, competitions, or notable infrastructure contributions
  • Familiarity with veRL, SkyRL, Slime, FSDP, DeepSpeed, or similar distributed training frameworks
Do Not Apply If

Candidates who match any of the following will not be considered, regardless of other qualifications.

  • No hands-on LLM or RL infrastructure experience
  • No meaningful GPU or distributed systems exposure
  • No evidence of strong technical ownership or independent work

Comp & Logistics

Base Salary
  • $300,000 – $500,000 base (average offer $400K–$425K)
  • Competitive equity at seed stage
Logistics
  • Location: San Francisco, CA (FiDi) — on-site 5 days/week
  • Visa: H1B transfer, OPT, and O-1 supported
  • 1 opening — urgent, target close within 1–3 months

Interview Process

1

AIONIA Intro Call

Background, fit, and logistics — confidential, no commitment required

2

Technical Interview

Deep-dive on infrastructure experience, systems design, and RL/LLM stack

3

On-site Loop

1–2 days on-site with researchers and engineering team in San Francisco