Member of Technical Staff, ML Systems
Make the model 10x faster.
Confidential AI Infrastructure Startup · Menlo Park, CA
A deep-tech AI infrastructure startup rebuilding the training and inference stack for world models. Today's ML infrastructure was built for language models — this team is rebuilding it for video, image, and world-model workloads, co-designing across three layers at once: low-level GPU kernel optimization, distributed systems, and the algorithms and models themselves.
The company came out of stealth with public benchmarks already in hand: a leading open video-generation model running roughly 10x faster at half the cost on its stack, a 2K image-generation model running in about four seconds for three cents, and a real-time video model running faster than real time. The founding team — with prior experience across leading AI labs, hyperscalers, and infrastructure companies — works out of Menlo Park in person and has an API already in production. The company raised a $10M seed round and is approaching a Series A.
"You report to the CEO. He runs every screen himself and makes the hiring decision — there is no layer between the work and the person who decides. The work sits below the application layer: kernels, runtimes, and distributed engines for video and world models. Nothing here is agents or RAG."
You'll own speed and efficiency across the full ML systems stack — low-level kernels, distributed inference engines, and multi-node training and serving systems for image, video, and world-model workloads. You'll work directly alongside a founding team that between them covers distributed systems, kernel optimization, cloud infrastructure, and research.
You'll feel at home here if you'd rather make a video model ten times faster than train one.
- Optimize GPU and system performance for training and inference across image, video, and world-model workloads
- Profile and remove bottlenecks at the kernel, memory, system, and cluster level using Nsight and related tooling
- Write low-level optimizations in CUDA and Triton on code paths that run in production
- Build distributed inference and training engines for diffusion models across multiple GPUs and nodes
- Own communication performance — NCCL, RDMA over InfiniBand or RoCE, and disaggregated serving
- Build benchmarking and regression harnesses so performance gains don't slide back in production
- Worked on inference or training performance — GPU kernels, runtime, or distributed execution
- Optimized diffusion, video, image, or other multimodal model workloads
- Degree in Computer Science or a related quantitative field
- 1+ years of experience in deep learning inference or training systems, or distributed systems
- Built inside or contributed to an inference engine or runtime — vLLM, SGLang, TensorRT-LLM, or equivalent
Hiring Manager Screen
30-minute conversation with the CEO on background, motivation, and a first read on GPU/distributed systems depth.
Domain Deep Dive
60-minute technical round with a member of the founding team on kernels, inference, or distributed execution.
System Design
60-minute systems design session with a member of the founding team.
Optional Onsite
If not already done in person, a chance to meet the full team on-site in Menlo Park.
