Waabi · Data · Staff+ · Posted 2026-09-23
Senior / Staff ML Ops Engineer
Waabi · Remote US & Canada · $184k–272k base
This range's midpoint is above 52% of posted data ranges at AI companies right now. See the salary index.
Apply on Waabi's site Watch Waabi for new roles
Waabi, founded by AI visionary Raquel Urtasun, is the leader in Physical AI. With a world-class team, we're unlocking the next era of autonomous transportation with technology that's powering commercial autonomous trucks and robotaxis. Waabi is backed by and partners with world leaders in AI, automotive, logistics, and deep tech.
With offices in Toronto, San Francisco, Dallas, and Pittsburgh, Waabi is growing quickly and looking for diverse, innovative and collaborative candidates who want to impact the world in a positive way. To learn more visit: www.waabi.ai
You will.. Build and evolve our training infrastructure on Kubernetes with Infrastructure — GPU scheduling, autoscaling, multi-node distributed jobs, capacity strategy, and the operators and workflow engines that keep long-running training reliable. Shape the developer-facing surface — CLIs, SDKs, job submission, templates, paved paths — designed with the teams who'll use them. Make the common case one command and keep the uncommon case possible. Shorten the inner loop. Time to first training run, edit-to-signal latency, local iteration before a job hits the cluster, fast failure over slow mystery. Measure it, publish it, drive it down. Evangelize best-in-class tooling and frameworks. Track what the ecosystem is shipping, evaluate honestly, and make the case with working prototypes and migration paths — or say plainly when a shiny thing isn't worth the switching cost. Strengthen the data and artifact layer. Dataset versioning, sharding, and high-throughput loading of large multimodal sensor data, so jobs saturate GPUs instead of waiting on I/O. Turn one-off Python into durable tooling — tested, documented, observable libraries, CLIs, and services with sane defaults, and deletions where they're overdue. Make experiments legible, with the teams who live in them: experiment hygiene, dashboards researchers trust, a real model registry, and lineage from dataset to checkpoint to simulation result. Ship CI/CD for models alongside autonomy and simulation, so a model change is validated the same way a code change is. Build observability across the ML stack — utilization, throughput, failure modes, queue times, cost per experiment. When a job fails at 3am on node 47, the researcher should find out why without you. Treat docs, onboarding, and support as product surface — golden-path guides, a new researcher productive on day two, office hours that turn repeat questions into shipped fixes. Drive adoption, not just availability. Prototype with real users, watch them work, iterate. A tool nobody adopts didn't ship. Make the platform boringly reliable — fewer failures, faster recovery, and none of the manual steps that quietly cost a team days. Build guardrails that don't feel like walls, with Security, IT, and Infrastructure: access controls, data handling, and cost governance that hold up in an IP-sensitive environment while staying self-serve.
Qualifications: 5+ years of software or infrastructure engineering, including tools or platforms used by other engineers and operating ML or data-intensive production systems. Hands-on Kubernetes expertise — GPU scheduling, autoscaling, Helm or equivalent, networking fundamentals, and the ability to debug a cluster under load rather than restart it. Excellent Python, and a track record of designing APIs and CLIs other people enjoy using. Practical AWS depth: object storage at scale, IAM, GPU compute, networking, cost management, and infrastructure as code (Terraform, Pulumi, or similar). Distributed training in PyTorch (DDP, FSDP, or similar), plus experiment tracking and model registry tooling — from the perspective of someone who made them pleasant for others to use. Fluency with containers, CI/CD, and modern build systems, including large monorepos. The ability to influence without authority: evaluate a framework on its merits, pilot it credibly, and persuade skeptical senior engineers to change how they work. A collaborative default …
More data roles at Waabi
-
Software Engineer, Labelling, Data & Automation
Waabi Toronto, ONDataCanadaCA$127k–225k5mo
-
Senior Software Engineer, Auto Labelling
Waabi Remote US & CanadaDataSeniorRemote US Canada$170k–220k6mo
See also: Machine Learning Engineer jobs · Waabi salaries · Python jobs · PyTorch jobs · NCCL jobs.
This listing is reproduced from Waabi's public careers feed and links to the original. AI Hiring Index is not the employer and does not accept applications. All Waabi roles · AI salaries.