AI Hiring Index

CoreWeave · Engineering · Unspecified · Posted 2026-08-18

Operations Engineer, MetalDev

CoreWeave · New York, NY/ Livingston, NJ · $109k–145k base

This range sits in the bottom 3% of posted engineering ranges at AI companies right now. See the salary index.

Apply on CoreWeave's site Watch CoreWeave for new roles

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. Learn more at  www.coreweave.com .

About the Role

CoreWeave's MetalDev team is seeking an Operations Engineer to support the services and automation used during data center bring-ups and in production. Reporting to the Engineering Manager, this hands-on role is designed for an engineer who can handle routine production issues with growing independence while developing deeper expertise in service reliability, observability, and hardware lifecycle management.

You will support the Redfish-based services and tools used to initialize, reboot, monitor, and recover infrastructure at scale. You will independently triage routine alerts and support requests, contribute to incident response and root-cause analysis, improve dashboards and runbooks, and deliver small automation or reliability changes with guidance on complex or high-risk work.

You will work with NVIDIA GPU servers, BMCs, DPUs, Cooling Distribution Units, NVLink switches, power shelves, and custom hardware while collaborating with Fleet Operations, Hardware Engineering, Service Engineering, and hardware and firmware vendors. Approximately 80% of the role focuses on production operations, troubleshooting, and incident response. The remaining 20% focuses on improving monitoring, documentation, operational processes, remediation capabilities, and automation.

Key Responsibilities

Production Support and Troubleshooting

Monitor team-owned services and fleet health, identify unhealthy devices, and coordinate remediation using established tools and procedures.

Independently troubleshoot routine initialization, reboot, provisioning, and service-health issues, escalating complex or high-risk problems appropriately.

Investigate problems across applications, Linux systems, networks, BMCs, servers, DPUs, power equipment, and cooling infrastructure using logs, metrics, and diagnostic data.

Perform and validate approved remediation, ensuring services and devices return to a healthy state.

Participate in incident response, maintain clear operational communications, and contribute to root-cause analysis and follow-up actions.

Participate in the team's on-call rotation after completing onboarding and a readiness review.

Observability and Reliability

Use Prometheus, Grafana, PromQL, logs, and service telemetry to investigate service and infrastructure problems.

Maintain and improve dashboards, alerts, and operational KPIs for team-owned services.

Identify noisy alerts, monitoring gaps, and recurring failure patterns, and recommend practical improvements.

Develop small scripts, tools, or automation that reduce manual effort and make remediation safer and more consistent.

Validate software, firmware, or process changes in test environments and support controlled production rollouts.

Documentation and Collaboration

Create and maintain run-books, troubleshooting guides, escalation procedures, and service-support documentation.

Capture incident findings and operational knowledge so the team can respond consistently and prevent recurrence.

Partner with Fleet Operations and engineering teams to reproduce issues, determine ownership, and track problems to resolution.

Support hardware and firmware vendor cases by collecting diagnostic evidence, tracking status, and validating vendor-provided fixes.

Communicate clearly, seek feedback, share knowledge, and escalate early when impact or risk is uncertain.

Minimum Qualifications

Two or more years of experienc …

New engineering roles at AI companies, every Monday. The week's openings in this function across 286 companies, plus the weekly index. Free.

More engineering roles at CoreWeave

See also: AI jobs in New York · CoreWeave salaries · Kubernetes jobs · Grafana jobs · Prometheus jobs.

This listing is reproduced from CoreWeave's public careers feed and links to the original. AI Hiring Index is not the employer and does not accept applications. All CoreWeave roles · AI salaries.