AI Hiring Index

Alembic · Infrastructure · Senior · Posted 2026-06-12

Senior Network & Site Reliability Engineer

Alembic · San Francisco HQ · $210k–240k base

This range's midpoint is above 54% of posted infrastructure ranges at AI companies right now. See the salary index.

Apply on Alembic's site Watch Alembic for new roles

ABOUT US

Alembic is the pioneering Causal AI platform. We help the world's largest enterprises move past correlation to prove what actually drives business outcomes — the question marketing and growth teams have never been able to answer with confidence. Fortune 100 companies including Nvidia, Delta Air Lines, and Mars use Alembic to make multimillion-dollar decisions on trusted, causal evidence.

We're backed by a $145M Series B from WndrCo (founded by Jeffrey Katzenberg), Jensen Huang, Joe Montana, Prysm Capital, and Accenture. Our models run on our own NVIDIA DGX SuperPOD built on Grace Blackwell infrastructure — one of the fastest private supercomputers in the world. (We've melted GPUs getting here.)

ABOUT THE ROLE

We're building infrastructure that has to perform under real-world scale, reliability, and security demands — and we're looking for an engineer who wants to own the foundation it runs on. This isn't a traditional "keep the lights on" role.

You'll design and operate the global network and reliability layer behind one of the world's fastest private supercomputers — the fabric powering distributed compute, ML workloads, real-time analytics, and mission-critical enterprise systems. You'll work across networking, systems, automation, observability, and reliability engineering to scale a platform where performance genuinely matters, with real influence over architecture decisions.

It's a strong fit if you like solving deep infrastructure problems, building resilient systems, automating everything repetitive, and owning architecture rather than just maintaining it.

WHAT YOU'LL DO

- Architect and operate scalable, secure network architecture for high-security requirements and large-scale machine learning workloads.

- Own network device configuration management end to end, ensuring consistency and reliability across the fleet.

- Improve system and network reliability and performance through automation, observability, and proactive capacity planning.

- Implement and manage complex network protocols and connectivity, including BGP, VPNs, and WAN circuits and external peering.

- Build and maintain comprehensive monitoring, alerting, and incident response — SLOs, runbooks, and on-call rotations — and drive post-incident analysis and continuous improvement.

- Ensure security, compliance, and operational readiness across our network and cloud infrastructure.

- Partner across engineering and data science to drive a culture of performance and reliability.

WHAT WILL HELP YOU SUCCEED

- 8+ years in network or infrastructure engineering, including 5+ years in datacenter operations and/or systems and network administration.

- A strong background in network security, architecture, design, and operations.

- Extensive hands-on experience with network devices (firewalls, switches, load balancers) and large-scale architectures and protocols — BGP, QoS, MPLS, and IPsec VPNs.

- Experience designing and operating modern datacenter network fabrics (spine-leaf, EVPN/VXLAN, ECMP).

- Network automation and IaC tooling (Ansible, Terraform, Nornir, or similar), plus IPAM/DCIM platforms (NetBox, Infoblox, or similar).

- WAN engineering — carrier circuit provisioning and external network peering.

- Familiarity with Kubernetes networking (CNI plugins, ingress, service networking, network policy) and strong operational experience with Linux-based production infrastructure.

- Experience with monitoring and observability stacks (Prometheus, Grafana, Datadog, ELK, OpenTelemetry).

- Solid scripting (Python, Bash) to debug complex network and system issues and automate solutions, plus excellent cross-functional communication.

ALSO HELPFUL

- NVIDIA networking technologies — Cumulus Linux, InfiniBand, Spectrum-X, and BlueField DPUs (this is the fabric behind our SuperPOD).

- Familiarity with data-intensive platforms (Spark, Airflow, Kafka) and storage network protocols (NFS, LustreFS, iSCSI).

- Security practices …

New infrastructure roles at AI companies, every Monday. The week's openings in this function across 286 companies, plus the weekly index. Free.

More infrastructure roles at Alembic

See also: Infrastructure Engineer jobs · AI jobs in San Francisco Bay Area · Alembic salaries · Python jobs · InfiniBand jobs · Kubernetes jobs.

This listing is reproduced from Alembic's public careers feed and links to the original. AI Hiring Index is not the employer and does not accept applications. All Alembic roles · AI salaries.