Ornn
Seed
Senior Engineer, Infrastructure Services
New York · Not specified
- Annual base salary
- Salary not listed
- Equity
- Not disclosed
- Commitment
- Full Time
- Company stage
- Seed
Senior Engineer, Infrastructure Services
New York\nEng\nIn office\nFull-time\nSenior Engineer, Infrastructure Services\n\n
About Ornn
\n\nOrnn is building the financial infrastructure for AI compute. Our price indices are live on Bloomberg Terminal. We structure and trade compute hedging instruments. And we're now building a platform that brings exchange-grade mechanics like order management, matching, scheduling, and settlement to how compute capacity gets reserved and allocated. We're a lean team in New York, backed by leading venture and strategic investors.\n\n
About the Role
\n\nYou would be one of the first infrastructure engineers responsible for the systems underneath Ornn Compute. Compute spans heterogeneous GPU clusters, datacenter infrastructure, schedulers, virtualization, networking, and the control plane that makes physical compute capacity available as a reliable financial product.\n\n\nThis is not a traditional cloud infrastructure role. You will work across Linux hosts, Kubernetes and Slurm clusters, containers and VMs, bare-metal provisioning, node management, networking, storage, and distributed control systems.\n\n\nYou will work directly with our Head of Engineering and have significant ownership over how Ornn operates infrastructure across multiple datacenters and compute providers. Expect to debug failures across the entire stack, from a process inside a container down to the physical node it is running on.\n\n
What You'll Do
\n\nBuild and operate the infrastructure layer powering Ornn Compute across heterogeneous GPU clusters and datacenters.\n\n\nDesign systems for provisioning, scheduling, isolating, monitoring, and managing compute across bare-metal, containerized, and virtualized environments.\n\n\nBuild infrastructure software in Rust and Python for node management, orchestration, telemetry, networking, and cluster operations.\n\n\nOperate and extend Kubernetes and Slurm environments, including technologies such as RKE2, Slinky, container runtimes, and GPU scheduling.\n\n\nDesign secure multi-tenant compute environments using containers, VMs, hypervisors, Kata Containers, device passthrough, virtio, and related isolation technologies.\n\n\nDebug complex failures across Linux, networking, storage, schedulers, virtualization, GPUs, and distributed systems.\n\n\nBuild systems for collecting and querying infrastructure state, telemetry, inventory, health, and utilization data.\n\n\nAutomate infrastructure deployment and lifecycle management across clusters rather than relying on one-off operational procedures.\n\n\nMake reliability a feature: design systems that degrade predictably, recover automatically, and expose enough information to understand failures when they occur.\n\n
What We're Looking For
\n\nStrong Linux systems knowledge. You are comfortable debugging processes, networking, memory, filesystems, devices, permissions, namespaces, cgroups, and system-level performance issues.\n\n\nStrong Rust knowledge and experience building production systems software.\n\n\nWorking knowledge of SQL and experience designing or interacting with data-intensive backend systems.\n\n\nExperience operating containerized and virtualized workloads in production.\n\n\nDeep familiarity with Kubernetes and/or Slurm and the systems surrounding them. Experience with technologies such as RKE2, Slinky, Docker/containerd, Kata Containers, or similar infrastructure is strongly preferred.\n\n\nUnderstanding of virtualization fundamentals including hypervisors, KVM/QEMU-style architectures, virtio, device passthrough, and workload isolation.\n\n\nStrong understanding of distributed systems, including failure handling, coordination, state management, idempotency, consistency, and observability.\n\n\nComfort working across abstraction boundaries. You should be willing to debug an application, kernel interaction, network path, scheduler, hypervisor, or physical server depending on where the problem actually is.\n\n\nAbility to operate independently in a lean engineering team and own systems from design through production.\n\n\nBS, MS, or equivalent experience in computer science, computer engineering, electrical engineering, or a related technical field.\n\n
Nice-to-Haves
\n\nExperience deploying or operating infrastructure inside datacenters, particularly large GPU or HPC clusters.\n\n\nKnowledge of modern GPU architecture, including CUDA execution, GPU memory, kernels, PCIe, NVLink, RDMA, GPUDirect, and multi-GPU communication.\n\n\nExperience debugging GPU workloads and performance at the kernel, runtime, or communication layer.\n\n\nUnderstanding of server management and node tenancy architectures, including BMCs, IPMI, Redfish, iDRAC/iLO, out-of-band management networks, and secure tenant access.\n\n\nExperience with high-performance networking such as InfiniBand, RoCE, RDMA, BGP, ECMP, or high-bandwidth Ethernet fabrics.\n\n\nExperience with distributed storage systems such as Ceph, WEKA, or other high-performance storage architectures.\n\n\nExperience building infrastructure for cloud providers, neoclouds, HPC environments, exchanges, or other systems where downtime and incorrect state have direct financial consequences.\n\n
Why This Role Matters
\n\nOrnn Compute ultimately turns physical compute infrastructure into something that can be scheduled, reserved, financed, and traded reliably.\n\n\nThat abstraction only works if the underlying machines can actually be provisioned, isolated, monitored, recovered, and delivered as promised. A failed node, broken network path, scheduler inconsistency, or tenancy issue is not merely an infrastructure problem when contractual capacity commitments depend on the system being correct.\n\n\nThe infrastructure layer is therefore foundational to everything Ornn is building. The systems you build will determine how reliably Ornn can onboard new datacenters, expose capacity to customers, and scale from individual clusters to a distributed compute market.\n\n
Compensation and Benefits
\n\nBenefits include competitive salary, meaningful equity, health coverage, free meals, and additional benefits. This is a high-ownership engineering role with significant influence over Ornn's infrastructure architecture and technical direction.\n\n
Equal Opportunity Statement
\n\nOrnn is committed to building a diverse team. We evaluate candidates based on their ability to do the work, not on pedigree or background. We encourage applications from people of all backgrounds and experiences.\n\n\nReady to apply?\nPowered by\nFirst name *\nLast name *\nEmail *\nLinkedIn URL\nResume *\nClick to upload or drag and drop here\nApply\nReq ID: R11
Source: a16z portfolio. Confirm availability with the employer.
Apply through the original posting.
View listing