HirusBrowse jobs
L

Luminal

Cloud Inference Engineer

San Francisco, CA, US · Not specified

Annual base salary
$150k – $250k USD
Equity
Not disclosed
Commitment
Full Time
Company stage
Not disclosed

Compensation as listed

$150K - $250K  •  0.15% - 0.75%

Qualifications

  • CUDA + GPU inference optimization
  • vLLM, SGLang, or TensorRT-LLM experience
  • KV caching, paged attention, batching, token streaming, etc.
  • Distributed compute (with GPUs is a super plus)
  • No degree required

Company

Luminal (YC S25) builds an AI compiler and serving stack that makes models 10x faster and production ready with one line.

Role

Founding, on site in downtown SF. Ship low latency, high throughput model serving on Luminal Cloud.

Day to day responsibilities:

  • Deploy and tune models with optimizations like KV caching, paged attention, sequence packing, etc.
  • Conducting model performance reviews
  • Improve scheduler, batcher, autoscaling; profile latency, cost, utilization
  • Sometimes write kernels and, yes, occasional tasteful shitposting

Technology

Luminal uses a search based approach to generate, tune, and verify GPU kernels so engineers do not have to hand write CUDA.

Search based approach

  • Express computations in a small IR, then generate candidate kernels via equality saturation rewrite rules (tiling, unrolling, vectorization, memory layout).
  • Guide exploration with cost models and bandit style search to find the fastest valid kernels for a target GPU.
  • Compile and benchmark candidates on real hardware, enforce correctness with property tests and equivalence checks, and keep strict shape and dtype constraints.
  • Cache, version, and reuse the best kernels across models and deployments with full reproducibility.

Tech stack

  • Compiler and runtime: Rust and egglog based compiler generating GPU kernels. Using a lightweight IR with e-graph style rewrites to search and benchmark kernels.
  • Backends: CUDA and Metal in production today. Other backends in progress.

Source: Y Combinator. Confirm availability with the employer.

Apply through the original posting.

View listing