Posted
01 Aug 2026
Last seen
07 Aug 2026
Location
San Francisco, California, US
Lifecycle
mature
Grade
C
yc-startup vc-backed
Morph was a 1 person company from 0 → 10M of revenue. We will be the first 10 person $10b company.
Every employee should contribute >30M of revenue/yr to the company.
The best candidates would be top 1% at multiple parts of the inference stack.
work on PD disaggregation research
Morph builds the inference infrastructure behind the fastest open models. Our stack spans kernels, model serving, routing, autoscaling, and capacity. We are hiring a performance engineer to make the entire system faster, cheaper, and more reliable.
What you’ll do
Find the gap between theoretical hardware performance and production performance
Trace latency and throughput regressions from the API layer down to individual kernels
Optimize batching, scheduling, routing, quantization, and distributed execution
Build benchmarks and observability that make bottlenecks obvious
Work with NVLink and RoCE
Validate that every optimization preserves model quality and correctness
You might be a fit if you
Have optimized complex production systems
Can juggle 8+ Codex/Claude/other coding agents concurrently
Understand GPU performance, memory bandwidth, collectives, and inference serving
Are strong in Python, CuTEdsl, and comfortable navigating unfamiliar codebases
Can turn profiling data into clear engineering decisions
Care about tokens per second, tokens per dollar, and correctness equally
You will work directly with the founders on problems that determine how efficiently frontier-scale models can be served. Small team, enormous compute, immediate production impact.
Every employee should contribute >30M of revenue/yr to the company.
The best candidates would be top 1% at multiple parts of the inference stack.
work on PD disaggregation research
Morph builds the inference infrastructure behind the fastest open models. Our stack spans kernels, model serving, routing, autoscaling, and capacity. We are hiring a performance engineer to make the entire system faster, cheaper, and more reliable.
What you’ll do
Find the gap between theoretical hardware performance and production performance
Trace latency and throughput regressions from the API layer down to individual kernels
Optimize batching, scheduling, routing, quantization, and distributed execution
Build benchmarks and observability that make bottlenecks obvious
Work with NVLink and RoCE
Validate that every optimization preserves model quality and correctness
You might be a fit if you
Have optimized complex production systems
Can juggle 8+ Codex/Claude/other coding agents concurrently
Understand GPU performance, memory bandwidth, collectives, and inference serving
Are strong in Python, CuTEdsl, and comfortable navigating unfamiliar codebases
Can turn profiling data into clear engineering decisions
Care about tokens per second, tokens per dollar, and correctness equally
You will work directly with the founders on problems that determine how efficiently frontier-scale models can be served. Small team, enormous compute, immediate production impact.