About MLSrvyn

We built MLSrvyn because we got tired of fixing the same cold-start problem ourselves

Before MLSrvyn, we ran ML infrastructure at companies where inference fleet ops was a daily fire drill. Cold-starts at 9am. P99 spikes on Friday afternoons. GPU costs growing faster than revenue. We know what breaks and why.

Why we started

In 2022, Alex Petrov and Jordan Wei left ML infrastructure roles at a major e-commerce company, where they had spent three years watching the same infrastructure gap appear at every org that deployed ML models at scale.

The tools available — Kubernetes HPA, custom RPS-based scripts, manual capacity planning spreadsheets — all shared the same fundamental assumption: that inference workloads behave like web services. They don't. A recommendation model during a marketing email blast, an LLM session at midnight, a fraud scorer during ACH settlement — each has a different traffic shape, a different resource profile, and a different failure mode.

MLSrvyn is the tool we wish had existed. Per-model traffic profiling. SLO-driven scaling policies instead of RPS thresholds. GPU bin-packing that actually knows what a model needs. A control plane that treats each inference workload as a distinct first-class thing.

2022
Year founded
San Francisco
Headquarters
Bootstrapped
Funding stage

The team

Alex Petrov, CEO and Co-Founder of MLSrvyn

Alex Petrov

CEO & Co-Founder

Previously ML Infrastructure Lead at a high-growth e-commerce company, where he managed GPU fleet scaling for a 200+ model production environment. Background in distributed systems and container orchestration.

Jordan Wei, CTO and Co-Founder of MLSrvyn

Jordan Wei

CTO & Co-Founder

Previously Staff ML Engineer at a fintech company, where she designed the real-time scoring infrastructure for fraud detection across 40M+ daily transactions. Deep expertise in latency-sensitive inference systems.

How we work

Measure before acting

Every scaling decision MLSrvyn makes is backed by real latency data, not heuristics. We hold ourselves to the same standard in our product work.

Published methodology

Our benchmarks and architecture decisions are documented and reproducible. We don't ask you to take our numbers on faith.

Inference-first engineering

We're not building general infrastructure that also happens to work for inference. We're building tools designed from the ground up for how inference workloads actually behave.

We're hiring engineers who've run ML fleets at scale

Small team, hard problems, real production systems. Get in touch if you've spent time on the wrong side of a 3am GPU incident.