stars 1 stars 2 stars 3

Cedana automatically saves, migrates, and resumes live GPU state continuously to maximize tokens per MW. It runs transparently and requires no code changes. Work survives interruptions and fleets scale elastically with demand. GPUs are expensive and mostly underused. When something goes wrong (a hardware failure, a maintenance window, a spot reclaim), work is lost and the job restarts from scratch. And inference demand is spiky while GPUs aren't: keeping models warm wastes money, and cold-starts are slow. Cedana’s job migration primitive resolves both. On failure, work resumes in seconds with no lost progress. For elasticity, Cedana resumes a fully loaded model from a live snapshot 10-20x faster, so provisioning can track demand closer to real-time instead of paying for peak around the clock. On real customer clusters, useful-work time rises from 30-40% to 80%+, and models that take minutes to load go live in seconds. Caltech: ~80% lower compute cost, results 2x faster. Works with Slurm, Kubernetes, Kueue, Ray, Armada, and Nvidia Dynamo.

Cedana Questions

Neel Master is the CEO and Co-founder of Cedana.

11 people are employed at Cedana.

Top Cedana Employees

View Similar People
G2 Leader Summer 2026 G2 Best Est ROI Mid-Market Summer 2026 G2 Easiest Admin Mid-Market Summer 2026 G2 Most Implementable Summer 2026 G2 Best Results Mid-Market Summer 2026 G2 Lead Capture Mid-Market Summer 2026 Inc Fastest Growing Private Companies 2026 Inc Best Workplace 2026
g2crowd
G2Crowd Trusted
chromestore
300K+ Plugin Users