Partnering with Sail: Agents Without Limits

Sail is building the inference stack and runtime for agents that work for hours, days and weeks. This is the token factory of the future.

Sail is building the inference stack and runtime for agents that work for hours, days and weeks. This is the token factory of the future.

For years, "using AI" has mostly meant humans typing into a box and waiting for an answer. The next chapter belongs to agents. Machines chase solutions to any problem with patience and tenacity, bound only by the compute and context they are given.

The demand for agents is here. The infrastructure to run them ambitiously is not. Sail is the platform to run agents ambitiously.

A new north star for AI infra

Sail is building for a world in which agents run for hours or days to solve our most complex problems. Sail purpose-built their inference platform to operate across a wide range of compute and flex to handle long running, agent tasks. To make this happen, the Sail team is pushing on two key technologies:

The first is their inference stack, the machinery that actually runs the models. Sail lives by a simple mantra, "there are no bad chips, only bad prices." The team took the inference stack apart and rebuilt it to squeeze far more out of every chip. They distribute work intelligently across providers so nothing stalls. They tap underused data-center capacity that would otherwise sit idle, particularly clusters too small or too old for anyone else to want. They're often among the first to try new accelerators as they hit the market, always searching for a way to find a productive place for them in their engine.

This improved stack allows agents to think longer and harder, spending billions of tokens on a single task, for a fraction of the cost. 

The second is Sailboxes, the workspace for agents that want to run forever. Sailboxes fluidly scale CPU and memory to match the user’s needs, so an agent waiting on an event or long-running tool call is billed almost nothing for that time. Customers can and do leave Sailboxes running for months; they’re basically immortal. Sailboxes underpin a wide range of customer use cases, from always-on security checks, monitoring every physical product on the web, and even powering a chess leaderboard that has been running for months at the Sail office.

The Sail team isn’t new to the world of compute infrastructure. Neil Movva started his career at NVIDIA pushing GPUs to their limits, then did the same for new generations of AI accelerators at Apple, and came full circle back to GPUs at Together AI. Samir Menon comes from Apple, where he built extremely reliable systems for billion-user scale. This is a team that knows and loves performance engineering, and will chase the speed of light for their inference engine to the ends of the earth. 

But it’s not just GPUs all day: the office combines technical brilliance with the occasional unicyclist intermission. The position at the top of their chess leaderboard is hotly contested by every new hire (h/t to the current champion, Nirvik for the streak so far!). Sail is an incredible team, expanding quickly. 

Making intelligence abundant

Sail’s raison d'etre is to make intelligence abundant. Every decision at Sail, from the chip level to the API, is about giving teams the tokens, scale, and runtime to build agents without limits. We’re thrilled to have been their partners from day 0. The future is bright. Let’s set sail!

share