Iterate.ai launches Lifeboat inference engine for private AI agents
Iterate.ai has launched Lifeboat for high-concurrency AI agent inference on customer-controlled infrastructure, targeting KV-cache pressure and workload scheduling.
Iterate.ai has launched Lifeboat as an enterprise inference engine for high-concurrency AI agent workloads on customer-controlled infrastructure. Organizations can run the software through containers, Kubernetes, desktop applications or a Python package, whether on premises, in a private cloud or in an air-gapped environment.
The product addresses a memory bottleneck that can emerge before a graphics processor runs out of computing capacity. An AI agent may call a model repeatedly, wait for tools and keep adding information to a long-running context. The inference system stores intermediate attention data in a key-value, or KV, cache to avoid recalculating the full context for every new token. Iterate.ai says those growing caches can exhaust available memory, forcing requests to queue or restart while memory-heavy sessions crowd out other work.
According to the company, Lifeboat’s optimization layer combines KV-cache compression and adaptive cache management with memory tiering, admission control and fair per-session scheduling. Iterate.ai also lists elastic memory, dynamic precision tuning, dynamic handling of mixture-of-experts models and separation of the initial prompt-processing phase from token-by-token generation. It positions these controls as a way to keep more long-lived agent sessions active on an existing GPU fleet.
The capacity and speed figures come from vendor benchmarks and have not been independently reproduced. Iterate.ai says one published test increased KV-cache capacity from 284,000 to 568,000 tokens and raised maximum concurrency at a 100% completion rate from 1,024 to 2,048 sessions. The company also claims throughput of 8,714 tokens per second, compared with 4,965 at 2,048 concurrent sessions, and says p99 time to first token at 128 sessions fell from 189 seconds to 1.5 seconds.
Iterate.ai says the test used full-precision BF16 model weights instead of quantizing the weights to obtain those gains. That quality claim has not been independently evaluated. The opened material also did not provide the complete baseline implementation or test configuration.
Deployment documentation lists NVIDIA, AMD and CPU-only container images, plus single-node and multi-node operation. Native desktop applications are offered for macOS and Linux, while the pip package supports macOS, Linux and Windows. Iterate.ai says the pip distribution includes only its GGUF engine; the GPU optimization layer behind its concurrency claims requires the containerized tensor engine. When the research was conducted, public artifacts were available through the company’s GitHub release repository and Docker Hub image.
Lifeboat focuses on model serving and inference scheduling. It therefore sits at a different layer of the agent stack from NVIDIA’s Open Agent Safety Platform, which focuses on runtime monitoring and containment.
For organizations that need to protect data while a model processes it, Iterate.ai offers a licensed confidential-computing option. The company says the option runs prompts and model weights inside a confidential virtual machine and can require successful attestation before the inference server starts. It requires an Intel Xeon host with TDX or an AMD EPYC host with SEV-SNP, plus a supported NVIDIA data-center GPU running in confidential mode; Iterate.ai says AMD GPUs are not supported for this option. No independent security audit or attestation assessment was available in the reviewed evidence.
More news

Meta sets 2027 deployments for MTIA 450 and MTIA 500 chips

AMD and Cerebras partner on disaggregated AI inference, claiming 5x efficiency gain

Google Research lays out unresolved privacy and security risks for AI agents
