Aolani and FriendliAI team up to scale AI inference infrastructure

Singapore-founded neocloud Aolani is partnering San Francisco-based FriendliAI to provide GPU cloud infrastructure for AI inference workloads, as demand for production-scale AI services accelerates globally and across Asia.

The collaboration positions Aolani as a key infrastructure provider for FriendliAI’s inference cloud, which serves developers and enterprises deploying open-weight and custom AI models in production.

FriendliAI is founded by researchers who invented continuous batching, a technique now standard in AI inference serving. It has built an end-to-end stack from optimised GPU kernels to global distribution to run workloads fast and reliably at scale.

“We’re seeing inference needs grow faster than companies can find compute to support and service their customers. To narrow the supply and demand gap, we actively partner with companies like FriendliAI to deliver compute capacity on time, at scale, and to rigorous standards,” said Nicholas Chia, Chief Executive Officer of Aolani.

“Businesses need the freedom to choose the AI models that best suit their applications and the ability to run them efficiently in production. Our job is to deliver high-performance, reliable inference so developers can focus on building their AI applications,” said Byung-Gon Chun, Founder and CEO of FriendliAI.