Interestana
Home/News/Prime Intellect Launches Prime Inference Serving Platform
MarkTechPost••3 min read

By Interestana AI Editorial — AI-drafted, human-overseen. How we report

Prime Intellect Launches Prime Inference Serving Platform

Prime Intellect has launched Prime Inference, a new serving platform designed for frontier open-source artificial intelligence models. This platform provides users with both serverless endpoints and reserved capacity, leveraging Prime Intellect's own Graphics Processing Units (GPUs) distributed across multiple datacenters. Prior to its public release, Prime Inference processed an internal volume of nearly one trillion tokens daily. This substantial internal traffic was generated from various applications including Reinforcement Learning (RL) rollouts, the creation of synthetic data, model evaluations, and the operation of long-running coding agents. Prime Inference functions as the serving layer within Prime Intellect's broader open training stack, which already includes post-training tools such as prime-rl, verifiers, and sandboxes. The introduction of serving capabilities completes this loop, enabling deployed models to generate production traces that can subsequently inform and improve the training process. Prime Intellect reports that its GLM-5.3 endpoint is among the fastest available on the OpenRouter platform, boasting a near-zero tool-call error rate and maintaining 100% uptime since its inception. The platform offers two primary serving modes: serverless endpoints are designed for handling variable and unpredictable demand, while reserved capacity is optimized for sustained and consistent workloads. Prime Inference is compatible with OpenAI's SDKs, allowing users to direct their OpenAI API calls to `https://api.pinference.ai/api/v1`. The system ensures high availability through automatic failover mechanisms that reroute traffic to healthy deployments across datacenters in the event of an issue. Future hardware support includes NVIDIA Blackwell GPUs, with Vera Rubin GPUs listed as coming soon. The platform features unified billing with team-level usage tracking, although per-model pricing details are still being finalized and published in the documentation. The underlying serving stack is a sophisticated integration of NVIDIA Dynamo, vLLM, Mooncake, and FlashInfer. This stack was developed in collaboration with Inferact and NVIDIA, with Prime Intellect actively contributing fixes back to the upstream projects. The primary target workload for Prime Inference is agentic AI, where a typical agent turn can add approximately 6,000 tokens to a prompt that already contains 140,000 tokens. Prime Intellect benchmarks this complex mix using the SemiAnalysis AgentX framework and simulates injected cold arrivals to test performance under load. A key architectural innovation is the disaggregation of prefill and decode operations, with each running on separate GPU groups. NVIDIA Dynamo manages the routing of tasks, while vLLM executes the models on their respective GPU groups. Decoders retrieve computed Key-Value (KV) data through NIXL. Prime Intellect claims this architecture results in nearly 40% lower p90 inter-token latency in their internal tests. Furthermore, the platform incorporates cache-aware routing, where Dynamo's KV-aware router dynamically weighs the overlap of cached KV prefixes against the volume of queued work to optimize performance and resource utilization.

Original source — read the full reporting at the publisher:

Read on MarkTechPost

Get the weekly AI digest

AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.

Read next