By Interestana AI Editorial — AI-drafted, human-overseen. How we report
Architect Launches Real-Time Auction for LLM Inference
Architect Financial Technologies launched Liquid Inference, a novel LLM inference marketplace that operates as a real-time auction for every request. This product allows LLM inference providers to bid on serving user prompts, with the buyer ultimately paying the lowest offer that adheres to their specified rules. The system is designed for developers to easily integrate by swapping a base URL, maintaining their existing code, and enabling providers to compete on price for inference services.
Liquid Inference functions as an exchange-style router for LLM inference. According to Architect, providers submit offers to serve specific models, and each incoming request is auctioned among all providers quoting that particular model. The winning bid is the lowest qualifying offer that meets the buyer's criteria. The product originates from Architect, a trading firm known for operating the AX perpetual futures exchange, rather than an AI research laboratory. In May 2026, Architect acquired a U.S. Designated Contract Market with the intention of listing GPU compute futures, pending regulatory approval. The company leveraged its expertise in building financial exchanges to develop a two-sided price discovery mechanism for LLM inference.
The inference auction process involves four key steps. First, a client initiates a request, which is formatted as a standard OpenAI or Anthropic API call. Second, the buyer's predefined constraints filter the eligible offers from providers. Third, an auction takes place where providers quoting the requested model compete, and the lowest offer meeting the buyer's rules is selected. Finally, the maximum price is locked before the generation process begins, and billing is based solely on metered usage. Buyers have the flexibility to set per-job cost caps, limits for time to first token, and minimum throughput requirements. Additional customization options include specifying approved regions, mandating zero data retention, and utilizing provider or model allow lists. An 'Auto' mode is also available to automatically select the most suitable model for a given task. Harrison's LinkedIn post further detailed the inclusion of routing rule presets and comprehensive multi-modal support. Account holders can access market data such as live order books, per-provider and per-model quotes, and cleared transactions, offering an unusual level of transparency for an LLM API.
For buyers, Liquid Inference offers drop-in compatibility with various agentic coding tools. The product's compatibility list includes popular tools such as Claude Code, Codex, OpenCode, Cursor, Pi, and Cline. The platform is accessible via a free email signup, and users can view detailed market data. This new offering aims to introduce greater efficiency and cost-effectiveness into the LLM inference market by applying financial market principles to AI infrastructure.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.