By Interestana AI Editorial — AI-drafted, human-overseen. How we report
PolyAI Releases Audio-Native Dialog Model Dialog-RSN-1
PolyAI has introduced Dialog-RSN-1, an audio-native dialog model designed to process caller audio directly rather than relying on pre-generated transcripts. This innovative approach integrates several key conversational AI components: turn-taking, speech recognition, function calling, and response generation into a single, unified model. Dialog-RSN-1 is currently operational and handling live production calls for enterprise clients. A significant architectural distinction of Dialog-RSN-1 is its audio-awareness on the input side, while text-to-speech (TTS) remains a separate component. This separation ensures that the output voice remains controllable, allowing for customization and brand consistency. The model operates on a request-based Large Language Model (LLM) architecture, meaning it is probed on demand rather than functioning as an always-on stream that would require continuous GPU resources. The model's first output token is dedicated to turn-taking, with possible classifications of EMPTY, ONGOING, or COMPLETE, indicating the status of the caller's speech. PolyAI reports substantial performance improvements with Dialog-RSN-1, including sub-300ms response times. Specific gains cited include an 11% relative containment increase for a restaurant group and a 37% latency reduction for an insurance company. At its launch, Dialog-RSN-1 supports English language interactions exclusively. Deployment is currently facilitated through PolyAI's proprietary platform, with no immediate plans for open-weight releases or a public API. Existing PolyAI customers can enable the model immediately, while new enterprise clients can request early access. The target audience for Dialog-RSN-1 comprises large, high-call-volume enterprises, aligning with PolyAI's reported 100+ enterprise customers and over 2,000 live deployments, as noted during its $86 million Series D funding round in December 2025. Self-serve developers and small to medium-sized businesses (SMBs) are not the primary focus for this initial release. The model is designed for a broad range of industries, including restaurants, insurance, financial services, healthcare, hotels, retail, telecommunications, travel, and utilities. Its applications span critical business functions such as booking and reservations, billing and payments, customer authentication, intelligent call routing, order management, and technical troubleshooting. The development of Dialog-RSN-1 addresses limitations found in existing conversational AI architectures. Traditional cascaded stacks send only the Automatic Speech Recognition (ASR) system's most confident transcription to the LLM, discarding nuances like tone, hesitation, and recognition uncertainty. Tuning these cascaded systems often involves manual adjustments to end-pointing parameters and ASR biasing, which can be difficult to generalize across diverse use cases. Speech-to-speech models, such as those exemplified by GPT Realtime and Gemini Live, retain audio but embed the voice directly into the model, thereby restricting pronunciation control. Furthermore, always-on, full-duplex variants of these models can consume significant computational resources by dedicating a GPU for the entirety of a call. Dialog-RSN-1's audio-native design aims to overcome these challenges by processing audio directly and maintaining flexibility in voice output.
Original source — read the full reporting at the publisher:
Read on MarkTechPostGet the weekly AI digest
AI news + new model releases, weekly. Drafted by our agents, reviewed by humans.