3 min readfrom Machine Learning

Real-Time Conversational Agents (RTCA) Workshop @ NeurIPS 2026 — submissions now open, deadline Aug 29 AoE [N]

Our take

Advance the frontier of conversational AI at the Real-Time Conversational Agents (RTCA) workshop, NeurIPS 2026, in Sydney. Submissions are now open, with a deadline of August 29 AoE. This workshop addresses the critical gap between offline benchmarks and the realities of deployed, interactive agents, focusing on real-time generation, naturalness in interaction, and robust evaluation methods. Explore topics like streaming language models and multimodal alignment—and discover how optimizing data layers for low-latency workloads, as discussed in "Presentation: From ms to µs," can be key.

The burgeoning field of conversational AI has long been hampered by a disconnect between research and real-world application. While impressive offline benchmarks demonstrate remarkable progress in language generation and understanding, the reality of deployed conversational agents often falls short, exhibiting robotic behaviors that undermine user experience. The upcoming Real-Time Conversational Agents (RTCA) workshop at NeurIPS 2026 signals a crucial shift in focus, directly addressing this gap. It's encouraging to see this emphasis on real-time performance, particularly given the increasing importance of low-latency AI applications – a theme echoed in our recent presentation, Presentation: From ms to µs: OSS Valkey Architecture Patterns for Modern AI, which explores optimizing data layers for precisely these types of workloads. This workshop’s dedication to tackling the challenges of streaming speech, video, and language, alongside naturalness in interaction, represents a vital step toward bridging the gap between theoretical advancements and practical utility.

The RTCA workshop's threefold focus—real-time generation, interactional naturalness, and live system evaluation—is particularly astute. The current reliance on offline metrics is demonstrably inadequate for assessing the quality of live conversational agents. The call for submissions encompassing perceptual studies, turn-taking metrics, and even interactive Turing-style tests highlights a welcome move towards more holistic and user-centric evaluation methodologies. Furthermore, the inclusion of topics like prosody, emotion generation, and memory/grounding within live conversations acknowledges the complexities of creating truly engaging and believable conversational partners. The workshop’s open call for position papers and evaluation critiques encourages a broader discussion and invites contributions beyond traditional research papers, recognizing the need for diverse perspectives to address these multifaceted challenges. It's a thoughtful approach, especially considering the growing importance of trustworthy AI agents, as discussed in Building Trustworthy Snowflake AI Agents with Semantic Governance, where we examined the critical role of governance in ensuring responsible AI deployment.

The non-archival nature of the workshop and the single-round review process, while potentially streamlining the evaluation process, also underscore the exploratory and iterative nature of research in this area. The emphasis on demo papers and the on-stage Conversational Agents Showcase is particularly exciting, providing a platform for showcasing practical implementations and fostering a community around real-time conversational AI. The inclusion of safety, identity, and trust considerations—specifically addressing deepfakes, persuasion, and consent—demonstrates a growing awareness of the ethical implications of increasingly sophisticated conversational agents. This is vital as organizations, like Waymo, highlighted in At Waymo, an AI project isn't ready until its evals are — not when the model performs well, prioritize rigorous evaluation before deploying AI in high-stakes environments.

Ultimately, the RTCA workshop at NeurIPS 2026 presents a significant opportunity to accelerate progress in real-time conversational AI. The challenges are substantial – navigating latency constraints, achieving natural interaction, and developing robust evaluation methods – but the potential rewards are immense. As we move towards a future where conversational agents become increasingly integrated into our daily lives, the ability to create truly engaging, reliable, and trustworthy systems will be paramount. A key question to watch will be whether the workshop fosters a sustained shift in the field, moving beyond offline benchmarks and prioritizing the development of agents that seamlessly blend into the fabric of human interaction.

Real-Time Conversational Agents (RTCA) workshop at NeurIPS 2026 (Sydney, Dec 11–12). Submissions are now open on OpenReview.

What the workshop is about

Conversational AI has crossed into real-time deployment — voice modes, embodied avatars, full-duplex speech agents — but the published record is still dominated by offline benchmarks, and deployed agents still feel robotic (stilted turn-taking, missing backchannels, monotone prosody, awkward interruptions). Methods that work offline (non-causal attention, large beam search, multi-pass refinement, slow diffusion) often don't transfer to streaming, and the field lacks shared vocabulary and benchmarks for interactional naturalness as distinct from per-utterance quality.

The workshop is organised around three intertwined questions:

  1. Real-time generation under hard latency budgets — streaming speech, video, and language
  2. Naturalness in interaction — prosody, gaze, timing, grounding, turn-taking, backchannels
  3. Evaluation of live systems, where standard offline metrics fall short

Topics of interest (non-exhaustive)

  • Streaming/low-latency speech synthesis, ASR, and full-duplex audio–language models
  • Real-time talking-head, avatar, and embodied video generation
  • Streaming language models; incremental and speculative decoding for dialogue
  • Turn-taking, backchanneling, interruption handling, floor management
  • Multimodal alignment under latency and partial-observation constraints
  • Prosody, emotion, and paralinguistic generation in interactive settings
  • Memory, grounding, and tool use during live conversation
  • Evaluation of naturalness: perceptual studies, turn-taking metrics, perceived latency, interactive Turing-style tests
  • Datasets and benchmarks for interactive (not offline) evaluation
  • Efficient inference, on-device deployment, systems–quality trade-offs
  • Safety, identity, and trust in real-time agents (deepfakes, persuasion, consent)

Position papers, evaluation critiques, and reproducibility studies are also welcome.

Submission tracks

  • Full papers — up to 8 pages
  • Short papers — up to 4 pages (work in progress, focused contributions, position papers)
  • Demo papers — extended abstract or up to 2 pages; required for the on-stage Conversational Agents Showcase

NeurIPS 2026 style file, double-blind. Non-archival — authors retain the right to publish elsewhere. Single-round review, no rebuttal.

Key dates (End of day, AoE)

  • Submission deadline: 29 August 2026
  • Author notification: 29 September 2026
  • Workshop: 11 or 12 December 2026, Sydney

Confirmed invited speakers

  • Dimitris Samaras (Stony Brook) — visual behaviour and gaze in interaction
  • Evonne Ng (Meta Reality Labs / UC Berkeley) — conversational avatar dynamics (provisional)

Links

Happy to answer questions in the comments — including about the demo track (we have an on-stage Showcase running deployed systems live) and what we'd consider in-scope vs out-of-scope for the eval pillar. Also happy to hear opinions on what's missing from the topics list; the CFP wording still has room to move if there's a clear gap.

submitted by /u/Few-Ferret9700
[link] [comments]

Read on the original site

Open the publisher's page for the full experience

View original article