CfP | RTCA @ NeurIPS 2026 [R]
Our take
The emergence of the Real-Time Conversational Agents (RTCA) workshop at NeurIPS 2026 signals a crucial shift in the trajectory of conversational AI. Moving beyond the familiar landscape of text-based chatbots, RTCA focuses on the far more complex challenge of creating agents capable of natural, multimodal interaction in real-time. This focus addresses a critical constraint in current conversational AI – the persistent reliance on offline generation models that simply aren’t suited for dynamic, responsive dialogues. As Intuit's experience shows [Intuit scrapped its own AI agent architecture twice in four months. At VB Transform 2026, its AI VP called that the fast path], building truly functional agents isn't always a smooth process, and RTCA directly tackles a core obstacle. The call for papers highlights the need for shared benchmarks and methodologies, acknowledging a gap in the field that is vital to overcome as we move towards more sophisticated and integrated AI experiences.
The workshop’s emphasis on real-time generation, naturalness, and evaluation is particularly compelling. The technical hurdles outlined – latency, turn-taking, multimodal alignment – are significant and demand innovative solutions. It’s fascinating to see this focus on even seemingly minor aspects of interaction, such as backchannels and interruptions, being elevated to first-class problems. The inclusion of topics like prosody, emotion, and paralinguistic generation underscores the ambition to create agents that feel genuinely human. The consideration of safety, identity, and trust, particularly in the context of deepfakes and persuasion, is also vital. Roblox’s move to incorporate AI-powered game creation [Roblox launches an AI-powered game-creation feature in its mobile app] demonstrates the expanding accessibility of AI tools, and ensuring these tools are ethically sound and trustworthy becomes increasingly important as they become more integrated into our lives. We've also seen the continued emphasis on short-paper submissions [short-paper at ACL/EMNLP/EACL [R]], showing a need for rapid iteration and sharing of research in this quickly evolving space.
The scope of the topics calls for a truly interdisciplinary approach, bringing together researchers from speech, vision, language, HCI, and ML systems. This cross-pollination of expertise is essential to address the multifaceted challenges of real-time multimodal conversation. The workshop’s commitment to being non-archival – allowing authors to publish elsewhere – is a pragmatic decision that encourages rapid dissemination of findings. The focus isn't just on building impressive demos, but on developing robust, evaluable systems, as evidenced by the inclusion of rigorous evaluation methods like perceptual studies and interactive Turing-style tests. The call for datasets and benchmarks specifically tailored for interactive evaluation is a critical step towards fostering progress and enabling meaningful comparisons between different approaches.
Ultimately, the RTCA workshop represents a significant investment in the future of conversational AI. The shift from offline to real-time processing is not merely a technical upgrade; it’s a fundamental reimagining of how we interact with AI. As these agents move beyond text and into the physical and digital worlds, their ability to engage in natural, fluid conversations will be paramount. The challenge now lies in translating the research presented at RTCA into practical, deployable systems that can truly empower users and enrich our interactions – what novel interaction paradigms will emerge, enabled by these increasingly responsive and multimodal agents, and how will we ensure these interactions are both beneficial and ethically sound?
Call for Papers and Demos
Real-Time Conversational Agents (RTCA): Toward Natural Multimodal Interaction
1st RTCA Workshop [@]() NeurIPS 2026, Sydney, Australia 11 or 12 December 2026
Website: https://rtcaneurips26.github.io/
We are pleased to share the Call for Papers and Demos for the inaugural RTCA Workshop at NeurIPS 2026, focused on real-time multimodal conversational agents: streaming speech, video, and language generation; naturalness in interaction; and evaluation of live systems.
Conversational AI has moved from text chat into the real world, voice modes that talk back, embodied avatars, agents that share our screens and tools. To feel natural, these systems must operate in real time, streaming while continuously listening, watching, and re-planning. This is fundamentally harder than offline generation: latency, turn-taking, backchannels, interruptions, and cross-modal alignment become first-class problems that the offline paradigm sidesteps. Recent progress on full-duplex speech–language models, real-time talking-head generation, and streaming ASR shows the regime is feasible, but the field still lacks shared benchmarks, vocabulary, and methodology for interactional naturalness.
RTCA brings together researchers across speech, vision, language, HCI, social-signal processing, and ML systems around three intertwined questions: real-time generation under hard latency budgets, naturalness in interaction, and evaluation of live systems.
Topics of Interest
We invite original contributions on topics including (but not limited to):
- Streaming/low-latency speech synthesis, ASR, and full-duplex audio–language models
- Real-time talking-head, avatar, and embodied video generation; lip-sync, gaze, expressivity under streaming
- Streaming language models; incremental and speculative decoding for dialogue
- Turn-taking, backchanneling, interruption handling, and floor management
- Multimodal alignment under latency and partial-observation constraints
- Prosody, emotion, and paralinguistic generation in interactive settings
- Memory, grounding, and tool use during live conversation
- Evaluation of naturalness: perceptual studies, turn-taking metrics, perceived latency, interactive Turing-style tests
- Datasets and benchmarks for interactive (not offline) evaluation
- Efficient inference, on-device deployment, and the systems–quality trade-off
- Safety, identity, and trust in real-time agents (deepfakes, persuasion, consent)
Submission Types
We welcome:
- Full papers (up to 8 pages) — may be presented as posters and/or contributed talks.
- Short papers (up to 4 pages) — work in progress or focused contributions.
- Demo papers (Extended Abstracts or up to 2 pages)
All submissions must use the NeurIPS 2026 style file and be formatted for double-blind review. Page limits exclude references and appendices. Papers must be submitted in PDF format via OpenReview (portal link to be published on the workshop website).
The workshop is non-archival; authors retain the right to publish elsewhere.
Important Dates (End of day, Anywhere on Earth)
- Call for papers opens: 18 July 2026
- Submission deadline (papers and demos): 29 August 2026
- Author notification: 29 September 2026
- Workshop date: 11 or 12 December 2026
Organisers
- Niki Foteinopoulou — Tavus, United Kingdom
- Alessandro Conti — Tavus, Italy
- Jack Saunders — Tavus, United Kingdom
- Oya Celiktutan — King's College London, United Kingdom
- Cigdem Beyan — University of Verona, Italy
- Ioannis Patras — Queen Mary University of London, United Kingdom
For more information, visit our website https://rtcaneurips26.github.io/ or contact us at [rtca-workshop@googlegroups.com](mailto:rtca-workshop@googlegroups.com).
We look forward to your contributions!
[link] [comments]
Read on the original site
Open the publisher's page for the full experience