DuplexAct-Bench

Broadening Full-Duplex Speech Evaluation toward Proactive Interaction across Diverse Behavioral Requirements

Keyue Xing1,2,†, Wentao Ding1,†, Mengmeng Wang1, Wenming Tu1,3, Zilong Zheng1,*, Yipeng Kang1,*

1 State Key Laboratory of General Artificial Intelligence, BIGAI, China
2 Peking University, China    3 X-LANCE Lab, Shanghai Jiao Tong University, China

†Equal contribution. *Corresponding authors.

Abstract

Existing full-duplex speech benchmarks cover only subsets of real-time interaction behaviors, often under limited contextual conditions. We introduce DuplexAct-Bench, a bilingual benchmark that systematically covers six complementary behaviors, from interruption and yielding to proactive initiation, active silence, and backchanneling, across Pre-session, In-session, and No-explicit conditions. Across 1,290 English and Chinese streaming trials, we evaluate 12 full-duplex speech systems on both Timing and Content.

Results reveal substantial variation across behaviors, conditions, and systems, as well as frequent mismatches between semantic quality and behavioral timing. These findings show that current systems remain far from robustly managing when, whether, and how to participate as real-time interaction unfolds.

Six interaction behaviors shown as user and agent timelines
Figure 1. Representative examples of six interaction behaviors under different contextual conditions. Blue and red waveforms denote user and agent speech, respectively; peach shading indicates behavior-specific latency for Timing evaluation.
Taxonomy wheel of six behavior families and their scenarios
Figure 2. Behavior taxonomy and interaction scenarios in DuplexAct-Bench. The inner ring shows the six behavior families, while the outer ring shows their corresponding interaction scenarios. Numbers in the behavior segments indicate trial counts, totaling 1,290 trials. Red, blue, and black arcs denote Pre-session, In-session, and No-explicit conditions, respectively.

Benchmark Overview

Agent Interruption

Intervene during an ongoing user turn when the context calls for it.

240 trials

User Interruption

Yield the floor when the user interrupts, then respond to that interruption.

240 trials

Interruption Resistance

Maintain an ongoing activity despite non-disruptive user input.

280 trials

Active Silence

Withhold speech when silence is appropriate.

160 trials

Proactive Initiation

Start speaking without an explicit request from the user.

160 trials

Agent Backchannel

Give brief floor-preserving responses during the user’s turn.

210 trials

Evaluation Protocol

Content (0–5)

Semantic fulfillment, constraint adherence, contextual consistency, and prosodic appropriateness. Not applied to Active Silence.

Timing

Behavior-specific success criteria and temporal metrics evaluate whether the intended behavior occurs at an appropriate time.

Benchmark Results

Table 1. Evaluation coverage and Behavioral Correctness Rate (BCR, %)

Lang. denotes the evaluated languages (en: English; zh: Chinese). P / I / N denote Pre-session / In-session / No-explicit; IInt. and ISing. denote the In-session simultaneous-interpretation and collaborative-singing scenarios. Active Silence is scored by its silence criterion alone; User Interruption is evaluated on Valid trials only. “—” indicates no reported result, not zero. Bold marks the highest reported BCR in each column.

System Lang. Agent Interruption User Interruption Interruption Resistance Active Silence Proactive Initiation Agent Backchannel
PIN N PIInt.ISing.N PI PN PIN
Locally deployed systems
Freeze-Omnien/zh17.51.32.59.952.50.00.075.088.85.00.00.00.01.30.0
MiniCPM-o 4.5en/zh28.88.82.524.867.520.010.095.056.310.01.30.028.08.80.0
Moshien/zh0.00.00.026.12.50.00.062.57.57.535.00.0—12.513.8
PersonaPlexen/zh20.00.00.035.712.50.02.577.52.50.015.00.0—15.06.3
VITA-1.5en—1.36.38.6—17.50.022.5—0.0—0.0—25.00.0
Raon-SpeechChaten2.50.010.031.237.50.00.067.50.00.035.00.0—25.040.0
DuplexCascadeen—0.02.57.3—0.00.027.5—2.5—0.0—15.01.3
Systems accessed through remote APIs
Nemotron 3 VoiceChaten50.00.015.010.00.00.00.00.00.037.50.00.0—20.03.8
Qwen3.5-Omnien/zh18.816.31.339.880.03.812.596.30.046.30.08.830.00.07.5
Grok Voiceen/zh0.00.00.09.891.30.05.0100.010.00.00.00.04.00.00.0
GPT Realtime 2.1en/zh2.51.31.347.696.30.00.0100.038.83.80.00.030.07.510.0
Gemini 2.5 Native Audioen/zh8.81.32.542.461.30.030.071.376.380.00.00.08.00.00.0

Table 2. Content and Timing results

C denotes the mean Content score (0–5), L the mean latency in seconds, and BAS the Backchannel Alignment Score. For User Interruption, both metrics are computed on Valid trials only. Active Silence is omitted because it is evaluated solely by its silence criterion in Table 1. Arrows indicate the preferred direction; bold marks the best reported value in each metric column.

System Agent Interruption User Interruption Interruption Resistance Proactive Initiation Agent Backchannel
C ↑L (s) ↓ C ↑L (s) ↓ C ↑L (s) ↓ C ↑L (s) ↓ C ↑BAS ↑
Locally deployed systems
Freeze-Omni1.669.671.5621.222.6915.980.7014.031.090.005
MiniCPM-o 4.52.329.261.9719.153.1313.600.7013.981.590.054
Moshi1.1510.411.6817.761.4716.652.0111.011.240.108
PersonaPlex2.059.862.0915.291.7916.741.9612.781.040.097
VITA-1.51.709.601.2920.811.9119.540.4810.051.210.042
Raon-SpeechChat2.2610.201.9715.532.4017.201.9811.132.110.052
DuplexCascade1.349.831.2221.821.3823.090.4910.051.240.024
Systems accessed through remote APIs
Nemotron 3 VoiceChat2.619.121.3621.461.3116.470.5914.031.400.041
Qwen3.5-Omni3.719.422.8016.383.5515.030.8813.822.460.040
Grok Voice3.6310.181.7721.093.1115.620.6914.030.760.002
GPT Realtime 2.14.1210.082.7615.034.0515.910.6614.031.920.032
Gemini 2.5 Native Audio3.949.892.9814.673.6514.510.6614.030.980.003

Interruption Resistance varies sharply by scenario

100% vs. 20–30% BCR

Under No-explicit Interruption Resistance, Grok Voice and GPT Realtime 2.1 both reach 100%. The same behavior drops to 20.0% for simultaneous interpretation and 30.0% for collaborative singing.

Active Silence is highly condition-sensitive

88.8% → 5.0% BCR

Freeze-Omni’s Active Silence BCR is 88.8% in Pre-session trials and 5.0% in In-session trials. Gemini 2.5 Native Audio stays high on both, at 76.3% and 80.0%.

Strong Content does not guarantee joint success

4.12 Content, ≤2.5% BCR

GPT Realtime 2.1 has the highest Agent Interruption content score (4.12), but its BCR is 2.5%, 1.3%, and 1.3% across Pre-session, In-session, and No-explicit.

No-explicit Proactive Initiation remains challenging

Best BCR: 8.8%

Only Qwen3.5-Omni exceeds zero on No-explicit Proactive Initiation. VITA-1.5 and DuplexCascade have the lowest latency (10.05 s) but content scores of 0.48 and 0.49, and 0.0% BCR.

Audio Examples

Each box is one trial. The profile or instruction sits above the user’s turn, and the three system responses sit together underneath.