StepFun Launches StepAudio 3: Five Audio Models for Real-Time Voice and Speech Recognition
Key Info
StepFun has launched StepAudio 3, a family of five audio models covering real-time voice, speech recognition, speech generation, audio generation, and music. The real-time model tops Artificial Analysis for conversational dynamics (98.9%) and speech reasoning (99.7%), while its ASR model reaches 1.7% WER.
Highlights
- Real-time voice model ranks #1 on Artificial Analysis for both Conversational Dynamics and Speech Reasoning.
- ASR achieves 1.7% word error rate, matching the best result on the leaderboard.
- Enables voice agents that handle interruptions, reason while speaking, and call tools.
- Also supports expressive voice generation, full audio scenes, and music creation.