StepFun Launches StepAudio 3: Five Audio Models for Real-Time Voice and Speech Recognition

StepFun ·

Key Info

StepFun has launched StepAudio 3, a family of five audio models covering real-time voice, speech recognition, speech generation, audio generation, and music. The real-time model tops Artificial Analysis for conversational dynamics (98.9%) and speech reasoning (99.7%), while its ASR model reaches 1.7% WER.

Highlights

  • Real-time voice model ranks #1 on Artificial Analysis for both Conversational Dynamics and Speech Reasoning.
  • ASR achieves 1.7% word error rate, matching the best result on the leaderboard.
  • Enables voice agents that handle interruptions, reason while speaking, and call tools.
  • Also supports expressive voice generation, full audio scenes, and music creation.
Loading...