Submitted by WangKe 24 VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing LLMs for Reasoning 19 2