USC SAIL · NSTC Fellow

Huang-Cheng Chou

周惶振

Open to full-time roles

Speech / multimodal researcher for voice assistants— speech emotion recognition, subjective evaluation, and systems that hear affect as carefully as words.

Huang-Cheng Chou (周惶振)

About

I am a Postdoctoral Scholar– NSTC Fellow at the Signal Analysis and Interpretation Laboratory (SAIL) at the University of Southern California, working with Prof. Shrikanth S. Narayanan. I received my Ph.D. in Electrical Engineering from National Tsing Hua University, advised by Prof. Chi-Chun Lee.

My research centers on three threads: (1) speech emotion recognition under subjectivity and ambiguity— multi-label learning, annotator disagreement, calibration, fairness, and open evaluation via EMO-SUPERB and a survey of bias and fairness in speech AI; (2) conversational social signal processing—deception and belongingness in dialogue / small groups, and dyadic clinical affect such as depression-related speech; and (3) speech systems for assistants and multilingual / low-resource settings—unified ASR + SER, speech LLMs, expressive TTS evaluation, edge SER / KWS, and collaborations such as TaigiSpeech. At USC SAIL / SPAN, I also work on speech enhancement for real-time MRI (rtMRI).

I have worked on speech and ML systems from research through deployment at Amazon Alexa, ITRI, and RealTek. I am motivated by building intelligent, expressive agents that integrate speech, language, and emotion for natural interaction. I am currently seeking full-time research and industry positions.

Seeking
Full-time research / industry roles in speech AI, voice assistants, and affective computing
Current
Postdoctoral Scholar (NSTC Fellow), USC SAIL / SPAN, Oct 2025 – Present
Education
Ph.D., Electrical Engineering, NTHU, 2024; B.S., Electrical Engineering, NTHU
Scholar
722 citations; h-index 15; i10-index 23 (Google Scholar, 29 Aug 2026)
Contact
huangchengchou@gmail.com

News

  1. First-author preprint on USC SAIL SPAN rtMRI speech enhancement is online: Navigating Speech Enhancement for Real-Time MRI (arXiv; submitted to JASA; demo).
  2. Revising VoxEmo, a reproducibility toolkit for speech-LLM SER (mentored; arXiv), now under review at IEEE Transactions on Affective Computing.
  3. Four papers accepted at INTERSPEECH 2026: The False Resonance, The Binding Effect, Speaker Identity in Non-Verbal Vocalizations, and Layer-wise Task Vector Merging (plus TaigiSpeech as a long paper).
  4. Survey Toward Fair Speech Technologies is on arXiv (submitted to TMLR), with code and a Hugging Face collection.
  5. Our paper TaigiSpeech: A Low-Resource Real-World Speech Intent Dataset… (coauthor) is available on arXiv (INTERSPEECH 2026 long paper).
  6. Do You Hear What I Mean? (instruction–perception gap in expressive TTS) is accepted by ICASSP 2026.
  7. DeSTA2.5-Audio is published in IEEE TASLP 2026.
  8. Started as an NSTC Postdoctoral Fellow at USC SAIL with Prof. Shrikanth Narayanan.
  9. INTERSPEECH 2025 papers: Mitigating subgroup disparities and Meta-PerSER.
  10. First-author work on stimulus modality for SER at ICASSP 2025.
  11. Received the NSTC Postdoctoral Research Abroad Fellowship (2025–2026).
  12. Co-led EMO-SUPERB / Open-Emotion (IEEE SLT 2024); Minority Views Matter in IEEE TAFFC.
  13. Ph.D. dissertation on SER under subjectivity received the ACLCLP Doctoral Dissertation Award — Honorable Mention.

Projects

Grouped by research theme. Selected papers under each; complete list on the CV.

USC SAIL · SPAN

rtMRI speech enhancement for SPAN

Contribution to SAIL’s Speech Production and Articulation Knowledge (SPAN) effort on real-time MRI of dynamic vocal-tract shaping. First-author work evaluates modern speech enhancement on rtMRI speech—objective quality plus downstream tasks (ASR, speaker, SER, demographics)— and releases practical guidance with an open demo so SPAN corpora are more usable beyond speech science. Preprint on arXiv (submitted to JASA).

Theme

Conversational social signals

Conversational deception detection pipeline
Conversational deception · APSIPA ASC 2019

Affect and social meaning in dialogue: deception and perceived lying in interactive games, belongingness / satisfaction in small-group conversation, multimodal group emotion, and dyadic depression-related speech (including mentored work).

Theme

Fairness & robust SER

Framework linking fairness definitions, measurement, bias sources, and mitigation along the speech pipeline
Toward Fair Speech Technologies · arXiv 2026

Subgroup disparity, debiasing without fragile demographic shortcuts, and confidence-oriented augmentation for fairer multi-label emotion recognition—plus a survey that maps fairness definitions to evaluation and mitigation across speech AI.

Open evaluation

Benchmarks & toolkits

EMO-SUPERB community evaluation loop
EMO-SUPERB · IEEE SLT 2024

Reproducible evaluation contracts for classical and speech-LLM SER: splits, metrics, prompts, parsing, and diagnostics so results stay auditable.

Theme

Assistant speech systems

Tiny Whisper-SER multitask architecture
Tiny Whisper-SER · APSIPA ASC 2024

Unified ASR + multi-label SER for voice-assistant workloads, emotional-speech ASR analysis, stimulus-modality effects, and on-device SER / KWS under compute budgets.

  • A Tiny Whisper-SER — APSIPA ASC 2024
  • Amazon Alexa Applied Scientist Intern — shared-layer ASR + SER on Whisper
  • Stimulus modality matters — ICASSP 2025
  • RealTek internship — streaming edge SER / KWS compression

Theme

Speech LLMs, generation & low-resource

DeSTA2.5 self-generated alignment training pipeline
DeSTA2.5-Audio · IEEE TASLP 2026

Large audio / speech language models, expressive TTS evaluation, codec emotion preservation, and low-resource speech resources such as TaigiSpeech.

Selected publications

Full list on the CV and Google Scholar (722 citations; h-index 15; i10-index 23, 29 Aug 2026). Name in bold indicates Huang-Cheng Chou.

Selected first-author

  1. Navigating Speech Enhancement for Real-Time MRI: A Systematic Assessment of Signal Quality, Source Preservation, and Downstream Tasks Huang-Cheng Chou, Sean Foley, Haley Hsu, Kevin Huang, Szu-Jui Chen, Rong Chao, Louis Goldstein, Khalil Iskarous, Dani Byrd, Yu Tsao, Sudarsana Reddy Kadiri, John H. L. Hansen, Shrikanth Narayanan arXiv 2026 (submitted to JASA) arXiv · Demo
  2. Minority Views Matter: Evaluating Speech Emotion Classifiers with Human Subjective Annotations by an All-Inclusive Aggregation Rule Huang-Cheng Chou, Lucas Goncalves, Seong-Gyun Leem, Ali N. Salman, Chi-Chun Lee, Carlos Busso IEEE Transactions on Affective Computing, 2024 DOI
  3. Open-Emotion: A Reproducible EMO-Superb For Speech Emotion Recognition Systems Huang-Cheng Chou, Haibin Wu, Kai-Wei Chang, Lucas Goncalves, Jiawei Du, Jyh-Shing Roger Jang, Chi-Chun Lee, Hung-yi Lee IEEE SLT 2024 DOI · Project
  4. Embracing Ambiguity And Subjectivity Using The All-Inclusive Aggregation Rule For Evaluating Multi-Label Speech Emotion Recognition Systems Huang-Cheng Chou, Lucas Goncalves, Haibin Wu, Hung-yi Lee, Chi-Chun Lee IEEE SLT 2024 DOI
  5. A Tiny Whisper-SER: Unifying Automatic Speech Recognition and Multi-label Speech Emotion Recognition Tasks Huang-Cheng Chou APSIPA ASC 2024 DOI
  6. Stimulus Modality Matters: Impact of Perceptual Evaluations from Different Modalities on Speech Emotion Recognition System Performance Huang-Cheng Chou, Haibin Wu, Chi-Chun Lee ICASSP 2025 DOI
  7. The Importance of Calibration: Rethinking Confidence and Performance of Speech Multi-label Emotion Classifiers Huang-Cheng Chou, Lucas Goncalves, Seong-Gyun Leem, Chi-Chun Lee, Carlos Busso INTERSPEECH 2023 DOI
  8. Every Rating Matters: Joint Learning of Subjective Labels and Individual Annotators for Speech Emotion Classification Huang-Cheng Chou, Chi-Chun Lee ICASSP 2019 DOI

Selected collaborative

  1. Toward Fair Speech Technologies: A Comprehensive Survey of Bias and Fairness in Speech AI Yi-Cheng Lin, Yun-Shao Tsai, Kuan-Yu Chen, Hsiao-Ying Huang, Huang-Cheng Chou, Shrikanth Narayanan, Yu Tsao, Jian-Jiun Ding, Hung-yi Lee arXiv 2026 (submitted to TMLR) arXiv · Code · Hub
  2. DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model With Self-Generated Cross-Modal Alignment Ke-Han Lu, …, Huang-Cheng Chou, …, Hung-yi Lee IEEE TASLP 2026 DOI
  3. TaigiSpeech: A Low-Resource Real-World Speech Intent Dataset and Preliminary Results with Scalable Data Mining In-the-Wild Kai-Wei Chang, Yi-Cheng Lin, Huang-Cheng Chou, …, Hung-yi Lee INTERSPEECH 2026 (long), arXiv arXiv
  4. The False Resonance: A Critical Examination of Emotion Embedding Similarity for Speech Generation Evaluation Yun-Shao Tsai, Yi-Cheng Lin, Huang-Cheng Chou, Tzu-Wen Hsu, Yun-Man Hsu, Chun Wei Chen, Shrikanth Narayanan, Hung-yi Lee INTERSPEECH 2026, arXiv arXiv
  5. The Binding Effect: Analyzing How Multi-Dimensional Cues Form Gender Bias in Instruction TTS Kuan-Yu Chen, Yi-Cheng Lin, Po-Chung Hsieh, Huang-Cheng Chou, Chih-Fan Hsu, Jeng-Lin Li, Hung-yi Lee, Jian-Jiun Ding INTERSPEECH 2026, arXiv arXiv
  6. Layer-wise Task Vector Merging: Leveraging ASR and SER Task Vectors for Enhanced Speech Emotion Representation Chia-Yu Lee, Huang-Cheng Chou (mentored), Tzu-Quan Lin, Yuanchao Li, Ya-Tse Wu, Shrikanth Narayanan, Chi-Chun Lee INTERSPEECH 2026, arXiv arXiv
  7. Speaker Identity in Non-Verbal Vocalizations: Conditional Distillation and Mixture of Experts Approach Tzu-Chieh Wei, Yi-Cheng Lin, Huang-Cheng Chou (mentored), Kuan-Yu Chen, Hsin-Yen Sung, Shrikanth Narayanan, Hung-yi Lee INTERSPEECH 2026, arXiv arXiv
  8. Do You Hear What I Mean? Quantifying the Instruction-Perception Gap in Instruction-Guided Expressive Text-to-Speech Systems Yi-Cheng Lin, Huang-Cheng Chou, Tzu-Chieh Wei, Kuan-Yu Chen, Hung-yi Lee ICASSP 2026 DOI
  9. VoxEmo: A Reproducible Toolkit for Speech Emotion Recognition with Speech LLMs Hezhao Zhang, Huang-Cheng Chou (mentored), Shrikanth Narayanan, Thomas Hain IEEE Transactions on Affective Computing (under review) arXiv
See the complete CV →

Experience

Selected awards

  • NSTC Postdoctoral Research Abroad Fellowship — 2025–2026
  • ACLCLP Doctoral Dissertation Award — Honorable Mention (2024)
  • Merry Electronics Electroacoustics Thesis Award — Silver (2025), Bronze (2021)
  • APSIPA ASC Best Regular Paper Award — 2019 (conversational deception detection)
  • NOVATEK Ph.D. Excellence Scholarship — 2022–2023
  • IEEE SPS ICASSP Travel Grant — 2025
  • IEEE SLT Travel Grant — 2024
  • Google East Asia Student Travel Grants — ICASSP 2022 & INTERSPEECH 2022
  • ISCA INTERSPEECH Grant — 2022
  • AAAC ACII Student Travel Grant — 2017
  • ACLCLP Outstanding Students Conference Travel Grant — 2019, 2022, 2024, 2025
  • FAOS Outstanding Students Conference Travel Grant — 2019, 2022, 2023

Full awards, reviewing, mentoring, and publication list are on the CV page.

Contact

Currently seeking full-time roles. Open to conversations on speech AI, voice assistants, and research collaboration.

huangchengchou@gmail.com