> For the complete documentation index, see [llms.txt](https://gzbin.gitbook.io/map/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://gzbin.gitbook.io/map/algorithm/voice.md).

# Audio

语音 音频 Voice speech Sound

autotune

[**基于CNN的歌声合成算法论文解读**](https://cloud.tencent.com/developer/article/1776840) [**NEUTRINO**](https://n3utrino.work)

[**linux下python文字转语音库pyttsx3用法**](https://www.bilibili.com/video/av67394684/)

Python实现语音合成 [zminoooooo](https://www.bilibili.com/video/BV1gu411R7L1) 使用百度api 35:00 开始

Python文字语音播报 [**书童**](https://xugaoxiang.com/2021/04/08/python-tts-chinese/)

LumenVox - Exceptional Voice Experiences [u](https://www.youtube.com/channel/UCe3Gq1lw94FjwDmy6aGVn_A/videos)

|                                                                                                                                                                                                                                                                                                                                                                                     |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Deep Learning for Human Language Processing (2020,Spring) [Hung-yi Lee](https://www.youtube.com/playlist?list=PLJV_el3uVTsO07RpBYFsXg-bN5Lu0nhdG)                                                                                                                                                                                                                                   |
|                                                                                                                                                                                                                                                                                                                                                                                     |
| Understanding Sound & Speakers [Branch Education](https://www.youtube.com/playlist?list=PL6rx9p3tbsMvYWeYUTNuRNMLXn0eyJvub)                                                                                                                                                                                                                                                         |
| Simple Voice Recorder with GUI in Python [NeuralNine](https://www.youtube.com/watch?v=u_xNvC9PpHA)                                                                                                                                                                                                                                                                                  |
| 深度学习-语音识别实战(Python) [网易云课堂](https://study.163.com/course/courseMain.htm?courseId=1210361865)                                                                                                                                                                                                                                                                                        |
| Valerio Velardo - The Sound of AI([u](https://www.youtube.com/c/ValerioVelardoTheSoundofAI/featured), [medium](https://medium.com/the-sound-of-ai), [linkedin](https://www.linkedin.com/in/valeriovelardo), [fb](https://www.facebook.com/TheSoundOfAI), [s](https://valeriovelardo.com), )                                                                                         |
| Deep Learning (for Audio) with Python([ulist](https://www.youtube.com/playlist?list=PL-wATfeyAMNrtbkCNsLcpoAyBBRJZVlnf), )                                                                                                                                                                                                                                                          |
| speech recognition [pixeldev](https://www.youtube.com/playlist?list=PLsaLbFPkNd55uUF-MpQ7RtldorYYR_CeT)                                                                                                                                                                                                                                                                             |
| [讯飞听见](https://www.iflyrec.com/)                                                                                                                                                                                                                                                                                                                                                    |
| Audacity [s](https://www.audacityteam.org/) [v](https://www.youtube.com/watch?v=P30suV1UdSY) [doc](https://manual.audacityteam.org) shop 音频处理软件 [v](https://www.youtube.com/playlist?list=PLlKpQrBME6xKm9iJlVHWbJd_xAvtAQy6W) [v](https://www.youtube.com/playlist?list=PLMoVDOzX4VcQGOySLj7TO3ab0g1eIzJYq)                                                                         |
| aeneas s [git](https://github.com/readbeyond/aeneas) [doc](https://www.readbeyond.it/aeneas/) 用于自动同步音频和文本（也称为强制对齐）的工具                                                                                                                                                                                                                                                               |
| 音频处理\|adobe audition cc 2020\|零基础教程 [长安老张](https://www.youtube.com/playlist?list=PLnIffqrKafFuF7r4aej24Zg2X0drpEc5e)                                                                                                                                                                                                                                                                |
| 音频编辑处理\|Adobe Audition CC 2019\|零基础教程 [长安老张](https://www.youtube.com/playlist?list=PLnIffqrKafFszIyQmLnHDhfpVExbtoYGp)                                                                                                                                                                                                                                                              |
| 文字转语音--让机器语音像人类一样自然真实 [长安老张](https://www.youtube.com/watch?v=SdhLawZEr1I)                                                                                                                                                                                                                                                                                                           |
| Desmos Shenanigans [Eric Tao](https://www.youtube.com/playlist?list=PLlYltssWVRe9vOW7BmghfExRg2jPMqiQH)                                                                                                                                                                                                                                                                             |
| Simple Voice Recorder in Python [NeuralNine](https://www.youtube.com/watch?v=av8E8qLZswU) pyaudio wave                                                                                                                                                                                                                                                                              |
| Guitars AI [u](https://www.youtube.com/c/GuitarsAI/playlists) [git](https://github.com/GuitarsAI/)                                                                                                                                                                                                                                                                                  |
| Multirate Signal Processing with Python [Guitars AI](https://www.youtube.com/playlist?list=PL6QnpHKwdPYiQmq8PcGvx-EpLFrVmYMOW) [git](https://github.com/GuitarsAI/MRSP_Notebooks)                                                                                                                                                                                                   |
| Python for Beginner: Basics of Digital Audio Processing and Machine Learning using Python [Guitars AI](https://www.youtube.com/playlist?list=PL6QnpHKwdPYi-600wCa4PPIrP2DpgKqFk)                                                                                                                                                                                                    |
| <p>Machine Learning for Audio Signals in Python - Full Course - Ilmenau University of Technology <a href="https://www.youtube.com/watch?v=9NDG-b6VMik">Guitars AI</a></p><p>Machine Learning for Audio Signals in Python <a href="https://www.youtube.com/playlist?list=PL6QnpHKwdPYjfCH2zkMGEHu2kv1HTICYA">Guitars AI</a> <a href="https://github.com/GuitarsAI/MLfAS">git</a></p> |
| Advanced Digital Signal Processing using Python [Guitars AI](https://www.youtube.com/playlist?list=PL6QnpHKwdPYhIt-zvYMSYXzLpQ0MMRVTK) [git](https://github.com/GuitarsAI/ADSP_Tutorials)                                                                                                                                                                                           |
| speechbrain/[speechbrain](https://github.com/speechbrain/speechbrain) A PyTorch-based Speech Toolkit [s](https://speechbrain.github.io/) [s](https://www.zhihu.com/question/266242493)                                                                                                                                                                                              |
| 从语音识别到语音交互 陈伟 [青云QingCloud](https://www.youtube.com/watch?v=AUKycdq6jig)                                                                                                                                                                                                                                                                                                            |
| Talk \| 字节跳动研究员董倩倩：基于单调切分的端到端同传 [将门-TechBeat技术社区](https://www.youtube.com/watch?v=uv_4sRg47hk)                                                                                                                                                                                                                                                                                      |
| Talk \| 中国科学技术大学和微软亚洲研究院联合培养博士生冷燚冲：语音识别的快速纠错模型FastCorrect [将门-TechBeat技术社区](https://www.youtube.com/watch?v=uYaeOM6_4Go)                                                                                                                                                                                                                                                            |
| How to do entity detection on audio files with Python [AssemblyAI](https://www.youtube.com/watch?v=7jrs3DrMVJQ)                                                                                                                                                                                                                                                                     |
| Audio Signal Processing for Machine Learning and Deep Learning [Prabhjot Gosal](https://www.youtube.com/playlist?list=PLvELbYeZ7GEHBaoCyZS63FhI0bbhnwZbB)                                                                                                                                                                                                                           |
| oscilloscope [search](https://www.douyin.com/search/oscilloscope)                                                                                                                                                                                                                                                                                                                   |
| Audio Spectrum Analyzer [Mark Jay](https://www.youtube.com/playlist?list=PLX-LrBk6h3wQVsrldsQdtKmeTygurKiuS)                                                                                                                                                                                                                                                                        |
| Python - Read WAV file SMN [DigiTech](https://www.youtube.com/watch?v=ZBS31QxVvSg)                                                                                                                                                                                                                                                                                                  |
| Audio processing in Python with Feature Extraction for machine learning [Prodramp](https://www.youtube.com/watch?v=vbhlEMcb7RQ)                                                                                                                                                                                                                                                     |
| Extract Features from Audio File \| MFCC \| Python [Hackers Realm](https://www.youtube.com/watch?v=hX2sOvrWC1Q\&list=PL_8jNcohs27XTz6sGUZaQjXULA1J-Vun7\&index=124)                                                                                                                                                                                                                 |
| Basic Sound Processing in Python \| SciPy 2015 \| Allen Downey [Enthought](https://www.youtube.com/watch?v=0ALKGR0I5MA)                                                                                                                                                                                                                                                             |
| Extract Musical Notes from Audio in Python with FFT [Jeff Heaton](https://www.youtube.com/watch?v=rj9NOiFLxWA)                                                                                                                                                                                                                                                                      |
| Music Production with FL Studio – Full Tutorial for Beginners [freeCodeCamp](https://www.youtube.com/watch?v=BUjdnxgBgzM)                                                                                                                                                                                                                                                           |
| Sorting Visualizer with Sound (JavaScript Tutorial) [Radu Mariescu-Istodor](https://www.youtube.com/watch?v=_AwSlHlpFuc\&t=27s)                                                                                                                                                                                                                                                     |
| #219 Generating sounds 🎵 from the PICO! But is it music to my 👂 ears? [Ralph S Bacon](https://www.youtube.com/watch?v=9H6zeZcrZ7M)                                                                                                                                                                                                                                                |
| AI sound bites series for media pickup [Dr Alan D. Thompson](https://www.youtube.com/playlist?list=PLqJbCeNOfEK-81jVeTcG9MHLfgeFkLlVm)                                                                                                                                                                                                                                              |
| 文字转语音、音频转文字软件！双向转换，完全免费开源！支持 Windows、macOS、Linux \| [零度解说](https://www.youtube.com/watch?v=MP_tjvFq3m0)                                                                                                                                                                                                                                                                             |
| 注意看，这些全是 AI 配音。 [Topbook](https://www.youtube.com/watch?v=yUuS-oTDSD4)                                                                                                                                                                                                                                                                                                              |
| 中国传媒大学\_\_播音主持普通话语音 [知识资源世界(KnowledgeWorld)](https://www.youtube.com/playlist?list=PLoEWjLHPG6Z4J3SZWxrDH-xpythKiXyim)                                                                                                                                                                                                                                                              |
| <p>Audio Slicer 音频切片机 一个简约的 GUI 应用程序，通过静音检测对音频进行切片。</p><p>flutydeer/<a href="https://github.com/flutydeer/audio-slicer">audio-slicer</a> <a href="https://www.youtube.com/watch?v=YAviCCA-Fic">v</a></p>                                                                                                                                                                            |
| Transcribe Audio Files with OpenAI Whisper [NeuralNine](https://www.youtube.com/watch?v=UAdX0cGuC28)                                                                                                                                                                                                                                                                                |
| Convert Videos To MP3 in Python [NeuralNine](https://www.youtube.com/watch?v=ucXTQ0V8qMA)                                                                                                                                                                                                                                                                                           |
|                                                                                                                                                                                                                                                                                                                                                                                     |
|                                                                                                                                                                                                                                                                                                                                                                                     |
|                                                                                                                                                                                                                                                                                                                                                                                     |

***

## Audio Classify

|                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Build a Deep Audio Classifier with Python and Tensorflow [Nicholas Renotte](https://www.youtube.com/watch?v=ZLIPkmmDJAc)                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                             |
| MIT 6.S191: Automatic Speech Recognition [Alexander Amini](https://www.youtube.com/watch?v=sR6_bZ6VkAg)                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                              |
| <p>Complete Deep Learning</p><p>Krish Naik</p><p><a href="https://www.youtube.com/watch?v=mHPpCXqQd7Y&#x26;list=PLZoTAELRMXVPGU70ZGsckrMdr0FteeRUi">Part 1</a>-EDA-Audio Classification Project Using Deep Learning</p><p><a href="https://www.youtube.com/watch?v=4F-cwOkMdTE&#x26;list=PLZoTAELRMXVPGU70ZGsckrMdr0FteeRUi&#x26;index=73">Part 2</a>-Data Preprocessing-Audio Classification Project Using Deep Learning</p><p><a href="https://www.youtube.com/watch?v=uTFU7qThylE&#x26;list=PLZoTAELRMXVPGU70ZGsckrMdr0FteeRUi&#x26;index=74">Part 3</a>-Model Creation-Audio Classification Project Using Deep Learning</p><p><a href="https://www.youtube.com/watch?v=cqndT517NcQ&#x26;list=PLZoTAELRMXVPGU70ZGsckrMdr0FteeRUi&#x26;index=75">Part 4</a>-Testing ANN Model-Audio Classification Project Using Deep Learning</p> |
|                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
|                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |
|                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                                      |

## 语音识别, ASR, Automatic Speech Recognition

**Speech to Text /** Speech Recognition **/** Voice Recognition / Speech Detection 语音转文字

|                                                                                                                                                                                                                                                                                                               |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| vosk                                                                                                                                                                                                                                                                                                          |
| paddle speech                                                                                                                                                                                                                                                                                                 |
| SpeechRecognition [pypi](https://pypi.org/project/SpeechRecognition/) [git](https://github.com/Uberi/speech_recognition) [v](https://www.youtube.com/watch?v=31DZfkYRvI4)                                                                                                                                     |
| PyAudio [**s**](http://people.csail.mit.edu/hubert/pyaudio/#downloads)                                                                                                                                                                                                                                        |
| Speech Recognition with Convolutional Neural Networks in Keras/TensorFlow [Weights & Biases](https://www.youtube.com/watch?v=Qf4YJcHXtcY)                                                                                                                                                                     |
| Automatic Speech Recognition - An Overview [Microsoft Research](https://www.youtube.com/watch?v=q67z7PTGRi8) [s](https://www.microsoft.com/en-us/research/video/automatic-speech-recognition-overview/)                                                                                                       |
| I Built a Personal Speech Recognition System for my AI Assistant [The A.I. Hacker - Michael Phi](https://www.youtube.com/watch?v=YereI6Gn3bM)                                                                                                                                                                 |
| A Basic Introduction to Speech Recognition (Hidden Markov Model & Neural Networks) [Hannes van Lier](https://www.youtube.com/watch?v=U0XtE4_QLXI)                                                                                                                                                             |
| Voice Recognition As Fast As Possible [Techquickie](https://www.youtube.com/watch?v=AWWjN1QqoYY)                                                                                                                                                                                                              |
| Phoneme Detection with CNN-RNN-CTC Loss Function - Machine Learning [Ali Yektaie](https://www.youtube.com/watch?v=x1IAPgvKUmM)                                                                                                                                                                                |
| how speech recognition works in under 4 minutes. [Sam Stork](https://www.youtube.com/watch?v=iNbOOgXjnzE)                                                                                                                                                                                                     |
| Automatic Speech Recognition: Chapter 1 [NEWTON Technologies](https://www.youtube.com/watch?v=vlr_3kNQZAk)                                                                                                                                                                                                    |
| <p>Lecture 9 - Speech Recognition (ASR) \[Andrew Senior] <a href="https://www.youtube.com/watch?v=HyUtT_z-cms">Zafarullah Mahmood</a></p><p>Deep Learning for NLP at Oxford with Deep Mind 2017 <a href="https://www.youtube.com/playlist?list=PL613dYIGMXoZBtZhbyiBqb0QtgK6oJbpm">Zafarullah Mahmood</a></p> |
| Python Speech Recognition Project Series [AssemblyAI](https://www.youtube.com/playlist?list=PLcWfeUsAys2nb0i79L_LqYVfwOWEYA4eD)                                                                                                                                                                               |
| <p>Speech Recognition in Python Tutorial – Full Course for Beginners <a href="https://www.youtube.com/watch?v=mYUyaKmvu6Y">freeCodeCamp</a></p><p>Python Speech Recognition Tutorial – Full Course for Beginners</p>                                                                                          |
| Speech Recognition in Python [NeuralNine](https://www.youtube.com/watch?v=9GJ6XeB-vMg)                                                                                                                                                                                                                        |
| Build your own real-time voice command recognition model with TensorFlow [AssemblyAI](https://www.youtube.com/watch?v=m-JzldXm9bQ)                                                                                                                                                                            |
| nobody132/[masr](https://github.com/nobody132/masr) 中文语音识别; Mandarin Automatic Speech Recognition;                                                                                                                                                                                                            |
| Build a Twitter Hate Speech Detection Model Using Machine Learning \| Python \| Project For Beginners [AI Sciences](https://www.youtube.com/watch?v=AEWty_y_WxE)                                                                                                                                              |
| Hate Speech Detection on Twitter Using Sentiment Analysis and Machine Learning Algorithms. [Vikas Mayurkumar Patel](https://www.youtube.com/watch?v=oJ_nX5jsfIY) colab                                                                                                                                        |
| JSALT 2020 Workshop Closing Ceremonies: Speech Recognition and Diarization for Unsegmented Multi-talker Recordings Team Presentation [Center for Language and Speech (CLSP) @ JHU](https://www.youtube.com/playlist?list=PLSeS0sl8xpTz705M-KK7s6VT19a99SsFT)                                                  |
| <p>Conformer-1: a new large scale/robust speech recognition model <a href="https://www.youtube.com/watch?v=hkChdbq7IQI">AssemblyAI</a></p><p>Conformer-2: A state-of-the-art speech recognition model <a href="https://www.youtube.com/watch?v=r-CEc_ZYV9E">AssemblyAI</a></p>                                |
| 【Whisper】免費開源語音辨識自動上字幕　字幕正確率比剪映還高！！！｜下載完後無須聯網　僅需使用自己電腦處理｜如何在Windows上使用Whisper [The walking fish 步行魚](https://www.youtube.com/watch?v=kFtrvdriLU8)                                                                                                                                                             |
| 【Code Gym】Python 文字轉語音教學，用短短幾行程式，做出會讀文字說話的自動化服務 [Code Gym](https://www.youtube.com/watch?v=0PuslZHJQes)                                                                                                                                                                                                       |
| 教你写一款简单的语音识别软件 [程序马](https://www.youtube.com/watch?v=4A3NAYgz-R8)                                                                                                                                                                                                                                             |
| 语音识别小程序 识别中英文语音转文字 附python code源码 [人工智障机器人](https://www.youtube.com/watch?v=dPiZG0yM5H4)                                                                                                                                                                                                                      |
| 【免費！極速！上字幕！】5分鐘完成8小時的工作量！ [木石mushi](https://www.youtube.com/watch?v=8tjbHfkO9bU\&list=PLi3JGBubTe7YtBHGuc5Gj9NfwIsWqnG4o) 剪映                                                                                                                                                                                  |
| 阿里达摩院paraformer 中文语音识别 [v](https://www.youtube.com/watch?v=8gd_WtBPxaw)                                                                                                                                                                                                                                       |
| Speech recognition in Python made easy \| Python Tutorial [AssemblyAI](https://www.youtube.com/watch?v=YdYTSxEW5bA)                                                                                                                                                                                           |
|                                                                                                                                                                                                                                                                                                               |
|                                                                                                                                                                                                                                                                                                               |

| pyTranscriber                                                                                                                              |                                                                                                               |   |
| ------------------------------------------------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------- | - |
| raryelcostasouza/[pyTranscriber](https://github.com/raryelcostasouza/pyTranscriber) generate automatic transcription / automatic subtitles | [069ba4a500](https://github.com/raryelcostasouza/pyTranscriber/tree/069ba4a500f9e37df2bfb31ea1a4627fbc82ca78) |   |
|                                                                                                                                            |                                                                                                               |   |
|                                                                                                                                            |                                                                                                               |   |

### **NVIDIA JARVIS**

|                                                                                                                                           |
| ----------------------------------------------------------------------------------------------------------------------------------------- |
| **NVIDIA JARVIS(**[**s**](https://developer.nvidia.com/nvidia-jarvis?ncid=partn-38053#cid=dl20_partn_en-us)**, )**                        |
| **NVIDIA Jarvis Conversational AI on Python - Dr. Ahmad Bazzi(**[**v**](https://www.youtube.com/watch?v=sbYolIax190)**, )**               |
| **Conversational AI w/ Jarvis - checking out the API(**[**sentdex**](https://www.youtube.com/watch?v=fQzjgaKSrkc)**, )**                  |
| **1. Live coding Jarvis Transcriptions for Speech to Text Dataset p.1(**[**sentdex**](https://www.youtube.com/watch?v=ubvgReZVf5g)**, )** |
| **2. Live coding Jarvis Transcriptions for Speech to Text Dataset p.2(**[**sentdex**](https://www.youtube.com/watch?v=BDl6fzhp2Ao)**, )** |

|                                                                                                                                                         |
| ------------------------------------------------------------------------------------------------------------------------------------------------------- |
| How to Convert Speech To Text In Python \| Speech Recognition Library \| 10 Lines Of Code [Know-How](https://www.youtube.com/watch?v=ifey3EnhB-g)       |
| How To Convert Text To Speech In Python \| pyttsx3 Library \| Basics Of Artificial Intelligence [Know-How](https://www.youtube.com/watch?v=6Za0ztPMr8g) |

## 语音合成 SpeechSynthesis, Audio Generation

Voice Synthesis

|                                                                                                                             |
| --------------------------------------------------------------------------------------------------------------------------- |
| AI Voice Synthesis [NanoNomad](https://www.youtube.com/playlist?list=PLn3gJVMORUsTiGp5Jc37XID9jS9SHKUV1)                    |
| ai实时变音，你的声卡能否一战！#人工智能 #aigc #有ai就有无限可能 #ai [AICK-KC](https://www.douyin.com/video/7249742859941776651)                      |
| How To Create AI Voice Over [How To In 5 Minutes](https://www.youtube.com/playlist?list=PLIi08al-Zd_kixb8cTdoOWsQvTTrNh0aR) |
|                                                                                                                             |

## **TTS Text to Speech**

|                                                                                                                                                                                                                      |
| -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| pdf to voice [v](https://www.douyin.com/video/7150605317326245151)                                                                                                                                                   |
| 魔音工坊 [s](https://www.moyin.com/)                                                                                                                                                                                     |
| 逗哥 云浩宇 解说模式                                                                                                                                                                                                          |
| VoiceMaker [s](https://voicemaker.in/) [v](https://www.youtube.com/watch?v=X-DacDX3W8M)                                                                                                                              |
| NaturalReaders [s](https://www.naturalreaders.com/online/) 无中文 英文发音好                                                                                                                                                 |
| narakeet [s](https://www.narakeet.com/app/text-to-audio) [s](https://www.narakeet.com/languages/chinese-text-to-speech/) Zihan Daoming                                                                               |
| play.ht [s](https://play.ht/text-to-speech-voices/chinese/) 下载付费 带情绪 Zhiyu Yunxi Xiaoyou Xiaoxuan Xiaoxiao                                                                                                           |
| IBM Watson Text to Speech [s](https://www.ibm.com/cloud/watson-text-to-speech)                                                                                                                                       |
| chineseedge [s](https://chineseedge.com/chinese-text-to-speech/) 单一 高亮                                                                                                                                               |
| imtranslator [s](https://text-to-speech.imtranslator.net/) 一般                                                                                                                                                        |
| Kukarella [v](https://www.youtube.com/watch?v=a6bVPAz2L3s) [s](https://kukarella.com/)                                                                                                                               |
| 百度 api paddlepaddle/[PaddleSpeech](https://gitee.com/paddlepaddle/PaddleSpeech)                                                                                                                                      |
| Balabolka [s](http://balabolka.site/cn/balabolka.htm) [v](https://www.youtube.com/watch?v=kfqpFKdDVMU) TTS程序 api调用 不支持linux                                                                                          |
| ranchlai/[mandarin-tts](https://github.com/ranchlai/mandarin-tts) 中文 (普通话) 语音 合成                                                                                                                                     |
| espeak [s](http://espeak.sourceforge.net/index.html)                                                                                                                                                                 |
| ubuntu完美安装espeak支持中文和粤语 [csdn](https://blog.csdn.net/qq_24406903/article/details/89811732) caixxiong/[espeak-data](https://github.com/caixxiong/espeak-data)                                                         |
| clipchamp [s](https://clipchamp.com/en/) [v](https://www.youtube.com/watch?v=TQDbtP5SB00) 语音需要从视频中提取                                                                                                                 |
| 【全球10大机器配音软件】有声书影视解说博主利器\|全球10大语音合成软件 软件侠何二                                                                                                                                                                          |
| 一款采用人工智能黑科技文字转流畅自然配音软件配音神器 [廉飞](https://www.youtube.com/watch?v=jd1WOkBIOTQ)                                                                                                                                         |
| 听书、看书、搜书（小说），一个视频教你全搞定！从此你就是大神！几千万书源，AI智能朗读 [嘴哥干货](https://www.youtube.com/watch?v=jI3yo0fXusw)                                                                                                                      |
| Harmonai, Dance Diffusion and The Audio Generation Revolution [Weights & Biases](https://www.youtube.com/watch?v=KmB8z2CYjZY)                                                                                        |
| [**语音合成-腾讯云**](https://cloud.tencent.com/developer/tag/10464)                                                                                                                                                        |
| How to make text to speech from your own voice with python [TheLazyBusyGamer](https://www.youtube.com/watch?v=h-M1kqL4C7c)                                                                                           |
| 我花了15小時，就為找到這3款AI語音生成工具｜AI朗读｜AI虚拟主播 [指令必达](https://www.youtube.com/watch?v=RsnU0Pjt7lo)                                                                                                                              |
| 【Anny讲Python】Python听书系统之语音合成技术 [Python学习指南](https://www.youtube.com/watch?v=A4qZzbFldyk) pytts 百度api                                                                                                                 |
|                                                                                                                                                                                                                      |
| <p>suno-ai/<a href="https://github.com/suno-ai/bark">bark</a></p><p>文字轉語音，堪比人聲！100%相似，完全找不到AI的痕跡！Bing Chat將擁有長期記憶😱 ChatGPT更新！ <a href="https://www.youtube.com/watch?v=9gp4dpH7maA">你的攝影大哥 Davidddodle</a> bark</p> |
| 2 Build Your Own Audiobooks App ! python projects 2023 ! Master Python by Building Real World Python [Self Study](https://www.youtube.com/watch?v=WBRL0gcZ1cg) pyttsx3                                               |
| 最强TTS（文本转语音）模型Bark发布 - 支持带有情感的语音，歌曲生成 -体验声音克隆功能 [小薇 Official Channel](https://www.youtube.com/watch?v=lz9YPR1vq60) huggingface space                                                                                 |
| 文本转语音的逼真程度再次突破天花板！Bark横空出世 The realism of text-to-speech conversion technology reached new heights! [工具狂Toolbuddy](https://www.youtube.com/watch?v=nUKXo_iLnZY)                                                      |
| AI工具 免费可商用的AI克隆声音 语音合成 文字转语音 手把手教学 bark #bark #AI语音 #声音克隆 #AI克隆 #创业工具 [AB视界](https://www.youtube.com/watch?v=LJaWU5yPJao)                                                                                            |
|                                                                                                                                                                                                                      |
|                                                                                                                                                                                                                      |

|                                                                                                                                                                                                                                                                                                                                                            |
| ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Ekho [s](https://www.eguidedog.net/ekho.php) [git](https://github.com/hgneng/ekho) 支持linux [为Ekho添加新的声音](http://eguidedog.net/doc/doc_make_new_voice_cn.php) [linuxInstall](https://www.eguidedog.net/doc/doc_install_ekho.php) [2](https://blog.csdn.net/cceking/article/details/51760732) [编译安装](https://blog.csdn.net/AMDS123/article/details/73825409) |
| TTS技术简单介绍和Ekho（余音）TTS的安装与编程 [csdn](https://blog.csdn.net/zouxy09/article/details/7909154)                                                                                                                                                                                                                                                                  |
| Cloud Text-to-Speech Google Cloud                                                                                                                                                                                                                                                                                                                          |
| gTTS (Google Text-to-Speech) [doc](https://gtts.readthedocs.io/en/latest/) [git](https://github.com/pndurette/gTTS)                                                                                                                                                                                                                                        |
| 自定义声音模型 📱文字转语音--先从克隆自己的声音开始\|20分钟合成自己的声音 [长安老张](https://www.youtube.com/watch?v=X3opZ6-tb-o)                                                                                                                                                                                                                                                              |
| How to Generate Your Own Voice - Text to Speech [Website Learners](https://www.youtube.com/watch?v=PzBdEW5WNGo)                                                                                                                                                                                                                                            |
| Can we CLONE my voice using ML? [Nicholas Renotte](https://www.youtube.com/watch?v=qVD1NBLoOJw)                                                                                                                                                                                                                                                            |
| I secretly used AI to clone my favourite Rapper [Nicholas Renotte](https://www.youtube.com/watch?v=HoMZxHNpM84)                                                                                                                                                                                                                                            |
| [**Real-Time-Voice-Cloning**](https://github.com/CorentinJ/Real-Time-Voice-Cloning)                                                                                                                                                                                                                                                                        |
| NVIDIA’s Amazing AI Clones Your Voice! 🤐 [Two Minute Papers](https://www.youtube.com/watch?v=k54cpsAbMn4)                                                                                                                                                                                                                                                 |
| Microsoft’s New AI Clones Your Voice In 3 Seconds! [Two Minute Papers](https://www.youtube.com/watch?v=F6HSsVIkqIU)                                                                                                                                                                                                                                        |
| My Vocal.AI s [v](https://www.youtube.com/watch?v=lz9YPR1vq60)                                                                                                                                                                                                                                                                                             |
| AI工具 五秒声音克隆 mockingbird 支持中文 [AB视界](https://www.youtube.com/watch?v=eHRHfYeoZ3g)                                                                                                                                                                                                                                                                           |
| 只需5秒！就能克隆你的声音，MockingBird 轻松实现 AI 文字转语音 ！附最新安装、使用教程！\| [零度解说](https://www.youtube.com/watch?v=dxIvkEB0XE0)                                                                                                                                                                                                                                                 |
| 我把自己做成了AI声库，效果太离谱了!! 给朋友打“诈骗电话”会被发现吗？ ｜ [LKs](https://www.youtube.com/watch?v=ezt3dQac3w4)                                                                                                                                                                                                                                                                 |
| AI声音克隆 [AI工坊#十个骑士](https://www.youtube.com/playlist?list=PL4HQtsL6idnZbBWq18aO5uXxgRd7M8zUh)                                                                                                                                                                                                                                                               |
|                                                                                                                                                                                                                                                                                                                                                            |
|                                                                                                                                                                                                                                                                                                                                                            |

|                                                                                                                                                                                                                                                                                                                                                             |
| ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Azure 文档 [microsoft](https://docs.microsoft.com/zh-cn/azure/?product=popular)                                                                                                                                                                                                                                                                               |
| azure 语音服务文档 [microsoft](https://docs.microsoft.com/zh-cn/azure/cognitive-services/speech-service/) 语音服务[定价](https://azure.microsoft.com/zh-cn/pricing/details/cognitive-services/speech-services/) [试用](https://azure.microsoft.com/zh-cn/services/cognitive-services/text-to-speech/#features) 试用[精简版](https://speech.microsoft.com/audiocontentcreation) |
| azure 文本转语音 [microsoft](https://azure.microsoft.com/zh-cn/services/cognitive-services/text-to-speech/#overview)                                                                                                                                                                                                                                             |
| Azure-Samples/[cognitive-services-speech-sdk](https://github.com/Azure-Samples/cognitive-services-speech-sdk) [Python](https://github.com/Azure-Samples/cognitive-services-speech-sdk/tree/master/quickstart/python/text-to-speech)                                                                                                                         |
| Build a natural custom voice for your brand [microsoft](https://techcommunity.microsoft.com/t5/ai-cognitive-services-blog/build-a-natural-custom-voice-for-your-brand/ba-p/2112777) Neural TTS [u](https://www.youtube.com/channel/UCxZ5iCrcXbpmdhZIKFKgMLA/videos)                                                                                         |
| Microsoft Azure Open Course (Mandarin) By MVP张诚 [Cheng Zhang](https://www.youtube.com/playlist?list=PLZmpc0o_yCMkHVSVdNSi4lfgzf40UyeiR)                                                                                                                                                                                                                     |
| articles with label Neural TTS [microsoft](https://techcommunity.microsoft.com/t5/ai-cognitive-services-blog/bg-p/CognitiveServicesBlog/label-name/Neural%20TTS)                                                                                                                                                                                            |
| articles with label Speech [microsoft](https://techcommunity.microsoft.com/t5/ai-cognitive-services-blog/bg-p/CognitiveServicesBlog/label-name/Speech)                                                                                                                                                                                                      |
| Artificial Intelligence and Machine Learning [microsoft](https://techcommunity.microsoft.com/t5/artificial-intelligence-and/ct-p/AI)                                                                                                                                                                                                                        |
| articles with label Cognitive Services [microsoft](https://techcommunity.microsoft.com/t5/ai-cognitive-services-blog/bg-p/CognitiveServicesBlog/label-name/Cognitive%20Services)                                                                                                                                                                            |
| Neural Text to Speech extends support to 15 more languages with state-of-the-art AI quality By Qinying Liao [microsoft](https://techcommunity.microsoft.com/t5/ai-cognitive-services-blog/neural-text-to-speech-extends-support-to-15-more-languages-with/ba-p/1505911)                                                                                     |
| Neural Speech Synthesis with Transformer Network [arxiv](https://arxiv.org/abs/1809.08895)                                                                                                                                                                                                                                                                  |
| FastSpeech: Fast, Robust and Controllable Text to Speech [arxiv](https://arxiv.org/abs/1905.09263)                                                                                                                                                                                                                                                          |
| 微软AI配音，免费的文字转音频在线工具 [便利空间](https://www.youtube.com/watch?v=I7TRD9mjW1c)                                                                                                                                                                                                                                                                                     |
|                                                                                                                                                                                                                                                                                                                                                             |
|                                                                                                                                                                                                                                                                                                                                                             |

## voice 2 voice

### AI音色替换技术(Sovits4.0)

|                                                                                                                                                               |     |     |                                                                                                                                                                                                  |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------- | --- | --- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| voice 2 voice                                                                                                                                                 |     |     |                                                                                                                                                                                                  |
| AI孙燕姿 [u](https://www.youtube.com/@AI-nx9lp/videos) 华语AI翻唱 [u](https://www.youtube.com/@AImusicCover4u)                                                       |     |     |                                                                                                                                                                                                  |
| 【AI翻唱】“超详细”！Sovits4.0手把手详细教学！ [StarPorridge小粥](https://www.youtube.com/watch?v=VTWTQ5l9eHI)                                                                   |     |     |                                                                                                                                                                                                  |
| 从孙燕姿的“漠河舞厅”谈生成式AI创作 [Jeff科技视角](https://www.youtube.com/watch?v=x3DwYWqlY5k) Jarret Charvis                                                                    |     |     |                                                                                                                                                                                                  |
| AI孙燕姿，再次席卷乐坛。版权律师们，愁白了头。 [老范讲故事](https://www.youtube.com/watch?v=322Al2O6Dz8)                                                                                 |     |     |                                                                                                                                                                                                  |
| 面向小白的低成本的VITS模型训练教程及注意事项 [欧皇张大千](https://www.youtube.com/watch?v=hnxPPDw5fWE) colab                                                                           |     |     |                                                                                                                                                                                                  |
| Voice Cloning Tutorial with Coqui TTS and Google Colab \| Fine Tune Your Own VITS Model for Free [NanoNomad](https://www.youtube.com/watch?v=6QAGk_rHipE)     |     |     |                                                                                                                                                                                                  |
| 一键完成Ai声音模拟和音色替换，无需训练模型 – 快速生成AI孙燕姿歌曲 [小薇 Official Channel](https://www.youtube.com/watch?v=YAviCCA-Fic)                                                       |     |     |                                                                                                                                                                                                  |
| 我们做了个能对话的AI派蒙，免费给大家玩！ [极客湾Geekerwan](https://www.youtube.com/watch?v=8gd_WtBPxaw)                                                                             |     |     |                                                                                                                                                                                                  |
| <p>变分推断 学习声线 VITS:</p><p>Variational Inference with adversarial learning for end-to-end Text-to-Speech</p>                                                    |     |     |                                                                                                                                                                                                  |
| Conditional Variational Autoencoder with Adversarial Learning for End-to-End Text-to-Speech                                                                   |     |     |                                                                                                                                                                                                  |
| 【警告】涉及敏感內容！未公開的教學流出！隨時會刪片... #ai孫燕姿 #tuneflow #svs [咖哩哥](https://www.youtube.com/watch?v=EEtQg9zJ0_g)                                                         |     |     |                                                                                                                                                                                                  |
| TuneFlow [s](https://tuneflow.com/)                                                                                                                           |     |     |                                                                                                                                                                                                  |
| svc-develop-team/[so-vits-svc](https://github.com/svc-develop-team/so-vits-svc)                                                                               |     |     |                                                                                                                                                                                                  |
| 【超火的 AI翻唱】克隆自己的声音！全网最详细的使用教程， So-VITS-SVC 4.1 声音复制神器，完全免费！！ \| [零度解说](https://www.youtube.com/watch?v=o6mKekgJDqw)                                            |     |     |                                                                                                                                                                                                  |
| 如何克隆自己的声音？如何克隆别人的声音？免安装1分钟克隆声音 关键还是免费的，支持中文英文日语韩语、AI文字转语音克隆 [峰哥探世界](https://www.youtube.com/watch?v=cli0ZfTxlJA)                                              |     |     |                                                                                                                                                                                                  |
| 【GPT-SoVITS 本地整合包0306】声音克隆新纪元！1分钟完美声音克隆，完美复刻任何语音、语调、语气！效果超神！ [十个骑士](https://www.youtube.com/watch?v=J9FwHp9egg0)                                              |     |     |                                                                                                                                                                                                  |
| 只需一小时，完美克隆你自己！最新中文AI数字人制作教程, 手机可搞定的AI数字人分身, 中文领域秒杀Heygen数字人 clone yourself perfectly within one hour [靜電的AI設計教室](https://www.youtube.com/watch?v=vBhEuQO2QFc) |     |     |                                                                                                                                                                                                  |
| <p>GPT-SoVITS语音克隆AI，只需一分钟素材训练模型，效果堪比商用。一键安装，附Colab脚本                                                                                                          | TTS | RVC | GPT-SoVITS Colab <a href="https://www.youtube.com/watch?v=BDC2aJJFSgE">AI探索与发现</a></p><p>声音克隆 <a href="https://www.youtube.com/playlist?list=PLZoPEoiia3kyEaE7aitCDGnwJEIJ9Knye">AI探索与发现</a></p> |
|                                                                                                                                                               |     |     |                                                                                                                                                                                                  |
|                                                                                                                                                               |     |     |                                                                                                                                                                                                  |
|                                                                                                                                                               |     |     |                                                                                                                                                                                                  |

### Transcribed speech（转录语音）

|                                                                                                                               |
| ----------------------------------------------------------------------------------------------------------------------------- |
| Transcribed speech 是将口语转录为书面语的过程，也可以理解为将已经存在的音频或视频转换为文本的过程。                                                                   |
| LeMUR: a framework for applying powerful LLMs to transcribed speech [AssemblyAI](https://www.youtube.com/watch?v=q2RShntLoW0) |
|                                                                                                                               |

## 音乐生成, Text to Music Generation

|                                                                                                                                                                     |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| MusicLM is a GAMECHANGER for ML Text to Music Generation [Nicholas Renotte](https://www.youtube.com/watch?v=F4WEGvFAwv8)                                            |
| Noise2music、MusicLM：文本生成音乐效果如何？ [跟李沐学AI](https://www.youtube.com/watch?v=XsE9WWP9awo)                                                                               |
| "壞了壞了，有了AI音樂人都沒飯吃了！“ [Devin Wang](https://www.youtube.com/watch?v=q8I4NCvmQJ8)                                                                                      |
| [AIVA](https://www.aiva.ai/) - The AI composing emotional soundtrack music [v](https://www.youtube.com/watch?v=8wHNIZWIOqg)                                         |
| Google's MusicLM: Text Generated Music & It's Absurdly Good [bycloud](https://www.youtube.com/watch?v=2CUKU2iAzAs)                                                  |
| Music Creation [John Tan Chong Min](https://www.youtube.com/playlist?list=PLcORGO1bTgjWqsb4sEqI9SFUUvLV7HNfh)                                                       |
| 资深音频人怎么看谷歌最新作曲AI，MusicLM（上集）. How does a composer look at Goolge Music AI MusicLM，Episode I? [Van Zeng \[Audio Man\]](https://www.youtube.com/watch?v=N_s0Tc_Tfa0)  |
| 资深音频人怎么看谷歌最新作曲AI（下集）读评论，How does a composer look at Goolge Music AI，Episode II, Read Comments [Van Zeng \[Audio Man\]](https://www.youtube.com/watch?v=JBqurD1KG_w) |
| 【[Code Gym](https://www.youtube.com/watch?v=0PuslZHJQes)】Python 文字轉語音教學，用短短幾行程式，做出會讀文字說話的自動化服務                                                                      |
| 借助ChatGPT，乐盲也能轻松驾驭 Google MusicLM，Midjourney来助阵，让音乐与图像完美结合 \| [回到Axton](https://www.youtube.com/watch?v=s3Lm_uff4PA)                                                |
| Advanced Music Production with FL Studio – Tutorial [freeCodeCamp](https://www.youtube.com/watch?v=I_ShMaNw0Rc)                                                     |
|                                                                                                                                                                     |

### text to song

|                                                                                                            |
| ---------------------------------------------------------------------------------------------------------- |
| Voicemod声音模拟 [s](https://tuna.voicemod.net/text-to-song/) [v](https://www.youtube.com/watch?v=YAviCCA-Fic) |
|                                                                                                            |
|                                                                                                            |

### X Studio 3

|                                                                                                                                                                                      |
| ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| X Studio 3 \| 小冰公司AI合成歌声软件 [最佳拍档](https://www.youtube.com/watch?v=GuuDAv0FwxI) [s](https://singer.xiaoice.com/) [doc](https://cpt17i0bkj.feishu.cn/docx/TZOsdaI5xoCunvxptO2ctQXInTg) |
|                                                                                                                                                                                      |
|                                                                                                                                                                                      |

## 人声分离

|                                                                                                                    |
| ------------------------------------------------------------------------------------------------------------------ |
| algoriddim [s](https://www.algoriddim.com/djay-pro-mac)                                                            |
| vocal remover [v](https://www.youtube.com/watch?v=YAviCCA-Fic) [s](https://www.media.io/online-vocal-remover.html) |
| Ultimate Vocal Remover [s](https://ultimatevocalremover.com/) [v](https://www.youtube.com/watch?v=YAviCCA-Fic)     |
| 除了提取伴奏和人声、还可以提取乐器的免费开源工具! \| Ultimate Vocal Remover \| UVR5 [忙嘞个芒](https://www.youtube.com/watch?v=W7nmoM9OIfE)    |
|                                                                                                                    |
|                                                                                                                    |

## Voice Assistant

|                                                                                                                                                       |
| ----------------------------------------------------------------------------------------------------------------------------------------------------- |
| Python Voice Assistant [Tech With Tim](https://www.youtube.com/playlist?list=PLzMcBGfZo4-mBungzp4GO4fswxO8wTEFx)                                      |
| [search\_query](https://www.youtube.com/results?search_query=Python+Voice+Assistant)=Python+Voice+Assistant                                           |
| Azure-Samples/[Cognitive-Services-Voice-Assistant](https://github.com/Azure-Samples/Cognitive-Services-Voice-Assistant)                               |
| Assistente Vocale Python - Riconoscimento Vocale con Speech Recognition e Pyttsx3 [PitoneProgrammatore](https://www.youtube.com/watch?v=kXiIK1qahsg)  |
| Build an AI Voice Assistant with PyTorch [The A.I. Hacker - Michael Phi](https://www.youtube.com/playlist?list=PL5rWfvZIL-NpFXM9nFr15RmEEh4F4ePZW)    |
| Build A Python Speech Assistant App [Traversy Media](https://www.youtube.com/watch?v=x8xjj6cR9Nc)                                                     |
| Intelligent Voice Assistant in Python [NeuralNine](https://www.youtube.com/watch?v=SXsyLdKkKX0)                                                       |
| Voice Assistant with Wake Word in Python [NeuralNine](https://www.youtube.com/watch?v=F62wb_jfUUw\&list=PL7yh-TELLS1G9mmnBN3ZSY8hYgJ5kBOg-\&index=21) |

<details>

<summary>Open Jarvis 大數軟體有限公司</summary>

\[[Open Jarvis](https://www.youtube.com/watch?v=31DZfkYRvI4)] 如何讓Python 自動將語音轉譯成文字?

\[[Open Jarvis](https://www.youtube.com/watch?v=xd_1rn89W2k)] 如何用Python 讓電腦說話?

\[[Open Jarvis](https://www.youtube.com/watch?v=Y2t68jDwfhc)] 如何用不到30行Python程式碼寫出「真‧對話機器人」?

\[[Open Jarvis](https://www.youtube.com/watch?v=T5UIySP9Owc)] 如何讓對話機器人利用 Wikipedia 回答專業知識?

\[[Open Jarvis](https://www.youtube.com/watch?v=9lpoYnWFXjQ)] 如何使用Python寫一個翻譯蒟蒻?

</details>
