docs+i18n: add Japanese, Spanish and Arabic localizations (#740)

- Add translated READMEs: docs/README_ja.md, docs/README_es.md,
  docs/README_ar.md (structure, code blocks and tables kept identical
  to the English README)
- Add WebUI locales: ja_JP.json, es_ES.json, ar_SA.json (110 keys
  each, covering every i18n key used in webui.py)
- Fill 3 missing keys in en_US.json (删除/情感向量/语言)
- Add language switcher links to the README headers

Co-authored-by: nanaoto <10inspiral@gmail.com>
This commit is contained in:
nanaoto
2026-08-11 16:35:05 +08:00
committed by GitHub
parent 9fced08672
commit 506e7ddfe8
7 changed files with 1941 additions and 1 deletions
+525
View File
@@ -0,0 +1,525 @@
<div align="center">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="../assets/indextts_icon_dark.png"/>
<img src="../assets/indextts_icon_light.png" width="300"/>
</picture>
**نظام صناعي المستوى لتحويل النص إلى كلام بتقنية الاستدلال الصفري (Zero-Shot)، قابل للتحكم وعالي الكفاءة**
[简体中文](README_zh.md) | [English](../README.md) | [日本語](README_ja.md) | [Español](README_es.md) | العربية
[![GitHub Stars](https://img.shields.io/github/stars/index-tts/index-tts?style=flat&logo=github)](https://github.com/index-tts/index-tts/stargazers)
[![arXiv](https://img.shields.io/badge/arXiv-2601.03888-b31b1b?logo=arxiv)](https://arxiv.org/abs/2601.03888)
[![Discord](https://img.shields.io/badge/Discord-join-5865F2?logo=discord&logoColor=white)](https://discord.gg/uT32E7KDmy)
</div>
IndexTTS هو نظام لتحويل النص إلى كلام بتقنية الاستدلال الصفري يستنسخ الصوت من
مقطع صوتي مرجعي واحد. يدعم أحدث إصدار، **IndexTTS-2.5**، اللغات الصينية
والإنجليزية واليابانية والإسبانية والعربية، مع تحكم دقيق في المشاعر، وتحكم
في سرعة الكلام، وتحكم في النطق (البينيين / أصوات CMU / الكانا اليابانية)،
واستدلال أسرع من IndexTTS-2.
---
## 🗂️ مجموعة النماذج
| النموذج | العروض | الورقة البحثية | ModelScope | HuggingFace |
| :--- | :---: | :---: | :---: | :---: |
| **IndexTTS-2.5** | [![Demo](https://img.shields.io/badge/Demo-Page-orange?logo=github)](https://index-tts.github.io/index-tts2-5.github.io/) | [![Paper](https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv)](https://arxiv.org/abs/2601.03888) | [![ModelScope](https://img.shields.io/badge/ModelScope-Model-purple?logo=modelscope)](https://modelscope.cn/models/IndexTeam/IndexTTS-2.5) | [![HuggingFace](https://img.shields.io/badge/HuggingFace-Model-blue?logo=huggingface)](https://huggingface.co/IndexTeam/IndexTTS-2.5) |
| **IndexTTS-2** | [![Demo](https://img.shields.io/badge/Demo-Page-orange?logo=github)](https://index-tts.github.io/index-tts2.github.io/) | [![Paper](https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv)](https://arxiv.org/abs/2506.21619) | [![ModelScope](https://img.shields.io/badge/ModelScope-Model-purple?logo=modelscope)](https://modelscope.cn/models/IndexTeam/IndexTTS-2) | [![HuggingFace](https://img.shields.io/badge/HuggingFace-Model-blue?logo=huggingface)](https://huggingface.co/IndexTeam/IndexTTS-2) |
| **IndexTTS-1.5** | [![Demo](https://img.shields.io/badge/Demo-Page-orange?logo=github)](https://index-tts.github.io/) | [![Paper](https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv)](https://arxiv.org/abs/2502.05512) | [![ModelScope](https://img.shields.io/badge/ModelScope-Model-purple?logo=modelscope)](https://modelscope.cn/models/IndexTeam/IndexTTS-1.5) | [![HuggingFace](https://img.shields.io/badge/HuggingFace-Model-blue?logo=huggingface)](https://huggingface.co/IndexTeam/IndexTTS-1.5) |
| **IndexTTS** | [![Demo](https://img.shields.io/badge/Demo-Page-orange?logo=github)](https://index-tts.github.io/) | [![Paper](https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv)](https://arxiv.org/abs/2502.05512) | [![ModelScope](https://img.shields.io/badge/ModelScope-Model-purple?logo=modelscope)](https://modelscope.cn/models/IndexTeam/Index-TTS) | [![HuggingFace](https://img.shields.io/badge/HuggingFace-Model-blue?logo=huggingface)](https://huggingface.co/IndexTeam/Index-TTS) |
## 📣 الأخبار
- `2026/08/10` 🔥 نطلق **IndexTTS-2.5**
- يدعم الآن الصينية والإنجليزية واليابانية والإسبانية والعربية، مع استدلال أسرع من IndexTTS-2، مع الحفاظ على قدرات التخليق عبر اللغات وفصل نبرة الصوت عن المشاعر.
- تحسين قابلية التحكم في البينيين الصيني وأصوات CMU الإنجليزية والكانا اليابانية.
- تحكم في سرعة الكلام عبر `duration_factor` (مدة من 0.5x إلى 2.0x).
- `2025/09/08` 🔥 نطلق **IndexTTS-2**
- أول نموذج TTS ذاتي الارتداد مع تحكم دقيق في مدة التخليق، يدعم الوضعين القابل للتحكم وغير القابل للتحكم. <i>هذه الميزة غير مفعّلة بعد في هذا الإصدار.</i>
- تخليق كلام عالي التعبير العاطفي، مع التحكم في المشاعر عبر وسائط إدخال متعددة.
- `2025/05/14` 🔥 نطلق **IndexTTS-1.5**، مع تحسين كبير في استقرار النموذج وأدائه في اللغة الإنجليزية.
- `2025/03/25` 🔥 نطلق **IndexTTS-1.0** مع أوزان النموذج وكود الاستدلال.
- `2025/02/12` 🎉 قدّمنا ورقتنا البحثية إلى arXiv، وأصدرنا العروض التوضيحية ومجموعات الاختبار.
## 🎬 العروض التوضيحية
<div align="center">
**IndexTTS-2.5: مستقبل الصوت، يُولَّد الآن**
[![IndexTTS2.5 Demo](../assets/index2.5_video_cover.png)](https://www.bilibili.com/video/BV1uvMk6ZEdK/)
**IndexTTS-2: مستقبل الصوت، يُولَّد الآن**
[![IndexTTS2 Demo](../assets/IndexTTS2-video-pic.png)](https://www.bilibili.com/video/BV136a9zqEk5)
</div>
## 🚀 البدء
### 1. المتطلبات الأساسية
تأكد من تثبيت [git](https://git-scm.com/downloads)، ثم نزّل هذا المستودع:
```bash
git clone https://github.com/index-tts/index-tts.git && cd index-tts
```
يتم تنزيل الملفات الصوتية النموذجية عند الطلب من HuggingFace/ModelScope عند
التشغيل الأول، لذا لم يعد Git LFS مطلوبًا.
### 2. تثبيت التبعيات
نستخدم [uv](https://docs.astral.sh/uv/getting-started/installation/) لإدارة
بيئة تبعيات المشروع. وهو **مطلوب** لتثبيت موثوق:
```bash
pip install -U uv # or see the link above for other install methods
```
```bash
uv sync --all-extras
```
يؤدي هذا تلقائيًا إلى إنشاء دليل المشروع `.venv` وتثبيت الإصدارات الصحيحة
من Python وجميع التبعيات المطلوبة.
إذا كان التنزيل بطيئًا، استخدم مرآة محلية، مثل إحدى هاتين المرآتين في الصين:
```bash
uv sync --all-extras --default-index "https://mirrors.aliyun.com/pypi/simple"
uv sync --all-extras --default-index "https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple"
```
> [!TIP]
> **الميزات الإضافية المتاحة:**
>
> - `--all-extras`: يضيف تلقائيًا *كل* ميزة إضافية مدرجة أدناه. يمكنك إزالة
> هذا الخيار إذا أردت تخصيص خيارات التثبيت.
> - `--extra webui`: يضيف دعم واجهة الويب WebUI (موصى به).
> - `--extra deepspeed`: يضيف دعم DeepSpeed (قد يسرّع الاستدلال على بعض
> الأنظمة).
> [!IMPORTANT]
> **Windows:** قد يكون تثبيت DeepSpeed صعبًا. يمكنك تخطيه بإزالة خيار
> `--all-extras` وإضافة خيارات الميزات الأخرى يدويًا.
>
> **Linux/Windows:** إذا ظهر خطأ CUDA أثناء التثبيت، تأكد من تثبيت الإصدار
> **12.8** (أو أحدث) من [CUDA Toolkit](https://developer.nvidia.com/cuda-toolkit)
> من NVIDIA على نظامك.
### 3. تنزيل النماذج
نزّل النماذج المطلوبة عبر [uv tool](https://docs.astral.sh/uv/guides/tools/#installing-tools):
عبر `huggingface-cli`:
```bash
uv tool install "huggingface-hub[cli,hf_xet]"
# IndexTTS-2.5
hf download IndexTeam/IndexTTS-2.5 --local-dir=checkpoints
# IndexTTS-2
hf download IndexTeam/IndexTTS-2 --local-dir=checkpoints_2
```
أو عبر `modelscope`:
```bash
uv tool install "modelscope"
# IndexTTS-2.5
modelscope download --model IndexTeam/IndexTTS-2.5 --local_dir checkpoints
# IndexTTS-2
modelscope download --model IndexTeam/IndexTTS-2 --local_dir checkpoints_2
```
> [!IMPORTANT]
> إذا لم تكن الأوامر أعلاه متاحة، اقرأ بعناية مخرجات `uv tool` —
> فهي ستخبرك بكيفية إضافة الأدوات إلى PATH في نظامك.
> [!NOTE]
> يتم تنزيل بعض النماذج الصغيرة تلقائيًا عند التشغيل الأول. إذا كانت شبكتك
> بطيئة في الوصول إلى HuggingFace، اضبط مرآة قبل تشغيل الكود:
>
> ```bash
> export HF_ENDPOINT="https://hf-mirror.com"
> ```
### 4. التحقق من تسريع GPU
لتشخيص بيئتك ومعرفة وحدات GPU المكتشفة، استخدم الأداة المرفقة:
```bash
uv run tools/gpu_check.py
```
## 💻 الاستخدام
### 🌐 عرض الويب
```bash
# IndexTTS-2.5 (default)
uv run webui.py --version 2.5 --model_dir ./checkpoints
# IndexTTS-2
uv run webui.py --version 2 --model_dir ./checkpoints_2
```
افتح متصفحك وزر `http://127.0.0.1:7860` لمشاهدة العرض.
يمكنك ضبط الإعدادات لتفعيل استدلال BF16 (IndexTTS-2.5) / FP16 (IndexTTS-2)
(استهلاك أقل لذاكرة VRAM)، وتسريع DeepSpeed، ونوى CUDA المجمّعة للسرعة،
وما إلى ذلك. يمكن الاطلاع على جميع الخيارات المتاحة عبر:
```bash
uv run webui.py -h
```
> [!IMPORTANT]
> استدلال **FP16/BF16** (نصف الدقة) أسرع ويستهلك ذاكرة VRAM أقل، مع فقدان
> ضئيل جدًا في الجودة.
>
> **DeepSpeed** *قد* يسرّع الاستدلال على بعض الأنظمة، لكنه قد يجعله أبطأ
> أيضًا — يعتمد ذلك على عتادك وتعريفاتك ونظام التشغيل. جرّب الطريقتين.
>
> جميع أوامر `uv` **تفعّل تلقائيًا** البيئة الافتراضية الصحيحة الخاصة
> بالمشروع. *لا* تفعّل أي بيئة يدويًا قبل تشغيل أوامر `uv`، لأن ذلك قد
> يسبب تعارضات في التبعيات.
### 🚀 التشغيل باستخدام vLLM
للنشر الإنتاجي، راجع [وصفة vLLM لـ IndexTTS](https://github.com/vllm-project/recipes/pull/772).
### 📝 واجهة Python البرمجية
لتشغيل السكربتات، استخدم `uv run <file.py>` ليعمل الكود داخل بيئة `uv`.
قد تحتاج أيضًا إلى إضافة الدليل الحالي إلى `PYTHONPATH`:
```bash
# IndexTTS2
PYTHONPATH="$PYTHONPATH:." uv run indextts/infer_v2.py
# IndexTTS2.5
PYTHONPATH="$PYTHONPATH:." uv run indextts/infer_v2_5.py \
--cfg_path checkpoints/config.yaml \
--model_dir checkpoints \
--text "Hello world" \
--lang EN
```
#### 0. تهيئة IndexTTS
```python
# IndexTTS2
from indextts.infer_v2 import IndexTTS2
tts = IndexTTS2(cfg_path="checkpoints_2/config.yaml", model_dir="checkpoints_2", use_fp16=False, use_cuda_kernel=False, use_deepspeed=False)
# IndexTTS2.5
from indextts.infer_v2_5 import IndexTTS2
tts = IndexTTS2(cfg_path="checkpoints/config.yaml", model_dir="checkpoints", use_bf16=True)
```
#### 1. استنساخ الصوت من مقطع صوتي مرجعي واحد
```python
text = "Translate for me, what is a surprise!"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, output_path="gen.wav", verbose=True)
# IndexTTS2.5 (multilingual, with language selection)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="EN", output_path="gen.wav", verbose=True)
```
#### 2. التحكم في المشاعر بصوت مرجعي عاطفي منفصل
```python
text = "酒楼丧尽天良,开始借机竞拍房间,哎,一群蠢货。"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", emo_audio_prompt="examples/emo_sad.wav", verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, lang="ZH", output_path="gen.wav", emo_audio_prompt="examples/emo_sad.wav", verbose=True)
```
#### 3. ضبط شدة المشاعر باستخدام `emo_alpha`
عند تحديد صوت مرجعي عاطفي، يضبط `emo_alpha` مقدار تأثيره على المخرجات.
النطاق الصالح: `0.0 - 1.0`، الافتراضي: `1.0` (100%).
```python
text = "酒楼丧尽天良,开始借机竞拍房间,哎,一群蠢货。"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", emo_audio_prompt="examples/emo_sad.wav", emo_alpha=0.9, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", lang="ZH", emo_audio_prompt="examples/emo_sad.wav", emo_alpha=0.9, verbose=True)
```
#### 4. التحكم في المشاعر باستخدام متجه مشاعر
يمكنك حذف الصوت المرجعي العاطفي وتقديم بدلًا منه قائمة من 8 أرقام عشرية
تحدد شدة كل شعور، بالترتيب
`[happy, angry, sad, afraid, disgusted, melancholic, surprised, calm]`.
استخدم `use_random` لإدخال العشوائية أثناء الاستدلال (الافتراضي: `False`).
> [!NOTE]
> تفعيل أخذ العينات العشوائي يقلل من دقة استنساخ الصوت.
```python
text = "对不起嘛!我的记性真的不太好,但是和你在一起的事情,我都会努力记住的~"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/09.wav', text=text, output_path="gen.wav", emo_vector=[0, 0, 0.8, 0, 0, 0, 0, 0], use_random=False, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/09.wav', text=text, lang="ZH", output_path="gen.wav", emo_vector=[0, 0, 0.8, 0, 0, 0, 0, 0], use_random=False, verbose=True)
```
#### 5. التحكم في المشاعر من النص نفسه (`use_emo_text`)
فعّل `use_emo_text` لتحويل نص `text` تلقائيًا إلى متجهات مشاعر. يُوصى
بقيمة `emo_alpha` حوالي 0.6 (أو أقل) للحصول على كلام أكثر طبيعية. يمكن
إدخال العشوائية باستخدام `use_random` (الافتراضي: `False`).
```python
text = "快躲起来!是他要来了!他要来抓我们了!"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, use_random=False, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, lang="ZH", output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, use_random=False, verbose=True)
```
#### 6. التحكم في المشاعر بوصف عاطفي صريح (`emo_text`)
قدّم وصفًا نصيًا محددًا للمشاعر عبر `emo_text`، والذي يتم تحويله إلى
متجهات مشاعر — مما يمنحك تحكمًا منفصلًا في النص ووصف المشاعر:
```python
text = "快躲起来!是他要来了!他要来抓我们了!"
emo_text = "你吓死我了!你是鬼吗?"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, emo_text=emo_text, use_random=False, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, lang="ZH", output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, emo_text=emo_text, use_random=False, verbose=True)
```
#### 7. التحكم في سرعة الكلام (`duration_factor`)
القيمة الأكبر من `1.0` تبطئ الكلام، والقيمة الأصغر من `1.0` تسرّعه.
الافتراضي: `1.0` (السرعة العادية). النطاق الصالح: `0.5 - 2.0`.
```python
text = "大家好,欢迎来到IndexTTS的语速控制演示。"
# IndexTTS2.5
# Slow down (1.2x duration)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="ZH", output_path="gen_slow.wav", duration_factor=1.2, verbose=True)
# Speed up (0.8x duration)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="ZH", output_path="gen_fast.wav", duration_factor=0.8, verbose=True)
```
### 🗣️ التحكم في النطق
**IndexTTS2.5 — البينيين / أصوات CMU / الكانا اليابانية:**
يدعم IndexTTS2.5 هذه الاستبدالات النصية مع قدرة أفضل على اتباع التعليمات.
للقائمة الكاملة بالإدخالات الصالحة، راجع `checkpoints/pinyin.vocab` للبينيين
و[قاموس CMU](https://svn.code.sf.net/p/cmusphinx/code/trunk/cmudict/cmudict-0.7b)
للأصوات الإنجليزية.
```
他在银<行|XING2>里<行|HANG2>走了半天,发现这笔业务办不<行|HANG2>。
He had a <minute|M IH1 . N AH0 T> to examine the <minute|M AY0 . N UW1 T> details of the contract.
彼は料理が<上手|じょうず>だが、囲碁では<上手|うわて>に負けた。
```
**IndexTTS2 — البينيين:**
يدعم IndexTTS2 النمذجة المختلطة للأحرف الصينية والبينيين. لتفعيل التحكم
في البينيين، قدّم نصًا مع تعليقات بينيين محددة. لاحظ أن التحكم في البينيين
لا يعمل مع كل تركيبة صامت–صائت ممكنة؛ فقط حالات البينيين الصيني الصالحة
مدعومة (راجع `checkpoints/pinyin.vocab`).
```
之前你做DE5很好,所以这一次也DEI3做DE2很好才XING2,如果这次目标完成得不错的话,我们就直接打DI1去银行取钱。
```
### 🕰️ IndexTTS-1.5 (الإصدار القديم)
يمكنك أيضًا استخدام نموذج IndexTTS1 السابق باستيراد وحدة مختلفة:
```python
from indextts.infer import IndexTTS
tts = IndexTTS(model_dir="checkpoints", cfg_path="checkpoints/config.yaml")
voice = "examples/voice_07.wav"
text = "大家好,我现在正在bilibili 体验 ai 科技,说实话,来之前我绝对想不到!AI技术已经发展到这样匪夷所思的地步了!比如说,现在正在说话的其实是B站为我现场复刻的数字分身,简直就是平行宇宙的另一个我了。如果大家也想体验更多深入的AIGC功能,可以访问 bilibili studio,相信我,你们也会吃惊的。"
tts.infer(voice, text, 'gen.wav')
```
لمزيد من التفاصيل، راجع [README_INDEXTTS_1_5](../archive/README_INDEXTTS_1_5.md
أو زر مستودع IndexTTS1 على [index-tts:v1.5.0](https://github.com/index-tts/index-tts/tree/v1.5.0).
## 📊 التقييم
**الجدول 1: تحويل النص إلى كلام بتقنية الاستدلال الصفري على CV3-Eval** (تستخدم العربية مجموعة اختبار داخلية). †مقتبس من الورقة الأصلية.
<table>
<thead>
<tr>
<th rowspan="2">النموذج</th>
<th rowspan="2">المعاملات</th>
<th colspan="2">zh</th>
<th colspan="2">en</th>
<th colspan="2">es</th>
<th colspan="2">ja</th>
<th colspan="2">ar</th>
<th colspan="2">Avg</th>
</tr>
<tr>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
</tr>
</thead>
<tbody>
<tr><td>VoxCPM2</td><td>2B</td><td>3.88</td><td>74.99</td><td>5.13</td><td>71.57</td><td>5.49</td><td>74.67</td><td>6.69</td><td>72.90</td><td>14.94</td><td>65.99</td><td>7.22</td><td>72.02</td></tr>
<tr><td>OmniVoice</td><td>0.8B</td><td>3.41</td><td>72.99</td><td>3.62</td><td>70.13</td><td>3.52</td><td>74.14</td><td>5.38</td><td>70.49</td><td>17.88</td><td>64.22</td><td>6.76</td><td>70.39</td></tr>
<tr><td>Moss-TTS 1.5</td><td>8B</td><td>4.02</td><td>72.68</td><td>4.45</td><td>67.46</td><td>3.83</td><td>71.75</td><td>10.97</td><td>68.71</td><td>23.71</td><td>62.21</td><td>9.40</td><td>68.56</td></tr>
<tr><td>CosyVoice3-0.5B</td><td>0.5B</td><td>3.84</td><td>80.01</td><td>4.88</td><td>74.16</td><td>4.04</td><td>78.85</td><td>-</td><td>76.36</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>CosyVoice3-1.5B</td><td>1.5B</td><td>3.91†</td><td>-</td><td>4.99†</td><td>-</td><td>4.47†</td><td>-</td><td>7.57†</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>FireRedTTS-2</td><td>1.5B</td><td>8.22</td><td>68.10</td><td>14.92</td><td>56.93</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>Fish Audio S2 Pro</td><td>4B</td><td>3.62</td><td>67.79</td><td>3.83</td><td>61.66</td><td>2.93</td><td>67.44</td><td>5.15</td><td>66.15</td><td>14.15</td><td>59.43</td><td>5.94</td><td>64.49</td></tr>
<tr><td>Qwen3-TTS</td><td>1.7B</td><td>3.27</td><td>73.02</td><td>5.06</td><td>67.17</td><td>2.87</td><td>73.17</td><td>5.89</td><td>70.18</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td><b>IndexTTS2.5</b></td><td>0.8B</td><td>4.36</td><td>77.10</td><td>5.12</td><td>68.06</td><td>3.75</td><td>76.39</td><td>5.66</td><td>74.62</td><td>14.88</td><td>69.74</td><td>6.75</td><td>73.18</td></tr>
<tr><td><b>IndexTTS2.5-RL</b></td><td>0.8B</td><td>3.93</td><td>77.92</td><td>3.89</td><td>67.79</td><td>3.33</td><td>76.68</td><td>5.30</td><td>75.41</td><td>13.58</td><td>70.36</td><td>6.00</td><td>73.63</td></tr>
</tbody>
</table>
**الجدول 2: تحويل النص إلى كلام عبر اللغات على CV3-Eval** (مطالبة صينية ← اللغة المستهدفة، تستخدم العربية مجموعة اختبار داخلية).
<table>
<thead>
<tr>
<th rowspan="2">النموذج</th>
<th rowspan="2">المعاملات</th>
<th colspan="2">zh→en</th>
<th colspan="2">zh→es</th>
<th colspan="2">zh→ja</th>
<th colspan="2">zh→ar</th>
<th colspan="2">Avg</th>
</tr>
<tr>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
</tr>
</thead>
<tbody>
<tr><td>VoxCPM2</td><td>2B</td><td>4.48</td><td>64.25</td><td>16.38</td><td>64.89</td><td>11.84</td><td>71.54</td><td>11.09</td><td>67.62</td><td>10.95</td><td>67.08</td></tr>
<tr><td>OmniVoice</td><td>0.8B</td><td>3.74</td><td>64.91</td><td>5.84</td><td>62.08</td><td>9.09</td><td>69.06</td><td>19.80</td><td>65.27</td><td>9.62</td><td>65.33</td></tr>
<tr><td>Moss-TTS 1.5</td><td>8B</td><td>6.13</td><td>59.23</td><td>4.32</td><td>56.63</td><td>11.52</td><td>65.54</td><td>17.03</td><td>62.93</td><td>9.75</td><td>61.08</td></tr>
<tr><td>CosyVoice3-0.5B</td><td>0.5B</td><td>3.23</td><td>62.79</td><td>4.58</td><td>64.04</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>CosyVoice3-1.5B</td><td>1.5B</td><td>4.32</td><td>-</td><td>-</td><td>-</td><td>13.70</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>FireRedTTS-2</td><td>1.5B</td><td>9.34</td><td>53.19</td><td>12.25</td><td>58.31</td><td>19.05</td><td>64.12</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>Fish Audio S2 Pro</td><td>4B</td><td>4.14</td><td>55.89</td><td>4.46</td><td>55.57</td><td>10.48</td><td>61.74</td><td>14.49</td><td>59.80</td><td>8.39</td><td>58.25</td></tr>
<tr><td>Qwen3-TTS</td><td>1.7B</td><td>5.74</td><td>63.04</td><td>5.15</td><td>68.02</td><td>36.09</td><td>65.71</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td><b>IndexTTS2.5</b></td><td>0.8B</td><td>3.62</td><td>63.83</td><td>5.17</td><td>65.48</td><td>6.57</td><td>74.16</td><td>9.51</td><td>71.02</td><td>6.22</td><td>68.62</td></tr>
<tr><td><b>IndexTTS2.5-RL</b></td><td>0.8B</td><td>3.55</td><td>67.47</td><td>4.86</td><td>64.47</td><td>6.38</td><td>75.82</td><td>9.89</td><td>73.05</td><td>6.17</td><td>70.20</td></tr>
</tbody>
</table>
## 🤝 المجتمع والتواصل
- **QQ Groups:** 663272642 (No.4), 1013410623 (No.5)
- **Discord:** https://discord.gg/uT32E7KDmy
- **Email:** indexspeech@bilibili.com
نرحب بانضمامك إلى مجتمعنا! 🌏 欢迎大家来交流讨论!
> [!CAUTION]
> شكرًا لدعمكم مشروع bilibili IndexTTS!
> يرجى ملاحظة أن **القناة الرسمية الوحيدة** التي يديرها الفريق الأساسي هي: [https://github.com/index-tts/index-tts](https://github.com/index-tts/index-tts).
> ***أي مواقع أو خدمات أخرى ليست رسمية***، ولا يمكننا ضمان أمانها أو دقتها أو تحديثها في الوقت المناسب.
> للاطلاع على آخر التحديثات، يرجى دائمًا الرجوع إلى هذا المستودع الرسمي.
للاستخدام التجاري والتعاون، يرجى التواصل عبر <u>indexspeech@bilibili.com</u>.
## 📚 الاستشهاد
🌟 إذا وجدت عملنا مفيدًا، يرجى منحنا نجمة والاستشهاد بأوراقنا البحثية.
IndexTTS2.5:
```bibtex
@misc{li2026indextts25technicalreport,
title={IndexTTS 2.5 Technical Report},
author={Yunpei Li and Xun Zhou and Jinchao Wang and Lu Wang and Yong Wu and Siyi Zhou and Yiquan Zhou and Yining Wang and Yaogen Yang and Zhetao Hu and Shiyao Duan and Jiacheng Xu and Bin Xia and Jingchen Shu},
year={2026},
eprint={2601.03888},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2601.03888},
}
```
IndexTTS2:
```bibtex
@article{zhou2025indextts2,
title={IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech},
author={Siyi Zhou and Yiquan Zhou and Yi He and Xun Zhou and Jinchao Wang and Wei Deng and Jingchen Shu},
journal={arXiv preprint arXiv:2506.21619},
year={2025}
}
```
IndexTTS:
```bibtex
@article{deng2025indextts,
title={IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System},
author={Wei Deng and Siyi Zhou and Jingchen Shu and Jinchao Wang and Lu Wang},
journal={arXiv preprint arXiv:2502.05512},
year={2025},
doi={10.48550/arXiv.2502.05512},
url={https://arxiv.org/abs/2502.05512}
}
```
## 🙏 شكر وتقدير
1. [tortoise-tts](https://github.com/neonbjb/tortoise-tts)
2. [XTTSv2](https://github.com/coqui-ai/TTS)
3. [BigVGAN](https://github.com/NVIDIA/BigVGAN)
4. [wenet](https://github.com/wenet-e2e/wenet/tree/main)
5. [icefall](https://github.com/k2-fsa/icefall)
6. [maskgct](https://github.com/open-mmlab/Amphion/tree/main/models/tts/maskgct)
7. [seed-vc](https://github.com/Plachtaa/seed-vc)
## 📄 الترخيص
هذا المشروع صادر بموجب [اتفاقية ترخيص استخدام نماذج bilibili](../LICENSE).
يرجى أيضًا قراءة [إخلاء المسؤولية](../DISCLAIMER) قبل الاستخدام.
+545
View File
@@ -0,0 +1,545 @@
<div align="center">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="../assets/indextts_icon_dark.png"/>
<img src="../assets/indextts_icon_light.png" width="300"/>
</picture>
**Un sistema de texto a voz zero-shot, controlable y eficiente, de nivel industrial**
[简体中文](README_zh.md) | [English](../README.md) | [日本語](README_ja.md) | Español | [العربية](README_ar.md)
[![GitHub Stars](https://img.shields.io/github/stars/index-tts/index-tts?style=flat&logo=github)](https://github.com/index-tts/index-tts/stargazers)
[![arXiv](https://img.shields.io/badge/arXiv-2601.03888-b31b1b?logo=arxiv)](https://arxiv.org/abs/2601.03888)
[![Discord](https://img.shields.io/badge/Discord-join-5865F2?logo=discord&logoColor=white)](https://discord.gg/uT32E7KDmy)
</div>
IndexTTS es un sistema de texto a voz zero-shot que clona una voz a partir de un
único clip de audio de referencia. La última versión, **IndexTTS-2.5**, admite
chino, inglés, japonés, español y árabe, con control de emociones de grano
fino, control de la velocidad de habla, control de la pronunciación (Pinyin /
fonemas CMU / Kana japonés) y una inferencia más rápida que IndexTTS-2.
---
## 🗂️ Catálogo de modelos
| Modelo | Demos | Artículo | ModelScope | HuggingFace |
| :--- | :---: | :---: | :---: | :---: |
| **IndexTTS-2.5** | [![Demo](https://img.shields.io/badge/Demo-Page-orange?logo=github)](https://index-tts.github.io/index-tts2-5.github.io/) | [![Paper](https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv)](https://arxiv.org/abs/2601.03888) | [![ModelScope](https://img.shields.io/badge/ModelScope-Model-purple?logo=modelscope)](https://modelscope.cn/models/IndexTeam/IndexTTS-2.5) | [![HuggingFace](https://img.shields.io/badge/HuggingFace-Model-blue?logo=huggingface)](https://huggingface.co/IndexTeam/IndexTTS-2.5) |
| **IndexTTS-2** | [![Demo](https://img.shields.io/badge/Demo-Page-orange?logo=github)](https://index-tts.github.io/index-tts2.github.io/) | [![Paper](https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv)](https://arxiv.org/abs/2506.21619) | [![ModelScope](https://img.shields.io/badge/ModelScope-Model-purple?logo=modelscope)](https://modelscope.cn/models/IndexTeam/IndexTTS-2) | [![HuggingFace](https://img.shields.io/badge/HuggingFace-Model-blue?logo=huggingface)](https://huggingface.co/IndexTeam/IndexTTS-2) |
| **IndexTTS-1.5** | [![Demo](https://img.shields.io/badge/Demo-Page-orange?logo=github)](https://index-tts.github.io/) | [![Paper](https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv)](https://arxiv.org/abs/2502.05512) | [![ModelScope](https://img.shields.io/badge/ModelScope-Model-purple?logo=modelscope)](https://modelscope.cn/models/IndexTeam/IndexTTS-1.5) | [![HuggingFace](https://img.shields.io/badge/HuggingFace-Model-blue?logo=huggingface)](https://huggingface.co/IndexTeam/IndexTTS-1.5) |
| **IndexTTS** | [![Demo](https://img.shields.io/badge/Demo-Page-orange?logo=github)](https://index-tts.github.io/) | [![Paper](https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv)](https://arxiv.org/abs/2502.05512) | [![ModelScope](https://img.shields.io/badge/ModelScope-Model-purple?logo=modelscope)](https://modelscope.cn/models/IndexTeam/Index-TTS) | [![HuggingFace](https://img.shields.io/badge/HuggingFace-Model-blue?logo=huggingface)](https://huggingface.co/IndexTeam/Index-TTS) |
## 📣 Novedades
- `2026/08/10` 🔥 Lanzamos **IndexTTS-2.5**
- Ahora admite chino, inglés, japonés, español y árabe, con una inferencia más rápida que IndexTTS-2, manteniendo las capacidades de síntesis multilingüe y de desacoplamiento timbre-emoción.
- Mayor controlabilidad del Pinyin chino, los fonemas CMU en inglés y el Kana japonés.
- Control de la velocidad de habla mediante `duration_factor` (0.5x2.0x de duración).
- `2025/09/08` 🔥 Lanzamos **IndexTTS-2**
- El primer modelo TTS autorregresivo con control preciso de la duración de la síntesis, compatible con los modos controlable y no controlable. <i>Esta funcionalidad aún no está habilitada en esta versión.</i>
- Síntesis de voz con emociones altamente expresivas, con control de emociones a través de múltiples modalidades de entrada.
- `2025/05/14` 🔥 Lanzamos **IndexTTS-1.5**, que mejora significativamente la estabilidad del modelo y su rendimiento en inglés.
- `2025/03/25` 🔥 Lanzamos **IndexTTS-1.0** con los pesos del modelo y el código de inferencia.
- `2025/02/12` 🎉 Enviamos nuestro artículo a arXiv y publicamos nuestras demos y conjuntos de prueba.
## 🎬 Demos
<div align="center">
**IndexTTS-2.5: El futuro de la voz, generándose ahora**
[![IndexTTS2.5 Demo](../assets/index2.5_video_cover.png)](https://www.bilibili.com/video/BV1uvMk6ZEdK/)
**IndexTTS-2: El futuro de la voz, generándose ahora**
[![IndexTTS2 Demo](../assets/IndexTTS2-video-pic.png)](https://www.bilibili.com/video/BV136a9zqEk5)
</div>
## 🚀 Primeros pasos
### 1. Requisitos previos
Asegúrate de tener [git](https://git-scm.com/downloads) instalado y luego descarga
este repositorio:
```bash
git clone https://github.com/index-tts/index-tts.git && cd index-tts
```
Los archivos de audio de ejemplo se descargan bajo demanda desde
HuggingFace/ModelScope en la primera ejecución, por lo que Git LFS ya no es
necesario.
### 2. Instalar dependencias
Usamos [uv](https://docs.astral.sh/uv/getting-started/installation/) para
gestionar el entorno de dependencias del proyecto. Es **obligatorio** para una
instalación fiable:
```bash
pip install -U uv # or see the link above for other install methods
```
```bash
uv sync --all-extras
```
Esto crea automáticamente un directorio de proyecto `.venv` e instala las
versiones correctas de Python y de todas las dependencias necesarias.
Si la descarga es lenta, usa un espejo local, por ejemplo uno de estos espejos
en China:
```bash
uv sync --all-extras --default-index "https://mirrors.aliyun.com/pypi/simple"
uv sync --all-extras --default-index "https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple"
```
> [!TIP]
> **Funciones adicionales disponibles:**
>
> - `--all-extras`: Añade automáticamente *todas* las funciones adicionales
> enumeradas a continuación. Puedes quitar este indicador si quieres
> personalizar tu instalación.
> - `--extra webui`: Añade compatibilidad con la WebUI (recomendado).
> - `--extra deepspeed`: Añade compatibilidad con DeepSpeed (puede acelerar la
> inferencia en algunos sistemas).
> [!IMPORTANT]
> **Windows:** DeepSpeed puede ser difícil de instalar. Puedes omitirlo
> quitando el indicador `--all-extras` y añadiendo manualmente los demás
> indicadores de funciones.
>
> **Linux/Windows:** Si ves un error de CUDA durante la instalación, asegúrate
> de que el [CUDA Toolkit](https://developer.nvidia.com/cuda-toolkit) de NVIDIA
> versión **12.8** (o superior) esté instalado en tu sistema.
### 3. Descargar los modelos
Descarga los modelos necesarios mediante [uv tool](https://docs.astral.sh/uv/guides/tools/#installing-tools):
Mediante `huggingface-cli`:
```bash
uv tool install "huggingface-hub[cli,hf_xet]"
# IndexTTS-2.5
hf download IndexTeam/IndexTTS-2.5 --local-dir=checkpoints
# IndexTTS-2
hf download IndexTeam/IndexTTS-2 --local-dir=checkpoints_2
```
O mediante `modelscope`:
```bash
uv tool install "modelscope"
# IndexTTS-2.5
modelscope download --model IndexTeam/IndexTTS-2.5 --local_dir checkpoints
# IndexTTS-2
modelscope download --model IndexTeam/IndexTTS-2 --local_dir checkpoints_2
```
> [!IMPORTANT]
> Si los comandos anteriores no están disponibles, lee atentamente la salida de
> `uv tool`: te indicará cómo añadir las herramientas al PATH de tu sistema.
> [!NOTE]
> Algunos modelos pequeños se descargan automáticamente en la primera
> ejecución. Si tu red accede lentamente a HuggingFace, configura un espejo
> antes de ejecutar el código:
>
> ```bash
> export HF_ENDPOINT="https://hf-mirror.com"
> ```
### 4. Comprobar la aceleración por GPU
Para diagnosticar tu entorno y ver qué GPU se detectan, usa la utilidad
incluida:
```bash
uv run tools/gpu_check.py
```
## 💻 Uso
### 🌐 Demo web
```bash
# IndexTTS-2.5 (default)
uv run webui.py --version 2.5 --model_dir ./checkpoints
# IndexTTS-2
uv run webui.py --version 2 --model_dir ./checkpoints_2
```
Abre tu navegador y visita `http://127.0.0.1:7860` para ver la demo.
Puedes ajustar la configuración para habilitar la inferencia en BF16
(IndexTTS-2.5) / FP16 (IndexTTS-2) (menor uso de VRAM), la aceleración con
DeepSpeed, núcleos CUDA compilados para mayor velocidad, etc. Todas las
opciones disponibles se pueden ver con:
```bash
uv run webui.py -h
```
> [!IMPORTANT]
> La inferencia en **FP16/BF16** (media precisión) es más rápida y usa menos
> VRAM, con una pérdida de calidad muy pequeña.
>
> **DeepSpeed** *puede* acelerar la inferencia en algunos sistemas, pero
> también podría hacerla más lenta: depende de tu hardware, controladores y
> sistema operativo. Pruébalo de ambas formas.
>
> Todos los comandos `uv` **activan automáticamente** el entorno virtual
> correcto del proyecto. *No* actives manualmente ningún entorno antes de
> ejecutar comandos `uv`, ya que eso puede causar conflictos de dependencias.
### 🚀 Servicio con vLLM
Para el despliegue en producción, consulta la [receta de vLLM para IndexTTS](https://github.com/vllm-project/recipes/pull/772).
### 📝 API de Python
Para ejecutar scripts, usa `uv run <file.py>` de modo que el código se ejecute
dentro del entorno de `uv`. Es posible que también necesites añadir el
directorio actual a `PYTHONPATH`:
```bash
# IndexTTS2
PYTHONPATH="$PYTHONPATH:." uv run indextts/infer_v2.py
# IndexTTS2.5
PYTHONPATH="$PYTHONPATH:." uv run indextts/infer_v2_5.py \
--cfg_path checkpoints/config.yaml \
--model_dir checkpoints \
--text "Hello world" \
--lang EN
```
#### 0. Inicializar IndexTTS
```python
# IndexTTS2
from indextts.infer_v2 import IndexTTS2
tts = IndexTTS2(cfg_path="checkpoints_2/config.yaml", model_dir="checkpoints_2", use_fp16=False, use_cuda_kernel=False, use_deepspeed=False)
# IndexTTS2.5
from indextts.infer_v2_5 import IndexTTS2
tts = IndexTTS2(cfg_path="checkpoints/config.yaml", model_dir="checkpoints", use_bf16=True)
```
#### 1. Clonación de voz con un único audio de referencia
```python
text = "Translate for me, what is a surprise!"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, output_path="gen.wav", verbose=True)
# IndexTTS2.5 (multilingual, with language selection)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="EN", output_path="gen.wav", verbose=True)
```
#### 2. Control de emociones con un audio de referencia emocional independiente
```python
text = "酒楼丧尽天良,开始借机竞拍房间,哎,一群蠢货。"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", emo_audio_prompt="examples/emo_sad.wav", verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, lang="ZH", output_path="gen.wav", emo_audio_prompt="examples/emo_sad.wav", verbose=True)
```
#### 3. Ajustar la intensidad de la emoción con `emo_alpha`
Cuando se especifica un audio de referencia emocional, `emo_alpha` ajusta
cuánto afecta al resultado. Rango válido: `0.0 - 1.0`, valor por defecto:
`1.0` (100%).
```python
text = "酒楼丧尽天良,开始借机竞拍房间,哎,一群蠢货。"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", emo_audio_prompt="examples/emo_sad.wav", emo_alpha=0.9, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", lang="ZH", emo_audio_prompt="examples/emo_sad.wav", emo_alpha=0.9, verbose=True)
```
#### 4. Control de emociones con un vector de emociones
Puedes omitir el audio de referencia emocional y, en su lugar, proporcionar una
lista de 8 valores flotantes que especifican la intensidad de cada emoción, en
el orden
`[alegre, enfadado, triste, asustado, disgustado, melancólico, sorprendido, tranquilo]`.
Usa `use_random` para introducir estocasticidad durante la inferencia
(valor por defecto: `False`).
> [!NOTE]
> Activar el muestreo aleatorio reduce la fidelidad de la clonación de voz.
```python
text = "对不起嘛!我的记性真的不太好,但是和你在一起的事情,我都会努力记住的~"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/09.wav', text=text, output_path="gen.wav", emo_vector=[0, 0, 0.8, 0, 0, 0, 0, 0], use_random=False, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/09.wav', text=text, lang="ZH", output_path="gen.wav", emo_vector=[0, 0, 0.8, 0, 0, 0, 0, 0], use_random=False, verbose=True)
```
#### 5. Control de emociones a partir del propio texto (`use_emo_text`)
Activa `use_emo_text` para convertir automáticamente tu guion `text` en
vectores de emociones. Se recomienda un `emo_alpha` en torno a 0.6 (o menor)
para una voz más natural. Se puede introducir aleatoriedad con `use_random`
(valor por defecto: `False`).
```python
text = "快躲起来!是他要来了!他要来抓我们了!"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, use_random=False, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, lang="ZH", output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, use_random=False, verbose=True)
```
#### 6. Control de emociones con una descripción explícita de la emoción (`emo_text`)
Proporciona una descripción textual específica de la emoción mediante
`emo_text`, que se convierte en vectores de emociones, lo que te permite
controlar por separado el guion del texto y la descripción de la emoción:
```python
text = "快躲起来!是他要来了!他要来抓我们了!"
emo_text = "你吓死我了!你是鬼吗?"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, emo_text=emo_text, use_random=False, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, lang="ZH", output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, emo_text=emo_text, use_random=False, verbose=True)
```
#### 7. Control de la velocidad de habla (`duration_factor`)
Un valor mayor que `1.0` ralentiza el habla; un valor menor que `1.0` la
acelera. Valor por defecto: `1.0` (velocidad normal). Rango válido:
`0.5 - 2.0`.
```python
text = "大家好,欢迎来到IndexTTS的语速控制演示。"
# IndexTTS2.5
# Slow down (1.2x duration)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="ZH", output_path="gen_slow.wav", duration_factor=1.2, verbose=True)
# Speed up (0.8x duration)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="ZH", output_path="gen_fast.wav", duration_factor=0.8, verbose=True)
```
### 🗣️ Control de la pronunciación
**IndexTTS2.5 — Pinyin / fonemas CMU / Kana japonés:**
IndexTTS2.5 admite estas sustituciones de caracteres con una mejor capacidad
de seguimiento de instrucciones. Para ver la lista completa de entradas
válidas, consulta `checkpoints/pinyin.vocab` para el Pinyin y el
[diccionario CMU](https://svn.code.sf.net/p/cmusphinx/code/trunk/cmudict/cmudict-0.7b)
para los fonemas en inglés.
```
他在银<行|XING2>里<行|HANG2>走了半天,发现这笔业务办不<行|HANG2>。
He had a <minute|M IH1 . N AH0 T> to examine the <minute|M AY0 . N UW1 T> details of the contract.
彼は料理が<上手|じょうず>だが、囲碁では<上手|うわて>に負けた。
```
**IndexTTS2 — Pinyin:**
IndexTTS2 admite el modelado mixto de caracteres chinos y Pinyin. Para activar
el control por Pinyin, proporciona texto con anotaciones específicas de Pinyin.
Ten en cuenta que el control por Pinyin no funciona para todas las
combinaciones posibles de consonante y vocal; solo se admiten los casos de
Pinyin chino válidos (consulta `checkpoints/pinyin.vocab`).
```
之前你做DE5很好,所以这一次也DEI3做DE2很好才XING2,如果这次目标完成得不错的话,我们就直接打DI1去银行取钱。
```
### 🕰️ IndexTTS-1.5 (heredado)
También puedes usar el modelo anterior IndexTTS1 importando un módulo
diferente:
```python
from indextts.infer import IndexTTS
tts = IndexTTS(model_dir="checkpoints", cfg_path="checkpoints/config.yaml")
voice = "examples/voice_07.wav"
text = "大家好,我现在正在bilibili 体验 ai 科技,说实话,来之前我绝对想不到!AI技术已经发展到这样匪夷所思的地步了!比如说,现在正在说话的其实是B站为我现场复刻的数字分身,简直就是平行宇宙的另一个我了。如果大家也想体验更多深入的AIGC功能,可以访问 bilibili studio,相信我,你们也会吃惊的。"
tts.infer(voice, text, 'gen.wav')
```
Para más detalles, consulta [README_INDEXTTS_1_5](../archive/README_INDEXTTS_1_5.md),
o visita el repositorio de IndexTTS1 en [index-tts:v1.5.0](https://github.com/index-tts/index-tts/tree/v1.5.0).
## 📊 Evaluación
**Tabla 1: TTS zero-shot en CV3-Eval** (el árabe usa un conjunto de prueba interno). †Citado del artículo original.
<table>
<thead>
<tr>
<th rowspan="2">Modelo</th>
<th rowspan="2">Parámetros</th>
<th colspan="2">zh</th>
<th colspan="2">en</th>
<th colspan="2">es</th>
<th colspan="2">ja</th>
<th colspan="2">ar</th>
<th colspan="2">Prom.</th>
</tr>
<tr>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
</tr>
</thead>
<tbody>
<tr><td>VoxCPM2</td><td>2B</td><td>3.88</td><td>74.99</td><td>5.13</td><td>71.57</td><td>5.49</td><td>74.67</td><td>6.69</td><td>72.90</td><td>14.94</td><td>65.99</td><td>7.22</td><td>72.02</td></tr>
<tr><td>OmniVoice</td><td>0.8B</td><td>3.41</td><td>72.99</td><td>3.62</td><td>70.13</td><td>3.52</td><td>74.14</td><td>5.38</td><td>70.49</td><td>17.88</td><td>64.22</td><td>6.76</td><td>70.39</td></tr>
<tr><td>Moss-TTS 1.5</td><td>8B</td><td>4.02</td><td>72.68</td><td>4.45</td><td>67.46</td><td>3.83</td><td>71.75</td><td>10.97</td><td>68.71</td><td>23.71</td><td>62.21</td><td>9.40</td><td>68.56</td></tr>
<tr><td>CosyVoice3-0.5B</td><td>0.5B</td><td>3.84</td><td>80.01</td><td>4.88</td><td>74.16</td><td>4.04</td><td>78.85</td><td>-</td><td>76.36</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>CosyVoice3-1.5B</td><td>1.5B</td><td>3.91†</td><td>-</td><td>4.99†</td><td>-</td><td>4.47†</td><td>-</td><td>7.57†</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>FireRedTTS-2</td><td>1.5B</td><td>8.22</td><td>68.10</td><td>14.92</td><td>56.93</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>Fish Audio S2 Pro</td><td>4B</td><td>3.62</td><td>67.79</td><td>3.83</td><td>61.66</td><td>2.93</td><td>67.44</td><td>5.15</td><td>66.15</td><td>14.15</td><td>59.43</td><td>5.94</td><td>64.49</td></tr>
<tr><td>Qwen3-TTS</td><td>1.7B</td><td>3.27</td><td>73.02</td><td>5.06</td><td>67.17</td><td>2.87</td><td>73.17</td><td>5.89</td><td>70.18</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td><b>IndexTTS2.5</b></td><td>0.8B</td><td>4.36</td><td>77.10</td><td>5.12</td><td>68.06</td><td>3.75</td><td>76.39</td><td>5.66</td><td>74.62</td><td>14.88</td><td>69.74</td><td>6.75</td><td>73.18</td></tr>
<tr><td><b>IndexTTS2.5-RL</b></td><td>0.8B</td><td>3.93</td><td>77.92</td><td>3.89</td><td>67.79</td><td>3.33</td><td>76.68</td><td>5.30</td><td>75.41</td><td>13.58</td><td>70.36</td><td>6.00</td><td>73.63</td></tr>
</tbody>
</table>
**Tabla 2: TTS multilingüe en CV3-Eval** (prompt en chino → idioma de destino; el árabe usa un conjunto de prueba interno).
<table>
<thead>
<tr>
<th rowspan="2">Modelo</th>
<th rowspan="2">Parámetros</th>
<th colspan="2">zh→en</th>
<th colspan="2">zh→es</th>
<th colspan="2">zh→ja</th>
<th colspan="2">zh→ar</th>
<th colspan="2">Prom.</th>
</tr>
<tr>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
</tr>
</thead>
<tbody>
<tr><td>VoxCPM2</td><td>2B</td><td>4.48</td><td>64.25</td><td>16.38</td><td>64.89</td><td>11.84</td><td>71.54</td><td>11.09</td><td>67.62</td><td>10.95</td><td>67.08</td></tr>
<tr><td>OmniVoice</td><td>0.8B</td><td>3.74</td><td>64.91</td><td>5.84</td><td>62.08</td><td>9.09</td><td>69.06</td><td>19.80</td><td>65.27</td><td>9.62</td><td>65.33</td></tr>
<tr><td>Moss-TTS 1.5</td><td>8B</td><td>6.13</td><td>59.23</td><td>4.32</td><td>56.63</td><td>11.52</td><td>65.54</td><td>17.03</td><td>62.93</td><td>9.75</td><td>61.08</td></tr>
<tr><td>CosyVoice3-0.5B</td><td>0.5B</td><td>3.23</td><td>62.79</td><td>4.58</td><td>64.04</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>CosyVoice3-1.5B</td><td>1.5B</td><td>4.32</td><td>-</td><td>-</td><td>-</td><td>13.70</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>FireRedTTS-2</td><td>1.5B</td><td>9.34</td><td>53.19</td><td>12.25</td><td>58.31</td><td>19.05</td><td>64.12</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>Fish Audio S2 Pro</td><td>4B</td><td>4.14</td><td>55.89</td><td>4.46</td><td>55.57</td><td>10.48</td><td>61.74</td><td>14.49</td><td>59.80</td><td>8.39</td><td>58.25</td></tr>
<tr><td>Qwen3-TTS</td><td>1.7B</td><td>5.74</td><td>63.04</td><td>5.15</td><td>68.02</td><td>36.09</td><td>65.71</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td><b>IndexTTS2.5</b></td><td>0.8B</td><td>3.62</td><td>63.83</td><td>5.17</td><td>65.48</td><td>6.57</td><td>74.16</td><td>9.51</td><td>71.02</td><td>6.22</td><td>68.62</td></tr>
<tr><td><b>IndexTTS2.5-RL</b></td><td>0.8B</td><td>3.55</td><td>67.47</td><td>4.86</td><td>64.47</td><td>6.38</td><td>75.82</td><td>9.89</td><td>73.05</td><td>6.17</td><td>70.20</td></tr>
</tbody>
</table>
## 🤝 Comunidad y contacto
- **Grupos de QQ:** 663272642 (n.º 4), 1013410623 (n.º 5)
- **Discord:** https://discord.gg/uT32E7KDmy
- **Correo electrónico:** indexspeech@bilibili.com
¡Te invitamos a unirte a nuestra comunidad! 🌏 欢迎大家来交流讨论!
> [!CAUTION]
> ¡Gracias por tu apoyo al proyecto IndexTTS de bilibili!
> Ten en cuenta que el **único canal oficial** mantenido por el equipo central es: [https://github.com/index-tts/index-tts](https://github.com/index-tts/index-tts).
> ***Cualquier otro sitio web o servicio no es oficial*** y no podemos garantizar su seguridad, exactitud ni actualidad.
> Para conocer las últimas novedades, consulta siempre este repositorio oficial.
Para uso comercial y colaboraciones, contacta con <u>indexspeech@bilibili.com</u>.
## 📚 Citación
🌟 Si nuestro trabajo te resulta útil, déjanos una estrella y cita nuestros artículos.
IndexTTS2.5:
```bibtex
@misc{li2026indextts25technicalreport,
title={IndexTTS 2.5 Technical Report},
author={Yunpei Li and Xun Zhou and Jinchao Wang and Lu Wang and Yong Wu and Siyi Zhou and Yiquan Zhou and Yining Wang and Yaogen Yang and Zhetao Hu and Shiyao Duan and Jiacheng Xu and Bin Xia and Jingchen Shu},
year={2026},
eprint={2601.03888},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2601.03888},
}
```
IndexTTS2:
```bibtex
@article{zhou2025indextts2,
title={IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech},
author={Siyi Zhou and Yiquan Zhou and Yi He and Xun Zhou and Jinchao Wang and Wei Deng and Jingchen Shu},
journal={arXiv preprint arXiv:2506.21619},
year={2025}
}
```
IndexTTS:
```bibtex
@article{deng2025indextts,
title={IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System},
author={Wei Deng and Siyi Zhou and Jingchen Shu and Jinchao Wang and Lu Wang},
journal={arXiv preprint arXiv:2502.05512},
year={2025},
doi={10.48550/arXiv.2502.05512},
url={https://arxiv.org/abs/2502.05512}
}
```
## 🙏 Agradecimientos
1. [tortoise-tts](https://github.com/neonbjb/tortoise-tts)
2. [XTTSv2](https://github.com/coqui-ai/TTS)
3. [BigVGAN](https://github.com/NVIDIA/BigVGAN)
4. [wenet](https://github.com/wenet-e2e/wenet/tree/main)
5. [icefall](https://github.com/k2-fsa/icefall)
6. [maskgct](https://github.com/open-mmlab/Amphion/tree/main/models/tts/maskgct)
7. [seed-vc](https://github.com/Plachtaa/seed-vc)
## 📄 Licencia
Este proyecto se publica bajo el [Acuerdo de Licencia de Uso de Modelos de bilibili](../LICENSE).
Lee también el [DESCARGO DE RESPONSABILIDAD](../DISCLAIMER) antes de usarlo.
+531
View File
@@ -0,0 +1,531 @@
<div align="center">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="../assets/indextts_icon_dark.png"/>
<img src="../assets/indextts_icon_light.png" width="300"/>
</picture>
**産業レベルの制御可能で効率的なゼロショット・テキスト読み上げシステム**
[简体中文](README_zh.md) | [English](../README.md) | 日本語 | [Español](README_es.md) | [العربية](README_ar.md)
[![GitHub Stars](https://img.shields.io/github/stars/index-tts/index-tts?style=flat&logo=github)](https://github.com/index-tts/index-tts/stargazers)
[![arXiv](https://img.shields.io/badge/arXiv-2601.03888-b31b1b?logo=arxiv)](https://arxiv.org/abs/2601.03888)
[![Discord](https://img.shields.io/badge/Discord-join-5865F2?logo=discord&logoColor=white)](https://discord.gg/uT32E7KDmy)
</div>
IndexTTS は、1つの参照音声クリップから声をクローンするゼロショット・テキスト
読み上げシステムです。最新リリースの **IndexTTS-2.5** は、中国語、英語、日本語、
スペイン語、アラビア語をサポートし、きめ細かな感情コントロール、話速コントロール、
発音コントロール(ピンイン / CMU 音素 / 日本語の仮名)を備え、IndexTTS-2 よりも
高速な推論を実現しています。
---
## 🗂️ モデル一覧
| モデル | デモ | 論文 | ModelScope | HuggingFace |
| :--- | :---: | :---: | :---: | :---: |
| **IndexTTS-2.5** | [![Demo](https://img.shields.io/badge/Demo-Page-orange?logo=github)](https://index-tts.github.io/index-tts2-5.github.io/) | [![Paper](https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv)](https://arxiv.org/abs/2601.03888) | [![ModelScope](https://img.shields.io/badge/ModelScope-Model-purple?logo=modelscope)](https://modelscope.cn/models/IndexTeam/IndexTTS-2.5) | [![HuggingFace](https://img.shields.io/badge/HuggingFace-Model-blue?logo=huggingface)](https://huggingface.co/IndexTeam/IndexTTS-2.5) |
| **IndexTTS-2** | [![Demo](https://img.shields.io/badge/Demo-Page-orange?logo=github)](https://index-tts.github.io/index-tts2.github.io/) | [![Paper](https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv)](https://arxiv.org/abs/2506.21619) | [![ModelScope](https://img.shields.io/badge/ModelScope-Model-purple?logo=modelscope)](https://modelscope.cn/models/IndexTeam/IndexTTS-2) | [![HuggingFace](https://img.shields.io/badge/HuggingFace-Model-blue?logo=huggingface)](https://huggingface.co/IndexTeam/IndexTTS-2) |
| **IndexTTS-1.5** | [![Demo](https://img.shields.io/badge/Demo-Page-orange?logo=github)](https://index-tts.github.io/) | [![Paper](https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv)](https://arxiv.org/abs/2502.05512) | [![ModelScope](https://img.shields.io/badge/ModelScope-Model-purple?logo=modelscope)](https://modelscope.cn/models/IndexTeam/IndexTTS-1.5) | [![HuggingFace](https://img.shields.io/badge/HuggingFace-Model-blue?logo=huggingface)](https://huggingface.co/IndexTeam/IndexTTS-1.5) |
| **IndexTTS** | [![Demo](https://img.shields.io/badge/Demo-Page-orange?logo=github)](https://index-tts.github.io/) | [![Paper](https://img.shields.io/badge/Paper-arXiv-red?logo=arxiv)](https://arxiv.org/abs/2502.05512) | [![ModelScope](https://img.shields.io/badge/ModelScope-Model-purple?logo=modelscope)](https://modelscope.cn/models/IndexTeam/Index-TTS) | [![HuggingFace](https://img.shields.io/badge/HuggingFace-Model-blue?logo=huggingface)](https://huggingface.co/IndexTeam/Index-TTS) |
## 📣 ニュース
- `2026/08/10` 🔥 **IndexTTS-2.5** をリリースしました
- 中国語、英語、日本語、スペイン語、アラビア語をサポートし、IndexTTS-2 よりも高速な推論を実現しながら、言語横断的および音色・感情分離の能力を維持しています。
- 中国語ピンイン、英語 CMU 音素、日本語仮名の制御性が向上しました。
- `duration_factor` による話速コントロール(0.5倍~2.0倍の長さ)。
- `2025/09/08` 🔥 **IndexTTS-2** をリリースしました
- 精密な合成時間制御を備えた初の自己回帰型 TTS モデルで、制御可能モードと非制御モードの両方をサポートします。<i>この機能は本リリースではまだ有効になっていません。</i>
- 高い表現力を持つ感情音声合成。複数の入力モダリティによる感情コントロールが可能です。
- `2025/05/14` 🔥 **IndexTTS-1.5** をリリースしました。モデルの安定性と英語での性能が大幅に向上しています。
- `2025/03/25` 🔥 **IndexTTS-1.0** をモデルウェイトおよび推論コードとともにリリースしました。
- `2025/02/12` 🎉 論文を arXiv に投稿し、デモとテストセットを公開しました。
## 🎬 デモ
<div align="center">
**IndexTTS-2.5: The Future of Voice, Now Generating**
[![IndexTTS2.5 Demo](../assets/index2.5_video_cover.png)](https://www.bilibili.com/video/BV1uvMk6ZEdK/)
**IndexTTS-2: The Future of Voice, Now Generating**
[![IndexTTS2 Demo](../assets/IndexTTS2-video-pic.png)](https://www.bilibili.com/video/BV136a9zqEk5)
</div>
## 🚀 はじめに
### 1. 前提条件
[git](https://git-scm.com/downloads) がインストールされていることを確認してから、
このリポジトリをダウンロードしてください:
```bash
git clone https://github.com/index-tts/index-tts.git && cd index-tts
```
サンプル音声ファイルは初回実行時に HuggingFace/ModelScope からオンデマンドで
ダウンロードされるため、Git LFS は不要になりました。
### 2. 依存関係のインストール
プロジェクトの依存環境の管理には [uv](https://docs.astral.sh/uv/getting-started/installation/) を使用しています。確実なインストールのために **必須** です:
```bash
pip install -U uv # or see the link above for other install methods
```
```bash
uv sync --all-extras
```
これにより `.venv` プロジェクトディレクトリが自動的に作成され、正しいバージョンの
Python と必要なすべての依存関係がインストールされます。
ダウンロードが遅い場合はローカルミラーを使用してください。例えば中国国内の
以下のミラーが利用できます:
```bash
uv sync --all-extras --default-index "https://mirrors.aliyun.com/pypi/simple"
uv sync --all-extras --default-index "https://mirrors.tuna.tsinghua.edu.cn/pypi/web/simple"
```
> [!TIP]
> **利用可能な追加機能:**
>
> - `--all-extras`: 以下に挙げる *すべての* 追加機能を自動的に追加します。
> インストール内容をカスタマイズしたい場合は、このフラグを外してください。
> - `--extra webui`: WebUI サポートを追加します(推奨)。
> - `--extra deepspeed`: DeepSpeed サポートを追加します(一部のシステムで推論が
> 高速化する場合があります)。
> [!IMPORTANT]
> **Windows:** DeepSpeed のインストールが難しい場合があります。`--all-extras`
> フラグを外し、他の機能フラグを個別に追加することでスキップできます。
>
> **Linux/Windows:** インストール中に CUDA エラーが表示された場合は、NVIDIA の
> [CUDA Toolkit](https://developer.nvidia.com/cuda-toolkit) バージョン **12.8**
> (以降)がシステムにインストールされていることを確認してください。
### 3. モデルのダウンロード
[uv tool](https://docs.astral.sh/uv/guides/tools/#installing-tools) を使って必要なモデルをダウンロードします:
`huggingface-cli` を使う場合:
```bash
uv tool install "huggingface-hub[cli,hf_xet]"
# IndexTTS-2.5
hf download IndexTeam/IndexTTS-2.5 --local-dir=checkpoints
# IndexTTS-2
hf download IndexTeam/IndexTTS-2 --local-dir=checkpoints_2
```
または `modelscope` を使う場合:
```bash
uv tool install "modelscope"
# IndexTTS-2.5
modelscope download --model IndexTeam/IndexTTS-2.5 --local_dir checkpoints
# IndexTTS-2
modelscope download --model IndexTeam/IndexTTS-2 --local_dir checkpoints_2
```
> [!IMPORTANT]
> 上記のコマンドが利用できない場合は、`uv tool` の出力をよく確認してください。
> ツールをシステムの PATH に追加する方法が表示されます。
> [!NOTE]
> 一部の小さなモデルは初回実行時に自動的にダウンロードされます。ネットワークから
> HuggingFace へのアクセスが遅い場合は、コードを実行する前にミラーを設定してください:
>
> ```bash
> export HF_ENDPOINT="https://hf-mirror.com"
> ```
### 4. GPU アクセラレーションの確認
環境を診断し、どの GPU が検出されているかを確認するには、付属のユーティリティを
使用してください:
```bash
uv run tools/gpu_check.py
```
## 💻 使い方
### 🌐 Web デモ
```bash
# IndexTTS-2.5 (default)
uv run webui.py --version 2.5 --model_dir ./checkpoints
# IndexTTS-2
uv run webui.py --version 2 --model_dir ./checkpoints_2
```
ブラウザを開いて `http://127.0.0.1:7860` にアクセスするとデモが表示されます。
設定を調整して、BF16IndexTTS-2.5/ FP16IndexTTS-2)推論(VRAM 使用量の削減)、
DeepSpeed アクセラレーション、高速化のためのコンパイル済み CUDA カーネルなどを
有効にできます。利用可能なすべてのオプションは以下で確認できます:
```bash
uv run webui.py -h
```
> [!IMPORTANT]
> **FP16/BF16**(半精度)推論は高速で VRAM 使用量も少なく、品質の低下は
> ごくわずかです。
>
> **DeepSpeed** は一部のシステムで推論を高速化する*可能性*がありますが、逆に
> 遅くなる場合もあります。ハードウェア、ドライバ、OS に依存します。両方を
> 試してみてください。
>
> すべての `uv` コマンドは、プロジェクトごとの正しい仮想環境を**自動的に
> アクティベート**します。`uv` コマンドを実行する前に手動で環境をアクティベート
> *しないでください*。依存関係の競合を引き起こす可能性があります。
### 🚀 vLLM によるサービング
本番環境へのデプロイについては、[IndexTTS 向け vLLM レシピ](https://github.com/vllm-project/recipes/pull/772)を参照してください。
### 📝 Python API
スクリプトを実行するには `uv run <file.py>` を使用し、コードが `uv` 環境内で
実行されるようにしてください。カレントディレクトリを `PYTHONPATH` に追加する
必要がある場合もあります:
```bash
# IndexTTS2
PYTHONPATH="$PYTHONPATH:." uv run indextts/infer_v2.py
# IndexTTS2.5
PYTHONPATH="$PYTHONPATH:." uv run indextts/infer_v2_5.py \
--cfg_path checkpoints/config.yaml \
--model_dir checkpoints \
--text "Hello world" \
--lang EN
```
#### 0. IndexTTS の初期化
```python
# IndexTTS2
from indextts.infer_v2 import IndexTTS2
tts = IndexTTS2(cfg_path="checkpoints_2/config.yaml", model_dir="checkpoints_2", use_fp16=False, use_cuda_kernel=False, use_deepspeed=False)
# IndexTTS2.5
from indextts.infer_v2_5 import IndexTTS2
tts = IndexTTS2(cfg_path="checkpoints/config.yaml", model_dir="checkpoints", use_bf16=True)
```
#### 1. 1つの参照音声によるボイスクローニング
```python
text = "Translate for me, what is a surprise!"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, output_path="gen.wav", verbose=True)
# IndexTTS2.5 (multilingual, with language selection)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="EN", output_path="gen.wav", verbose=True)
```
#### 2. 別の感情参照音声による感情コントロール
```python
text = "酒楼丧尽天良,开始借机竞拍房间,哎,一群蠢货。"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", emo_audio_prompt="examples/emo_sad.wav", verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, lang="ZH", output_path="gen.wav", emo_audio_prompt="examples/emo_sad.wav", verbose=True)
```
#### 3. `emo_alpha` による感情強度の調整
感情参照音声が指定されている場合、`emo_alpha` でそれが出力に与える影響の
大きさを調整できます。有効範囲:`0.0 - 1.0`、デフォルト:`1.0`100%)。
```python
text = "酒楼丧尽天良,开始借机竞拍房间,哎,一群蠢货。"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", emo_audio_prompt="examples/emo_sad.wav", emo_alpha=0.9, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_07.wav', text=text, output_path="gen.wav", lang="ZH", emo_audio_prompt="examples/emo_sad.wav", emo_alpha=0.9, verbose=True)
```
#### 4. 感情ベクトルによる感情コントロール
感情参照音声を省略し、代わりに各感情の強度を指定する 8 要素の浮動小数点数リストを
指定できます。順序は
`[happy, angry, sad, afraid, disgusted, melancholic, surprised, calm]`
です。`use_random` を使うと推論時に確率的な揺らぎを導入できます(デフォルト:`False`)。
> [!NOTE]
> ランダムサンプリングを有効にすると、ボイスクローニングの忠実度が低下します。
```python
text = "对不起嘛!我的记性真的不太好,但是和你在一起的事情,我都会努力记住的~"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/09.wav', text=text, output_path="gen.wav", emo_vector=[0, 0, 0.8, 0, 0, 0, 0, 0], use_random=False, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/09.wav', text=text, lang="ZH", output_path="gen.wav", emo_vector=[0, 0, 0.8, 0, 0, 0, 0, 0], use_random=False, verbose=True)
```
#### 5. テキスト自体からの感情コントロール(`use_emo_text`
`use_emo_text` を有効にすると、`text` のスクリプトが自動的に感情ベクトルに
変換されます。より自然な音声にするためには、`emo_alpha` を 0.6 前後(または
それ以下)にすることを推奨します。`use_random` でランダム性を導入できます
(デフォルト:`False`)。
```python
text = "快躲起来!是他要来了!他要来抓我们了!"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, use_random=False, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, lang="ZH", output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, use_random=False, verbose=True)
```
#### 6. 明示的な感情記述による感情コントロール(`emo_text`)
`emo_text` で特定の感情記述テキストを指定すると、それが感情ベクトルに変換
されます。これにより、テキストスクリプトと感情記述を別々にコントロールできます:
```python
text = "快躲起来!是他要来了!他要来抓我们了!"
emo_text = "你吓死我了!你是鬼吗?"
# IndexTTS2
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, emo_text=emo_text, use_random=False, verbose=True)
# IndexTTS2.5
tts.infer(spk_audio_prompt='examples/voice_12.wav', text=text, lang="ZH", output_path="gen.wav", emo_alpha=0.6, use_emo_text=True, emo_text=emo_text, use_random=False, verbose=True)
```
#### 7. 話速コントロール(`duration_factor`
`1.0` より大きい値を指定すると音声が遅くなり、`1.0` より小さい値を指定すると
速くなります。デフォルト:`1.0`(通常速度)。有効範囲:`0.5 - 2.0`
```python
text = "大家好,欢迎来到IndexTTS的语速控制演示。"
# IndexTTS2.5
# Slow down (1.2x duration)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="ZH", output_path="gen_slow.wav", duration_factor=1.2, verbose=True)
# Speed up (0.8x duration)
tts.infer(spk_audio_prompt='examples/voice_01.wav', text=text, lang="ZH", output_path="gen_fast.wav", duration_factor=0.8, verbose=True)
```
### 🗣️ 発音コントロール
**IndexTTS2.5 — ピンイン / CMU 音素 / 日本語仮名:**
IndexTTS2.5 は、より優れた指示追従能力により、これらの文字置換をサポートします。
有効なエントリの完全なリストについては、ピンインは `checkpoints/pinyin.vocab`
英語音素は [CMU 辞書](https://svn.code.sf.net/p/cmusphinx/code/trunk/cmudict/cmudict-0.7b)を参照してください。
```
他在银<行|XING2>里<行|HANG2>走了半天,发现这笔业务办不<行|HANG2>。
He had a <minute|M IH1 . N AH0 T> to examine the <minute|M AY0 . N UW1 T> details of the contract.
彼は料理が<上手|じょうず>だが、囲碁では<上手|うわて>に負けた。
```
**IndexTTS2 — ピンイン:**
IndexTTS2 は、漢字とピンインの混合モデリングをサポートします。ピンイン
コントロールを有効にするには、特定のピンイン注釈を付けたテキストを入力します。
なお、ピンインコントロールはすべての子音・母音の組み合わせで機能するわけでは
ありません。有効な中国語ピンインの場合のみサポートされます
`checkpoints/pinyin.vocab` を参照)。
```
之前你做DE5很好,所以这一次也DEI3做DE2很好才XING2,如果这次目标完成得不错的话,我们就直接打DI1去银行取钱。
```
### 🕰️ IndexTTS-1.5(レガシー)
別のモジュールをインポートすることで、以前の IndexTTS1 モデルを使用することも
できます:
```python
from indextts.infer import IndexTTS
tts = IndexTTS(model_dir="checkpoints", cfg_path="checkpoints/config.yaml")
voice = "examples/voice_07.wav"
text = "大家好,我现在正在bilibili 体验 ai 科技,说实话,来之前我绝对想不到!AI技术已经发展到这样匪夷所思的地步了!比如说,现在正在说话的其实是B站为我现场复刻的数字分身,简直就是平行宇宙的另一个我了。如果大家也想体验更多深入的AIGC功能,可以访问 bilibili studio,相信我,你们也会吃惊的。"
tts.infer(voice, text, 'gen.wav')
```
詳細については [README_INDEXTTS_1_5](../archive/README_INDEXTTS_1_5.md) を参照するか、
[index-tts:v1.5.0](https://github.com/index-tts/index-tts/tree/v1.5.0) の IndexTTS1 リポジトリをご覧ください。
## 📊 評価
**表 1: CV3-Eval におけるゼロショット TTS**(アラビア語は社内テストセットを使用)。†原論文より引用。
<table>
<thead>
<tr>
<th rowspan="2">Model</th>
<th rowspan="2">Params</th>
<th colspan="2">zh</th>
<th colspan="2">en</th>
<th colspan="2">es</th>
<th colspan="2">ja</th>
<th colspan="2">ar</th>
<th colspan="2">Avg</th>
</tr>
<tr>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
</tr>
</thead>
<tbody>
<tr><td>VoxCPM2</td><td>2B</td><td>3.88</td><td>74.99</td><td>5.13</td><td>71.57</td><td>5.49</td><td>74.67</td><td>6.69</td><td>72.90</td><td>14.94</td><td>65.99</td><td>7.22</td><td>72.02</td></tr>
<tr><td>OmniVoice</td><td>0.8B</td><td>3.41</td><td>72.99</td><td>3.62</td><td>70.13</td><td>3.52</td><td>74.14</td><td>5.38</td><td>70.49</td><td>17.88</td><td>64.22</td><td>6.76</td><td>70.39</td></tr>
<tr><td>Moss-TTS 1.5</td><td>8B</td><td>4.02</td><td>72.68</td><td>4.45</td><td>67.46</td><td>3.83</td><td>71.75</td><td>10.97</td><td>68.71</td><td>23.71</td><td>62.21</td><td>9.40</td><td>68.56</td></tr>
<tr><td>CosyVoice3-0.5B</td><td>0.5B</td><td>3.84</td><td>80.01</td><td>4.88</td><td>74.16</td><td>4.04</td><td>78.85</td><td>-</td><td>76.36</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>CosyVoice3-1.5B</td><td>1.5B</td><td>3.91†</td><td>-</td><td>4.99†</td><td>-</td><td>4.47†</td><td>-</td><td>7.57†</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>FireRedTTS-2</td><td>1.5B</td><td>8.22</td><td>68.10</td><td>14.92</td><td>56.93</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>Fish Audio S2 Pro</td><td>4B</td><td>3.62</td><td>67.79</td><td>3.83</td><td>61.66</td><td>2.93</td><td>67.44</td><td>5.15</td><td>66.15</td><td>14.15</td><td>59.43</td><td>5.94</td><td>64.49</td></tr>
<tr><td>Qwen3-TTS</td><td>1.7B</td><td>3.27</td><td>73.02</td><td>5.06</td><td>67.17</td><td>2.87</td><td>73.17</td><td>5.89</td><td>70.18</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td><b>IndexTTS2.5</b></td><td>0.8B</td><td>4.36</td><td>77.10</td><td>5.12</td><td>68.06</td><td>3.75</td><td>76.39</td><td>5.66</td><td>74.62</td><td>14.88</td><td>69.74</td><td>6.75</td><td>73.18</td></tr>
<tr><td><b>IndexTTS2.5-RL</b></td><td>0.8B</td><td>3.93</td><td>77.92</td><td>3.89</td><td>67.79</td><td>3.33</td><td>76.68</td><td>5.30</td><td>75.41</td><td>13.58</td><td>70.36</td><td>6.00</td><td>73.63</td></tr>
</tbody>
</table>
**表 2: CV3-Eval における言語横断 TTS**(中国語プロンプト → ターゲット言語。アラビア語は社内テストセットを使用)。
<table>
<thead>
<tr>
<th rowspan="2">Model</th>
<th rowspan="2">Params</th>
<th colspan="2">zh→en</th>
<th colspan="2">zh→es</th>
<th colspan="2">zh→ja</th>
<th colspan="2">zh→ar</th>
<th colspan="2">Avg</th>
</tr>
<tr>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
<th>WER↓</th><th>SS↑</th>
</tr>
</thead>
<tbody>
<tr><td>VoxCPM2</td><td>2B</td><td>4.48</td><td>64.25</td><td>16.38</td><td>64.89</td><td>11.84</td><td>71.54</td><td>11.09</td><td>67.62</td><td>10.95</td><td>67.08</td></tr>
<tr><td>OmniVoice</td><td>0.8B</td><td>3.74</td><td>64.91</td><td>5.84</td><td>62.08</td><td>9.09</td><td>69.06</td><td>19.80</td><td>65.27</td><td>9.62</td><td>65.33</td></tr>
<tr><td>Moss-TTS 1.5</td><td>8B</td><td>6.13</td><td>59.23</td><td>4.32</td><td>56.63</td><td>11.52</td><td>65.54</td><td>17.03</td><td>62.93</td><td>9.75</td><td>61.08</td></tr>
<tr><td>CosyVoice3-0.5B</td><td>0.5B</td><td>3.23</td><td>62.79</td><td>4.58</td><td>64.04</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>CosyVoice3-1.5B</td><td>1.5B</td><td>4.32</td><td>-</td><td>-</td><td>-</td><td>13.70</td><td>-</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>FireRedTTS-2</td><td>1.5B</td><td>9.34</td><td>53.19</td><td>12.25</td><td>58.31</td><td>19.05</td><td>64.12</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td>Fish Audio S2 Pro</td><td>4B</td><td>4.14</td><td>55.89</td><td>4.46</td><td>55.57</td><td>10.48</td><td>61.74</td><td>14.49</td><td>59.80</td><td>8.39</td><td>58.25</td></tr>
<tr><td>Qwen3-TTS</td><td>1.7B</td><td>5.74</td><td>63.04</td><td>5.15</td><td>68.02</td><td>36.09</td><td>65.71</td><td>-</td><td>-</td><td>-</td><td>-</td></tr>
<tr><td><b>IndexTTS2.5</b></td><td>0.8B</td><td>3.62</td><td>63.83</td><td>5.17</td><td>65.48</td><td>6.57</td><td>74.16</td><td>9.51</td><td>71.02</td><td>6.22</td><td>68.62</td></tr>
<tr><td><b>IndexTTS2.5-RL</b></td><td>0.8B</td><td>3.55</td><td>67.47</td><td>4.86</td><td>64.47</td><td>6.38</td><td>75.82</td><td>9.89</td><td>73.05</td><td>6.17</td><td>70.20</td></tr>
</tbody>
</table>
## 🤝 コミュニティ & お問い合わせ
- **QQ グループ:** 663272642 (No.4), 1013410623 (No.5)
- **Discord:** https://discord.gg/uT32E7KDmy
- **メール:** indexspeech@bilibili.com
ぜひコミュニティにご参加ください!🌏 皆様のご参加・ご意見をお待ちしております!
> [!CAUTION]
> bilibili IndexTTS プロジェクトをご支援いただきありがとうございます!
> コアチームがメンテナンスしている**唯一の公式チャンネル**は [https://github.com/index-tts/index-tts](https://github.com/index-tts/index-tts) です。
> ***その他のウェブサイトやサービスは公式ではありません***。その安全性、正確性、適時性について当方は保証できません。
> 最新情報については、常にこの公式リポジトリをご参照ください。
商用利用および協業については、<u>indexspeech@bilibili.com</u> までお問い合わせください。
## 📚 引用
🌟 私たちの成果がお役に立ちましたら、スターを付け、論文を引用していただけると幸いです。
IndexTTS2.5:
```bibtex
@misc{li2026indextts25technicalreport,
title={IndexTTS 2.5 Technical Report},
author={Yunpei Li and Xun Zhou and Jinchao Wang and Lu Wang and Yong Wu and Siyi Zhou and Yiquan Zhou and Yining Wang and Yaogen Yang and Zhetao Hu and Shiyao Duan and Jiacheng Xu and Bin Xia and Jingchen Shu},
year={2026},
eprint={2601.03888},
archivePrefix={arXiv},
primaryClass={cs.SD},
url={https://arxiv.org/abs/2601.03888},
}
```
IndexTTS2:
```bibtex
@article{zhou2025indextts2,
title={IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech},
author={Siyi Zhou and Yiquan Zhou and Yi He and Xun Zhou and Jinchao Wang and Wei Deng and Jingchen Shu},
journal={arXiv preprint arXiv:2506.21619},
year={2025}
}
```
IndexTTS:
```bibtex
@article{deng2025indextts,
title={IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System},
author={Wei Deng and Siyi Zhou and Jingchen Shu and Jinchao Wang and Lu Wang},
journal={arXiv preprint arXiv:2502.05512},
year={2025},
doi={10.48550/arXiv.2502.05512},
url={https://arxiv.org/abs/2502.05512}
}
```
## 🙏 謝辞
1. [tortoise-tts](https://github.com/neonbjb/tortoise-tts)
2. [XTTSv2](https://github.com/coqui-ai/TTS)
3. [BigVGAN](https://github.com/NVIDIA/BigVGAN)
4. [wenet](https://github.com/wenet-e2e/wenet/tree/main)
5. [icefall](https://github.com/k2-fsa/icefall)
6. [maskgct](https://github.com/open-mmlab/Amphion/tree/main/models/tts/maskgct)
7. [seed-vc](https://github.com/Plachtaa/seed-vc)
## 📄 ライセンス
このプロジェクトは [bilibili モデル使用許諾契約](../LICENSE) の下で公開されています。
ご使用前に [DISCLAIMER](../DISCLAIMER) もお読みください。
+112
View File
@@ -0,0 +1,112 @@
{
"本软件以自拟协议开源, 作者不对软件具备任何控制力, 使用软件者、传播软件导出的声音者自负全责.": "هذا البرنامج مفتوح المصدر بموجب ترخيص مخصص. لا يملك المؤلف أي سيطرة على البرنامج، ويتحمل مستخدمو البرنامج وكذلك من يوزعون الأصوات المولدة بواسطته كامل المسؤولية.",
"如不认可该条款, 则不能使用或引用软件包内任何代码和文件. 详见根目录LICENSE.": "إذا كنت لا توافق على هذه الشروط، فلا يُسمح لك باستخدام أو الاستشهاد بأي كود أو ملفات داخل حزمة البرنامج. لمزيد من التفاصيل، يرجى الرجوع إلى ملفات LICENSE في الدليل الجذر.",
"时长必须为正数": "يجب أن تكون المدة رقمًا موجبًا",
"请输入有效的浮点数": "يرجى إدخال رقم عشري صالح",
"使用情感参考音频": "استخدام صوت مرجعي للمشاعر",
"使用情感向量控制": "استخدام متجهات المشاعر",
"使用情感描述文本控制": "استخدام وصف نصي للتحكم في المشاعر",
"上传情感参考音频": "رفع صوت مرجعي للمشاعر",
"情感权重": "وزن التحكم في المشاعر",
"喜": "سعيد",
"怒": "غاضب",
"哀": "حزين",
"惧": "خائف",
"厌恶": "مشمئز",
"低落": "مكتئب",
"惊喜": "متفاجئ",
"平静": "هادئ",
"情感描述文本": "وصف المشاعر",
"请输入情绪描述(或留空以自动使用目标文本作为情绪描述)": "يرجى إدخال وصف للمشاعر (أو اتركه فارغًا لاستخدام النص الرئيسي تلقائيًا)",
"高级生成参数设置": "إعدادات معلمات التوليد المتقدمة",
"情感向量之和不能超过1.5,请调整后重试。": "لا يمكن أن يتجاوز مجموع متجهات المشاعر 1.5. يرجى التعديل والمحاولة مرة أخرى.",
"时长系数": "معامل المدة",
"快": "سريع",
"慢": "بطيء",
"不变": "عادي",
"音色参考音频": "الصوت المرجعي",
"音频生成": "توليد الكلام",
"文本": "النص",
"生成语音": "توليد",
"生成结果": "نتيجة التوليد",
"功能设置": "الإعدادات",
"分句设置": "إعدادات تقسيم النص",
"参数会影响音频质量和生成速度": "تؤثر هذه المعلمات على جودة الصوت وسرعة التوليد.",
"分句最大Token数": "الحد الأقصى للرموز (Tokens) لكل مقطع توليد",
"建议80~200之间,值越大,分句越长;值越小,分句越碎;过小过大都可能导致音频质量不高": "النطاق الموصى به: 80 - 200. القيم الأكبر تتطلب ذاكرة VRAM أكبر لكنها تحسّن انسيابية الكلام، بينما تتطلب القيم الأصغر ذاكرة VRAM أقل لكنها تعني جملًا أكثر تجزؤًا. القيم الصغيرة أو الكبيرة جدًا قد تؤدي إلى كلام أقل ترابطًا.",
"预览分句结果": "معاينة مقاطع توليد الصوت",
"序号": "الرقم",
"分句内容": "المحتوى",
"Token数": "عدد الرموز",
"情感控制方式": "طريقة التحكم في المشاعر",
"GPT2 采样设置": "إعدادات أخذ عينات GPT-2",
"参数会影响音频多样性和生成速度详见": "يؤثر على تنوع الصوت المولد وسرعة التوليد. لمزيد من التفاصيل، راجع",
"是否进行采样": "تفعيل أخذ عينات GPT-2",
"生成Token最大数量,过小导致音频被截断": "الحد الأقصى لعدد الرموز المولدة. إذا تجاوز النص هذا الحد، سيتم قطع الصوت.",
"请上传情感参考音频": "يرجى رفع الصوت المرجعي للمشاعر",
"当前模型版本": "إصدار النموذج الحالي: ",
"请输入目标文本": "يرجى إدخال النص المراد توليده",
"例如:委屈巴巴、危险在悄悄逼近": "مثال: \"حزين للغاية\"، \"الخطر يقترب ببطء\"",
"与音色参考音频相同": "نفس الصوت المرجعي",
"情感随机采样": "أخذ عينات عشوائية للمشاعر",
"显示实验功能": "إظهار الميزات التجريبية",
"提示:此功能为实验版,结果尚不稳定,我们正在持续优化中。": "ملاحظة: هذه الميزة تجريبية حاليًا وقد لا تعطي نتائج مرضية. نحن نعمل على تحسين أدائها في إصدار مستقبلي.",
"自定义个别专业术语的读音": "تخصيص نطق كل مصطلح أو طريقة نطقه.",
"请至少输入一种读法": "يرجى إدخال نطق واحد على الأقل",
"已添加": "تمت الإضافة",
"开启术语词汇读音": "تفعيل نطق المصطلحات المخصص",
"添加术语": "إضافة مصطلح",
"暂无术语": "لم تتم إضافة مصطلحات مخصصة بعد",
"请输入术语": "يرجى إدخال المصطلح",
"术语": "المصطلح",
"已删除": "تم الحذف",
"英文读法": "النطق الإنجليزي",
"自定义术语词汇读音": "تخصيص نطق المصطلحات",
"中文读法": "النطق الصيني",
"词汇表已更新": "تم تحديث قاموس المصطلحات بنجاح",
"保存词汇表时出错": "خطأ أثناء حفظ قاموس المصطلحات",
"加载词汇表时出错": "خطأ أثناء تحميل قاموس المصطلحات",
"预设": "إعداد مسبق",
"预设名称": "اسم الإعداد المسبق",
"保存预设": "حفظ الإعداد المسبق",
"删除预设": "حذف الإعداد المسبق",
"请选择预设": "اختر الإعداد المسبق",
"请输入预设名称": "يرجى إدخال اسم الإعداد المسبق",
"预设已保存": "تم حفظ الإعداد المسبق",
"预设已删除": "تم حذف الإعداد المسبق",
"预设不存在": "الإعداد المسبق غير موجود",
"预设名称不能为空": "لا يمكن أن يكون اسم الإعداد المسبق فارغًا",
"预设名称已存在,已覆盖": "اسم الإعداد المسبق موجود بالفعل، تم الاستبدال",
"参考音频文件缺失,已跳过": "الصوت المرجعي مفقود، تم التخطي",
"加载预设失败": "فشل تحميل الإعداد المسبق",
"预设管理": "إدارة الإعدادات المسبقة",
"从预设加载": "تحميل من إعداد مسبق",
"加载": "تحميل",
"或": "أو",
"保存为预设": "حفظ كإعداد مسبق",
"应用": "تطبيق",
"刷新": "تحديث",
"预设详情": "تفاصيل الإعداد المسبق",
"名称": "الاسم",
"音色音频": "صوت النبرة",
"情感音频": "صوت المشاعر",
"从当前状态创建": "إنشاء من الحالة الحالية",
"创建": "إنشاء",
"无": "لا شيء",
"从预设加载音色和参数": "تحميل النبرة والمعلمات من إعداد مسبق",
"预设列表": "قائمة الإعدادات المسبقة",
"操作": "الإجراءات",
"请选择要管理的预设": "يرجى اختيار إعداد مسبق لإدارته",
"预设已应用": "تم تطبيق الإعداد المسبق",
"预设列表已刷新": "تم تحديث قائمة الإعدادات المسبقة",
"属性": "الخاصية",
"值": "القيمة",
"未知": "غير معروف",
"确认": "تأكيد",
"取消": "إلغاء",
"预设预览": "معاينة الإعداد المسبق",
"关闭": "إغلاق",
"删除": "حذف",
"情感向量": "متجه المشاعر",
"语言": "اللغة"
}
+4 -1
View File
@@ -105,5 +105,8 @@
"确认": "Confirm",
"取消": "Cancel",
"预设预览": "Preset Preview",
"关闭": "Close"
"关闭": "Close",
"删除": "Delete",
"情感向量": "Emotion Vector",
"语言": "Language"
}
+112
View File
@@ -0,0 +1,112 @@
{
"本软件以自拟协议开源, 作者不对软件具备任何控制力, 使用软件者、传播软件导出的声音者自负全责.": "Este software es de código abierto bajo una licencia personalizada. El autor no tiene control alguno sobre el software; los usuarios del software y quienes distribuyan el audio generado por él asumen toda la responsabilidad.",
"如不认可该条款, 则不能使用或引用软件包内任何代码和文件. 详见根目录LICENSE.": "Si no aceptas estos términos, no puedes usar ni hacer referencia a ningún código o archivo del paquete de software. Para más detalles, consulta el archivo LICENSE en el directorio raíz.",
"时长必须为正数": "La duración debe ser un número positivo",
"请输入有效的浮点数": "Introduce un número de punto flotante válido",
"使用情感参考音频": "Usar audio de referencia emocional",
"使用情感向量控制": "Usar vector de emociones",
"使用情感描述文本控制": "Usar texto de descripción de la emoción",
"上传情感参考音频": "Subir audio de referencia emocional",
"情感权重": "Peso del control de emoción",
"喜": "Alegre",
"怒": "Enfadado",
"哀": "Triste",
"惧": "Asustado",
"厌恶": "Disgustado",
"低落": "Melancólico",
"惊喜": "Sorprendido",
"平静": "Tranquilo",
"情感描述文本": "Texto de descripción de la emoción",
"请输入情绪描述(或留空以自动使用目标文本作为情绪描述)": "Introduce una descripción de la emoción (o déjalo en blanco para usar automáticamente el texto de destino como descripción)",
"高级生成参数设置": "Configuración avanzada de parámetros de generación",
"情感向量之和不能超过1.5,请调整后重试。": "La suma de los vectores de emociones no puede superar 1.5. Ajústala e inténtalo de nuevo.",
"时长系数": "Factor de duración",
"快": "Rápido",
"慢": "Lento",
"不变": "Normal",
"音色参考音频": "Audio de referencia de voz",
"音频生成": "Síntesis de voz",
"文本": "Texto",
"生成语音": "Sintetizar",
"生成结果": "Resultado de la síntesis",
"功能设置": "Configuración",
"分句设置": "Configuración de segmentación de texto",
"参数会影响音频质量和生成速度": "Estos parámetros afectan a la calidad del audio y a la velocidad de generación.",
"分句最大Token数": "Máximo de tokens por segmento de generación",
"建议80~200之间,值越大,分句越长;值越小,分句越碎;过小过大都可能导致音频质量不高": "Rango recomendado: 80 - 200. Valores mayores requieren más VRAM pero mejoran la fluidez del habla; valores menores requieren menos VRAM pero generan frases más fragmentadas. Valores demasiado pequeños o demasiado grandes pueden reducir la calidad del audio.",
"预览分句结果": "Vista previa de los segmentos de generación de audio",
"序号": "Índice",
"分句内容": "Contenido",
"Token数": "N.º de tokens",
"情感控制方式": "Método de control de emoción",
"GPT2 采样设置": "Configuración de muestreo de GPT-2",
"参数会影响音频多样性和生成速度详见": "Influye tanto en la diversidad del audio generado como en la velocidad de generación. Para más detalles, consulta",
"是否进行采样": "Activar muestreo de GPT-2",
"生成Token最大数量,过小导致音频被截断": "Número máximo de tokens a generar. Si el texto lo supera, el audio se cortará.",
"请上传情感参考音频": "Sube el audio de referencia emocional",
"当前模型版本": "Versión actual del modelo: ",
"请输入目标文本": "Introduce el texto a sintetizar",
"例如:委屈巴巴、危险在悄悄逼近": "p. ej. \"profundamente triste\", \"el peligro se acerca sigilosamente\"",
"与音色参考音频相同": "Igual que la referencia de voz",
"情感随机采样": "Muestreo aleatorio de emoción",
"显示实验功能": "Mostrar funciones experimentales",
"提示:此功能为实验版,结果尚不稳定,我们正在持续优化中。": "Nota: esta función es actualmente experimental y puede no producir resultados satisfactorios. Trabajamos continuamente para mejorarla en una futura versión.",
"自定义个别专业术语的读音": "Personaliza la pronunciación de cada término o cómo se \"lee en voz alta\".",
"请至少输入一种读法": "Introduce al menos una pronunciación",
"已添加": "Añadido",
"开启术语词汇读音": "Activar pronunciaciones personalizadas de términos",
"添加术语": "Añadir término",
"暂无术语": "Aún no hay términos personalizados",
"请输入术语": "Introduce el término",
"术语": "Término",
"已删除": "Eliminado",
"英文读法": "Pronunciación en inglés",
"自定义术语词汇读音": "Personalizar pronunciaciones de términos",
"中文读法": "Pronunciación en chino",
"词汇表已更新": "Glosario actualizado correctamente",
"保存词汇表时出错": "Error al guardar el glosario",
"加载词汇表时出错": "Error al cargar el glosario",
"预设": "Preajuste",
"预设名称": "Nombre del preajuste",
"保存预设": "Guardar preajuste",
"删除预设": "Eliminar preajuste",
"请选择预设": "Seleccionar preajuste",
"请输入预设名称": "Introduce un nombre de preajuste",
"预设已保存": "Preajuste guardado",
"预设已删除": "Preajuste eliminado",
"预设不存在": "El preajuste no existe",
"预设名称不能为空": "El nombre del preajuste no puede estar vacío",
"预设名称已存在,已覆盖": "El nombre del preajuste ya existe, se ha sobrescrito",
"参考音频文件缺失,已跳过": "Audio de referencia no encontrado, se ha omitido",
"加载预设失败": "Error al cargar el preajuste",
"预设管理": "Gestión de preajustes",
"从预设加载": "Cargar desde preajuste",
"加载": "Cargar",
"或": "o",
"保存为预设": "Guardar como preajuste",
"应用": "Aplicar",
"刷新": "Actualizar",
"预设详情": "Detalles del preajuste",
"名称": "Nombre",
"音色音频": "Audio de voz",
"情感音频": "Audio de emoción",
"从当前状态创建": "Crear desde el estado actual",
"创建": "Crear",
"无": "Ninguno",
"从预设加载音色和参数": "Cargar voz y parámetros desde un preajuste",
"预设列表": "Lista de preajustes",
"操作": "Acciones",
"请选择要管理的预设": "Selecciona un preajuste para gestionar",
"预设已应用": "Preajuste aplicado",
"预设列表已刷新": "Lista de preajustes actualizada",
"属性": "Propiedad",
"值": "Valor",
"未知": "Desconocido",
"确认": "Confirmar",
"取消": "Cancelar",
"预设预览": "Vista previa del preajuste",
"关闭": "Cerrar",
"删除": "Eliminar",
"情感向量": "Vector de emociones",
"语言": "Idioma"
}
+112
View File
@@ -0,0 +1,112 @@
{
"本软件以自拟协议开源, 作者不对软件具备任何控制力, 使用软件者、传播软件导出的声音者自负全责.": "本ソフトウェアは独自ライセンスの下でオープンソースとして公開されています。作者はソフトウェアに対していかなる制御力も持たず、ソフトウェアの利用者および生成された音声を配布する者が全責任を負います。",
"如不认可该条款, 则不能使用或引用软件包内任何代码和文件. 详见根目录LICENSE.": "これらの条項に同意されない場合、ソフトウェアパッケージ内のいかなるコードやファイルも使用または引用することはできません。詳細はルートディレクトリの LICENSE をご参照ください。",
"时长必须为正数": "長さは正の数である必要があります",
"请输入有效的浮点数": "有効な浮動小数点数を入力してください",
"使用情感参考音频": "感情参照音声を使用",
"使用情感向量控制": "感情ベクトルで制御",
"使用情感描述文本控制": "感情記述テキストで制御",
"上传情感参考音频": "感情参照音声をアップロード",
"情感权重": "感情制御の重み",
"喜": "喜び",
"怒": "怒り",
"哀": "悲しみ",
"惧": "恐れ",
"厌恶": "嫌悪",
"低落": "憂鬱",
"惊喜": "驚き",
"平静": "穏やか",
"情感描述文本": "感情記述テキスト",
"请输入情绪描述(或留空以自动使用目标文本作为情绪描述)": "感情記述を入力してください(空欄の場合は対象テキストが自動的に感情記述として使用されます)",
"高级生成参数设置": "高度な生成パラメータ設定",
"情感向量之和不能超过1.5,请调整后重试。": "感情ベクトルの合計は 1.5 を超えることはできません。調整してから再度お試しください。",
"时长系数": "長さ係数",
"快": "速い",
"慢": "遅い",
"不变": "通常",
"音色参考音频": "音色参照音声",
"音频生成": "音声生成",
"文本": "テキスト",
"生成语音": "音声を生成",
"生成结果": "生成結果",
"功能设置": "機能設定",
"分句设置": "文分割設定",
"参数会影响音频质量和生成速度": "これらのパラメータは音声品質と生成速度に影響します",
"分句最大Token数": "文分割あたりの最大トークン数",
"建议80~200之间,值越大,分句越长;值越小,分句越碎;过小过大都可能导致音频质量不高": "80〜200 の範囲を推奨します。値が大きいほど分割される文が長くなり、値が小さいほど文が細かく分割されます。小さすぎても大きすぎても音声品質が低下する可能性があります",
"预览分句结果": "文分割結果のプレビュー",
"序号": "番号",
"分句内容": "分割内容",
"Token数": "トークン数",
"情感控制方式": "感情制御方法",
"GPT2 采样设置": "GPT-2 サンプリング設定",
"参数会影响音频多样性和生成速度详见": "これらのパラメータは生成音声の多様性と生成速度の両方に影響します。詳細はこちらを参照:",
"是否进行采样": "GPT-2 サンプリングを有効化",
"生成Token最大数量,过小导致音频被截断": "生成するトークンの最大数。小さすぎると音声が途中で切れます",
"请上传情感参考音频": "感情参照音声をアップロードしてください",
"当前模型版本": "現在のモデルバージョン:",
"请输入目标文本": "合成するテキストを入力してください",
"例如:委屈巴巴、危险在悄悄逼近": "例:「深い悲しみ」「危険が忍び寄っている」",
"与音色参考音频相同": "音色参照音声と同じ",
"情感随机采样": "感情のランダムサンプリング",
"显示实验功能": "実験的機能を表示",
"提示:此功能为实验版,结果尚不稳定,我们正在持续优化中。": "注意:この機能は実験版であり、結果が安定しない場合があります。継続的に改善に取り組んでいます。",
"自定义个别专业术语的读音": "各用語の読み方や発音をカスタマイズ",
"请至少输入一种读法": "少なくとも 1 つの読み方を入力してください",
"已添加": "追加済み",
"开启术语词汇读音": "カスタム用語の発音を有効化",
"添加术语": "用語を追加",
"暂无术语": "カスタム用語はまだ追加されていません",
"请输入术语": "用語を入力してください",
"术语": "用語",
"已删除": "削除済み",
"英文读法": "英語の読み方",
"自定义术语词汇读音": "用語の発音をカスタマイズ",
"中文读法": "中国語の読み方",
"词汇表已更新": "用語集が正常に更新されました",
"保存词汇表时出错": "用語集の保存中にエラーが発生しました",
"加载词汇表时出错": "用語集の読み込み中にエラーが発生しました",
"预设": "プリセット",
"预设名称": "プリセット名",
"保存预设": "プリセットを保存",
"删除预设": "プリセットを削除",
"请选择预设": "プリセットを選択",
"请输入预设名称": "プリセット名を入力してください",
"预设已保存": "プリセットが保存されました",
"预设已删除": "プリセットが削除されました",
"预设不存在": "プリセットが存在しません",
"预设名称不能为空": "プリセット名は空にできません",
"预设名称已存在,已覆盖": "プリセット名はすでに存在するため、上書きされました",
"参考音频文件缺失,已跳过": "参照音声ファイルが見つからないため、スキップされました",
"加载预设失败": "プリセットの読み込みに失敗しました",
"预设管理": "プリセット管理",
"从预设加载": "プリセットから読み込む",
"加载": "読み込む",
"或": "または",
"保存为预设": "プリセットとして保存",
"应用": "適用",
"刷新": "更新",
"预设详情": "プリセット詳細",
"名称": "名前",
"音色音频": "音色音声",
"情感音频": "感情音声",
"从当前状态创建": "現在の状態から作成",
"创建": "作成",
"无": "なし",
"从预设加载音色和参数": "プリセットから音色とパラメータを読み込む",
"预设列表": "プリセット一覧",
"操作": "操作",
"请选择要管理的预设": "管理するプリセットを選択してください",
"预设已应用": "プリセットが適用されました",
"预设列表已刷新": "プリセット一覧が更新されました",
"属性": "プロパティ",
"值": "値",
"未知": "不明",
"确认": "確認",
"取消": "キャンセル",
"预设预览": "プリセットプレビュー",
"关闭": "閉じる",
"删除": "削除",
"情感向量": "感情ベクトル",
"语言": "言語"
}