utter-project
/

mHuBERT-147-base-2nd-iter

@@ -123,33 +123,54 @@ language:
 - zh
 ---
-## mHuBERT-147 models
-mHuBERT-147 are compact and competitive multilingual HuBERT models trained on 90K hours of open-license data in 147 languages.
-This repository contains:
 * Fairseq checkpoint (original);
-* HuggingFace checkpoint;
 * Faiss index for continuous pre-training (OPQ16_64,IVF1000_HNSW32,PQ16x4fsr).
-# Additional Information
-**Manifest list:** https://huggingface.co/utter-project/mHuBERT-147-base-3rd-iter/tree/main/manifest
-Please note that since training, there were CommonVoice removal requests. This means that some of the listed files are no longer available.
-**Fairseq fork:** https://github.com/utter-project/fairseq
-**Scripts for pre-processing/faiss clustering:** https://github.com/utter-project/mHuBERT-147-scripts
-**Languages present not indexed by Huggingface:** Asturian (ast), Basaa (bas), Cebuano (ceb), Central Kurdish/Sorani (ckb), Hakha Chin (cnh), Hawaiian (haw), Upper Sorbian (hsb) Kabyle (kab), Moksha (mdf), Meadow Mari (mhr), Hill Mari (mrj), Erzya (myv), Taiwanese Hokkien (nan-tw), Sursilvan (rm-sursilv), Vallader (rm-vallader), Sakha (sah), Santali (sat), Scots (sco), Saraiki (skr), Tigre (tig), Tok Pisin (tpi), Akwapen Twi (tw-akuapem), Asante Twi (tw-asante), Votic (vot), Waray (war), Cantonese (yue).
-# Datasets Included
-For ASR/ST/TTS datasets, only train set is used.
 * [Aishell](https://www.openslr.org/33/) and [AISHELL-3](https://www.openslr.org/93/)
 * [BibleTTS](https://www.openslr.org/129/)
 * [ClovaCall](https://github.com/clovaai/ClovaCall)
@@ -166,8 +187,10 @@ For ASR/ST/TTS datasets, only train set is used.
 * [VoxLingua107](https://bark.phon.ioc.ee/voxlingua107/)
 * [VoxPopuli](https://github.com/facebookresearch/voxpopuli/)
-# Citing
 ```
 @inproceedings{boito2024mhubert,
@@ -178,9 +201,6 @@ booktitle={Interspeech 2024},
 }
 ```
-# Funding
 <img src="https://cdn-uploads.huggingface.co/production/uploads/62262e19d36494a6f743a28d/HbzC1C-uHe25ewTy2wyoK.png" width=7% height=7%>
 This is an output of the European Project UTTER (Unified Transcription and Translation for Extended Reality) funded by European Union’s Horizon Europe Research and Innovation programme under grant agreement number 101070631.

 - zh
 ---
+**This repository contains the SECOND ITERATION mHuBERT-147 model.**
+**The best mHuBERT-147 model is available [here](https://huggingface.co/utter-project/mHuBERT-147).**
+**MODEL DETAILS:** 2nd iteration, K=1000, HuBERT base architecture (95M parameters), 147 languages.
+# Table of Contents:
+1. [Summary](https://huggingface.co/utter-project/mHuBERT-147#mhubert-147-models)
+2. [Training Data and Code](https://huggingface.co/utter-project/mHuBERT-147#training)
+3. [ML-SUPERB Scores](https://huggingface.co/utter-project/mHuBERT-147#ml-superb-scores)
+4. [Languages and Datasets](https://huggingface.co/utter-project/mHuBERT-147#languages-and-datasets)
+6. [Citing and Funding Information](https://huggingface.co/utter-project/mHuBERT-147#citing-and-funding-information)
+# mHuBERT-147 models
+mHuBERT-147 are compact and competitive multilingual HuBERT models trained on 90K hours of open-license data in 147 languages.
+Different from *traditional* HuBERTs, mHuBERT-147 models are trained using faiss IVF discrete speech units.
+Training employs a two-level language, data source up-sampling during training. See more information in [our paper](https://arxiv.org/pdf/2406.06371).
+**This repository contains:**
 * Fairseq checkpoint (original);
+* HuggingFace checkpoint (conversion using transformers library);
 * Faiss index for continuous pre-training (OPQ16_64,IVF1000_HNSW32,PQ16x4fsr).
+**Related Models:**
+* [3rd Iteration mHuBERT-147](https://huggingface.co/utter-project/mHuBERT-147) (best)
+* [1st Iteration mHuBERT-147](https://huggingface.co/utter-project/mHuBERT-147-base-1st-iter)
+* [HUTTER-12 CommonVoice Prototype (12 languages)](https://huggingface.co/utter-project/hutter-12-3rd-base)
+# Training
+* **[Manifest list available here.](https://huggingface.co/utter-project/mHuBERT-147-base-3rd-iter/tree/main/manifest)** Please note that since training, there were CommonVoice removal requests. This means that some of the listed files are no longer available.
+* **[Fairseq fork](https://github.com/utter-project/fairseq)** contains the scripts for training with multilingual batching with two-level up-sampling.
+* **[Scripts for pre-processing/faiss clustering available here.](https://github.com/utter-project/mHuBERT-147-scripts)**
+# ML-SUPERB Scores
+mHubert-147 reaches second and first position in the 10min and 1h leaderboards respectively. We achieve new SOTA scores for three LID tasks.
+See more information in [our paper](https://arxiv.org/pdf/2406.06371).
+![image/png](https://cdn-uploads.huggingface.co/production/uploads/62262e19d36494a6f743a28d/chXjExnWc3rhhtdsyiU-W.png)
+# Languages and Datasets
+**Datasets:** For ASR/ST/TTS datasets, only train set is used.
 * [Aishell](https://www.openslr.org/33/) and [AISHELL-3](https://www.openslr.org/93/)
 * [BibleTTS](https://www.openslr.org/129/)
 * [ClovaCall](https://github.com/clovaai/ClovaCall)
 * [VoxLingua107](https://bark.phon.ioc.ee/voxlingua107/)
 * [VoxPopuli](https://github.com/facebookresearch/voxpopuli/)
+**Languages present not indexed by Huggingface:** Asturian (ast), Basaa (bas), Cebuano (ceb), Central Kurdish/Sorani (ckb), Hakha Chin (cnh), Hawaiian (haw), Upper Sorbian (hsb) Kabyle (kab), Moksha (mdf), Meadow Mari (mhr), Hill Mari (mrj), Erzya (myv), Taiwanese Hokkien (nan-tw), Sursilvan (rm-sursilv), Vallader (rm-vallader), Sakha (sah), Santali (sat), Scots (sco), Saraiki (skr), Tigre (tig), Tok Pisin (tpi), Akwapen Twi (tw-akuapem), Asante Twi (tw-asante), Votic (vot), Waray (war), Cantonese (yue).
+# Citing and Funding Information
 ```
 @inproceedings{boito2024mhubert,
 }
 ```
 <img src="https://cdn-uploads.huggingface.co/production/uploads/62262e19d36494a6f743a28d/HbzC1C-uHe25ewTy2wyoK.png" width=7% height=7%>
 This is an output of the European Project UTTER (Unified Transcription and Translation for Extended Reality) funded by European Union’s Horizon Europe Research and Innovation programme under grant agreement number 101070631.