PereLluis13
/

Wav2Vec2-Large-XLSR-53-catalan

@@ -28,7 +28,7 @@ model-index:
 # Wav2Vec2-Large-XLSR-53-ca
-Fine-tuned [facebook/wav2vec2-large-xlsr-53](https://huggingface.co/facebook/wav2vec2-large-xlsr-53) on catalan using the [Common Voice](https://huggingface.co/datasets/common_voice)dataset.
 When using this model, make sure that your speech input is sampled at 16kHz.
 ## Usage
@@ -113,15 +113,32 @@ def evaluate(batch):
 	return batch
 result = test_dataset.map(evaluate, batched=True, batch_size=8)
-print("WER: {:2f}".format(100 * wer.compute(predictions=result["pred_strings"], references=result["sentence"])))
 ```
-**Test Result**: XX.XX %  # TODO: write output of print here. IMPORTANT: Please remember to also replace {wer_result_on_test} at the top of with this value here. tags.
 ## Training
-The Common Voice `train`, `validation`, and ... datasets were used for training as well as ... and ...  # TODO: adapt to state all the datasets that were used for training.
-The script used for training can be found [here](...) # TODO: fill in a link to your training script here. If you trained your model in a colab, simply fill in the link here. If you trained the model locally, it would be great if you could upload the training script on github and paste the link here.

 # Wav2Vec2-Large-XLSR-53-ca
+Fine-tuned [facebook/wav2vec2-large-xlsr-53](https://huggingface.co/facebook/wav2vec2-large-xlsr-53) on catalan using the [Common Voice](https://huggingface.co/datasets/common_voice) dataset.
 When using this model, make sure that your speech input is sampled at 16kHz.
 ## Usage
 	return batch
 result = test_dataset.map(evaluate, batched=True, batch_size=8)
+import jiwer
+# Chunk WER computation due to memory issues, taken from https://huggingface.co/pcuenq/wav2vec2-large-xlsr-53-es
+def chunked_wer(targets, predictions, chunk_size=None):
+	if chunk_size is None: return jiwer.wer(targets, predictions)
+	start = 0
+	end = chunk_size
+	H, S, D, I = 0, 0, 0, 0
+	while start < len(targets):
+		chunk_metrics = jiwer.compute_measures(targets[start:end], predictions[start:end])
+		H = H + chunk_metrics["hits"]
+		S = S + chunk_metrics["substitutions"]
+		D = D + chunk_metrics["deletions"]
+		I = I + chunk_metrics["insertions"]
+		start += chunk_size
+		end += chunk_size
+	return float(S + D + I) / float(H + S + D)
+print("WER: {:2f}".format(100 * chunked_wer(result["sentence"], result["pred_strings"], chunk_size=4000)))
 ```
+**Test Result**: 15.20 %  # TODO: write output of print here. IMPORTANT: Please remember to also replace {wer_result_on_test} at the top of with this value here. tags.
 ## Training
+The Common Voice `train`, `validation` datasets were used for training. At the second epoch training was halted due to a memory issue, and was continued with lower batch size, but acc. gradient steps were scaled to keep it at 32 batch size during all training.
+The script used for training can be found [here](https://github.com/huggingface/transformers/blob/master/examples/research_projects/wav2vec2/run_common_voice.py). Slight modifications were done in order to speed up the ordering by length during training, which can be found [here](https://discuss.huggingface.co/t/spanish-asr-fine-tuning-wav2vec2/4586/6). Another version trained for catalan can be found [here](https://huggingface.co/ccoreilly/wav2vec2-large-xlsr-catala), which may be better than this one since it was trained with extra data and for longer time. Whoever, since it used different splits that include part of the Common Voice test set, this version can be used to get a baseline on the Common Voice dataset.

pytorch_model.bin CHANGED Viewed

@@ -1,3 +1,3 @@
 version https://git-lfs.github.com/spec/v1
-oid sha256:09a799507e51872c6c4be7ea8e404b5b3613a30c806a9909897b14cc87f934a2
 size 1262282327

 version https://git-lfs.github.com/spec/v1
+oid sha256:ad46cc2cbbcb1cef52c8bc48c61155e79d08e54afa3a6385ea8baf54dccb7bfa
 size 1262282327