about multilingual label

#8
by StatPan - opened

It is stated to be a multilingual model, but there is no specific mention of which languages are included, so I am curious.

Can you share which languages are included and to what extent for each country?

Microsoft org

Thank you for your interest! The Phi models are trained primarily on English text. In the Small and Medium we have some multilingual training data. We see the Small and Medium models are doing well on some languages such as German, Spanish, French, Japanese, Russian, Chinese.

did you have any Vietnamese?

Thank you for the clarification on the multilingual capabilities of the Phi models. It's great to know they perform well in several languages. I appreciate your help!

Microsoft org

Small and Medium have some Vietnamese data, but the Phi model family are trained primarily on English. We are eager to learn more from the community on language-specific finetuning experiments.

Sign up or log in to comment