Jorge Zozaya

yozozaya

AI & ML interests

LLM's and music generation for inspiration.

Recent Activity

liked a Space about 1 hour ago
moondream/gaze-demo
liked a Space about 1 hour ago
vikhyatk/contemplative-moondream
liked a Space 20 days ago
vikhyatk/moondream2
View all activity

Organizations

Montalvo Wire's profile picture

yozozaya's activity

reacted to manu's post with โค๏ธ 12 months ago
view post
Post
These past months, I've been busy baking a special sort of Croissant ๐Ÿฅ with an awesome team !

๐Ÿฅ CroissantLLM is a truly bilingual language model trained on 3 trillion tokens of French and English data. In its size category (<2B), it is the best model in French, but it also rivals the best monolingual English models !

๐Ÿ’พ To train it, we collected, filtered and cleaned huge quantities of permissively licensed French data, across various domains (legal, administrative, cultural, scientific), and different text modalities (speech transcriptions, movie subtitles, encyclopedias, forums, webpages)...

โš–๏ธ Assessing LLM performance is not easy, especially outside of English, and to this end we crafted a novel evaluation benchmark, FrenchBench, aiming to assess reasoning, factual knowledge, and linguistic capabilities of models in French !

๐Ÿ”Ž The best current LLMs are hidden behind a shroud of mystery, trained with undisclosed training data mixes or strategies. We go the opposite way, releasing all of the project's artefacts (model checkpoints, data, training details, evaluation benchmarks...) We obtain 81 % of the Stanford FMTI transparency criterias, far ahead of even most open initiatives !

๐ŸงชBeyond a powerful industrial resource, our transparent initiative is a stepping stone for many scientific questions ! How does teaching a model two languages instead of one splits its monolingual ability ? Does training on so much French help the model integrate French-centric knowledge and cultural biases ? How does the model memorize the training data ?

Many more things to say, for those interested, I recommend checking out:

๐Ÿ—ž๏ธ The blogpost: https://huggingface.co/blog/manu/croissant-llm-blog
๐Ÿ“– The 45 page report with lots of gems: https://arxiv.org/abs/2402.00786
๐Ÿค– Models, Data, Demo: https://huggingface.co/croissantllm
ยท
updated a Space 12 months ago