Occiglot-7B-FR-EN

A polyglot language model for the Occident.

Occiglot-7B-FR-EN is a generative language model with 7B parameters for French and English and trained by the Occiglot Research Collective. It is based on Mistral-7B-v0.1 and trained on 113B tokens of additional multilingual and code data with a block size of 8,192 tokens per sample. Note that the model is a general-purpose base model and was not instruction-fine-tuned nor optimized for chat or other applications. We make an instruction tuned variant available as occiglot-7b-fr-en-instruct

This is the first release of an ongoing open research project for multilingual language models. If you want to train a model for your own language or are working on evaluations, please contact us or join our Discord server. We are open for collaborations!

Model details

Continued-pretraining from: Mistral-7B-v0.1
Model type: Causal decoder-only transformer language model
Languages: English, French, and code.
License: Apache 2.0
Compute resources: HessianAI's 42
Contributors: Manuel Brack, Patrick Schramowski, Pedro Ortiz, Malte Ostendorff, Fabio Barth, Georg Rehm, Kristian Kersting
Research labs: Occiglot with support from SAINT and SLT
Contact: Discord

How to use

You can use this model directly with a pipeline for text generation. Since the generation relies on some randomness, we set a seed for reproducibility:

>>> from transformers import pipeline, set_seed
>>> generator = pipeline('text-generation', model='occiglot/occiglot-7b-fr-en')
>>> set_seed(42)
>>> generator("Bonjour, Je suis un modèle linguistique,", max_length=40, num_return_sequences=1)
[{'generated_text': 'Bonjour, Je suis un modèle linguistique qui peut t'aider à traduire des textes entre le français et l'anglais. Si tu me donnes un texte en français'}]

Dataset

The training data is the respective subset of the data used for occiglot-7b-eu5, i.e. French plus English and Code.

The data distribution by language (estimated) is as follows:

English: ~34%
Code: ~13%
French: ~52%

The training data was prepared using lm-datasets. The exact data configuration is here.

Training settings

Continual pre-training on 128 x A100-80GB on HessianAI's 42.
Framework: Determined
Precision: bf16
Optimizer: AdamW (lr: 0.00001, warmup_steps: 420)
Global batch size: 512 (with 8192 blocksize) split over 128 GPUs
Cosine Annealing with Warmup

Tokenizer

Tokenizer is unchanged from Mistral-7B-v0.1.

Evaluation

Preliminary evaluation results can be found below. Please note that the non-English results are based on partially machine-translated datasets and English prompts (Belebele and Okapi framework) and thus should be interpreted with caution, e.g., biased towards English model performance. Currently, we are working on more suitable benchmarks for Spanish, French, German, and Italian.

Evaluation results

All 5 Languages

	avg	arc_challenge	belebele	hellaswag	mmlu	truthfulqa
Occiglot-7b-eu5	0.516895	0.508109	0.675556	0.718963	0.402064	0.279782
Occiglot-7b-eu5-instruct	0.537799	0.53632	0.691111	0.731918	0.405198	0.32445
Occiglot-7b-fr-en	0.509209	0.496806	0.691333	0.667475	0.409129	0.281303
Occiglot-7b-fr-en-instruct	0.52884	0.515613	0.723333	0.67371	0.413024	0.318521
Claire-mistral-7b-0.1	0.514226	0.502773	0.705111	0.666871	0.412128	0.284245
Mistral-7b-v0.1	0.547111	0.528937	0.768444	0.682516	0.448253	0.307403
Mistral-7b-instruct-v0.2	0.56713	0.547228	0.741111	0.69455	0.422501	0.430262

English

	avg	arc_challenge	belebele	hellaswag	mmlu	truthfulqa
Occiglot-7b-eu5	0.59657	0.530717	0.726667	0.789882	0.531904	0.403678
Occiglot-7b-eu5-instruct	0.617905	0.558874	0.746667	0.799841	0.535109	0.449
Occiglot-7b-fr-en	0.621947	0.568259	0.771111	0.804919	0.570716	0.394726
Occiglot-7b-fr-en-instruct	0.646571	0.586177	0.794444	0.808305	0.569862	0.474064
Claire-mistral-7b-0.1	0.651798	0.59727	0.817778	0.827126	0.600912	0.415906
Mistral-7b-v0.1	0.668385	0.612628	0.844444	0.834097	0.624555	0.426201
Mistral-7b-instruct-v0.2	0.713657	0.637372	0.824444	0.846345	0.59201	0.668116

French

	avg	arc_challenge_fr	belebele_fr	hellaswag_fr	mmlu_fr	truthfulqa_fr
Occiglot-7b-eu5	0.525017	0.506416	0.675556	0.712358	0.495684	0.23507
Occiglot-7b-eu5-instruct	0.554216	0.541488	0.7	0.724245	0.499122	0.306226
Occiglot-7b-fr-en	0.542903	0.532934	0.706667	0.718891	0.51333	0.242694
Occiglot-7b-fr-en-instruct	0.567079	0.542344	0.752222	0.72553	0.52051	0.29479
Claire-mistral-7b-0.1	0.515127	0.486741	0.694444	0.642964	0.479566	0.271919
Mistral-7b-v0.1	0.558129	0.525235	0.776667	0.66481	0.543121	0.280813
Mistral-7b-instruct-v0.2	0.575821	0.551754	0.758889	0.67916	0.506837	0.382465

Acknowledgements

The model training was supported by a compute grant at the 42 supercomputer which is a central component in the development of hessian AI, the AI Innovation Lab (funded by the Hessian Ministry of Higher Education, Research and the Art (HMWK) & the Hessian Ministry of the Interior, for Security and Homeland Security (HMinD)) and the AI Service Centers (funded by the German Federal Ministry for Economic Affairs and Climate Action (BMWK)). The curation of the training data is partially funded by the German Federal Ministry for Economic Affairs and Climate Action (BMWK) through the project OpenGPT-X (project no. 68GX21007D).

License

Apache 2.0

occiglot
/

occiglot-7b-fr-en