uploading model files

Browse files

Files changed (15) hide show

README.md +143 -0
config.json +26 -0
generation_config.json +6 -0
lm-datasets-config.yml +398 -0
model-00001-of-00006.safetensors +3 -0
model-00002-of-00006.safetensors +3 -0
model-00003-of-00006.safetensors +3 -0
model-00004-of-00006.safetensors +3 -0
model-00005-of-00006.safetensors +3 -0
model-00006-of-00006.safetensors +3 -0
model.safetensors.index.json +298 -0
special_tokens_map.json +5 -0
tokenizer.json +0 -0
tokenizer.model +3 -0
tokenizer_config.json +42 -0

README.md CHANGED Viewed

@@ -1,3 +1,146 @@
 ---
 license: apache-2.0
 ---

 ---
 license: apache-2.0
+language:
+- en
+- es
+- de
+- fr
+- it
+pipeline_tag: text-generation
 ---
+![image/png](https://huggingface.co/datasets/malteos/images/resolve/main/occiglot.medium.png)
+# Occiglot-7B-EU5
+> A [polyglot](https://en.wikipedia.org/wiki/Multilingualism#In_individuals) language model for the [Occident](https://en.wikipedia.org/wiki/Occident).
+>
+**Occiglot-7B-EU5** is a generative language model with 7B parameters supporting the top-5 EU languages (English, Spanish, French, German, and Italian) and trained by the [German Research Center for Artificial Intelligence (DFKI)](https://www.dfki.de/en/web).
+It is based on [Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1) and trained on 293B tokens of additional multilingual and code data with a block size of 8,192 tokens per sample.
+Note that the model is a general-purpose base model and was not instruction-fine-tuned nor optimized for chat or other applications.
+This is the first release of an ongoing open research project for multilingual language models.
+If you want to train a model for your own language or are working on evaluations, please contact us. **We are open for collaborations!**
+### Model details
+- **Continued-pretraining from:** [Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1)
+- **Model type:** Causal decoder-only transformer language model
+- **Languages:** English, Spanish, French, German, Italian, and code.
+- **License:** [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0.html)
+- **Compute resources:** [HessianAI's 42](https://hessian.ai/)
+- **Contributors:** Manuel Brack, Patrick Schramowski, Pedro Ortiz, Malte Ostendorff, Fabio Barth, Georg Rehm, Kristian Kersting
+- **Research labs:** [SAINT](https://www.dfki.de/en/web/research/research-departments/foundations-of-systems-ai) and [SLT](https://www.dfki.de/en/web/research/research-departments/speech-and-language-technology)
+### How to use
+You can use this model directly with a pipeline for text generation. Since the generation relies on some randomness, we
+set a seed for reproducibility:
+```python
+>>> from transformers import pipeline, set_seed
+>>> generator = pipeline('text-generation', model='occiglot/occiglot-7b-eu5')
+>>> set_seed(42)
+>>> generator("Hallo, Ich bin ein Sprachmodell,", max_length=40, num_return_sequences=1)
+[{'generated_text': 'Hallo, Ich bin ein Sprachmodell, das dir bei der Übersetzung von Texten zwischen Deutsch und Englisch helfen kann. Wenn du mir einen Text in Deutsch'}]
+```
+## Dataset
+The training data was split amongst the 4 target languages (de, es, fr, it) and the continuous training in English and code.
+The data distribution by language (estimated) is as follows:
+- English: ~13%
+- Code: ~5%
+- German: ~20%
+- Spanish: ~20%
+- French: ~20%
+- Italian: ~20%
+The training data was prepared using [lm-datasets](https://github.com/malteos/lm-datasets).
+The exact data config will be released soon.
+## Training settings
+- Continual pre-training on 128 x A100-80GB on [HessianAI's 42](https://hessian.ai/).
+- Framework: [Determined](https://www.determined.ai/)
+- Precision: bf16
+- Optimizer: AdamW (lr: 0.00001, warmup_steps: 420)
+- Global batch size: 512 (with 8192 blocksize) split over 128 GPUs
+- Cosine Annealing with Warmup
+## Tokenizer
+Tokenizer is unchanged from [Mistral-7B-v0.1](https://huggingface.co/mistralai/Mistral-7B-v0.1).
+## Evaluation
+Preliminary evaluation results can be found below.
+Please note that the non-English results are based on partially machine-translated datasets and English prompts ([Belebele](https://huggingface.co/datasets/facebook/belebele) and [Okapi framework](https://github.com/nlp-uoregon/Okapi)) and thus should be interpreted with caution, e.g., biased towards English model performance.
+Currently, we are working on more suitable benchmarks for Spanish, French, German, and Italian.
+### All languages
+| **model_name**           | **arc_challenge** | **hellaswag** | **belebele** | **mmlu** | **avg** |
+|--------------------------|-------------------|---------------|--------------|----------|---------|
+| Mistral-7B-v0.1          |            0.5277 |        0.6825 |       0.7687 |   0.6287 |  0.6519 |
+| leo-mistral-hessianai-7b |            0.4614 |        0.6423 |       0.6524 |   0.5440 |  0.5750 |
+| Occiglot-7B-EU5          |            0.5083 |        0.7191 |       0.6758 |   0.5432 |  0.6116 |
+### English
+| **model_name**           | **arc_challenge** | **hellaswag** | **belebele** | **mmlu** | **avg** |
+|--------------------------|-------------------|---------------|--------------|----------|---------|
+| Mistral-7B-v0.1          |            0.6143 |        0.8344 |       0.8444 |   0.6351 |  0.7321 |
+| leo-mistral-hessianai-7b |            0.5213 |        0.7779 |       0.7356 |   0.5508 |  0.6464 |
+| Occiglot-7B-EU5          |            0.5307 |        0.7900 |       0.7267 |   0.5467 |  0.6485 |
+### German
+| **model_name**           | **arc_challenge** | **hellaswag** | **belebele** | **mmlu** | **avg** |
+|--------------------------|-------------------|---------------|--------------|----------|---------|
+| Mistral-7B-v0.1          |            0.4765 |        0.6101 |       0.7411 |   0.5274 |  0.5888 |
+| leo-mistral-hessianai-7b |            0.4739 |        0.6818 |       0.6900 |   0.4887 |  0.5836 |
+| Occiglot-7B-EU5          |            0.4944 |        0.6667 |       0.6467 |   0.4833 |  0.5728 |
+### Spanish
+| **model_name**           | **arc_challenge** | **hellaswag** | **belebele** | **mmlu** | **avg** |
+|--------------------------|-------------------|---------------|--------------|----------|---------|
+| Mistral-7B-v0.1          |            0.5256 |        0.6728 |       0.7478 |   0.5432 |  0.6224 |
+| leo-mistral-hessianai-7b |            0.4436 |        0.5970 |       0.6178 |   0.4359 |  0.5236 |
+| Occiglot-7B-EU5          |            0.5085 |        0.7255 |       0.6778 |   0.4997 |  0.6029 |
+### French
+| **model_name**           | **arc_challenge** | **hellaswag** | **belebele** | **mmlu** | **avg** |
+|--------------------------|-------------------|---------------|--------------|----------|---------|
+| Mistral-7B-v0.1          |            0.5244 |        0.6651 |       0.7744 |   0.5413 |  0.6263 |
+| leo-mistral-hessianai-7b |            0.4354 |        0.5967 |       0.6222 |   0.4326 |  0.5217 |
+| Occiglot-7B-EU5          |            0.5064 |        0.7125 |       0.6756 |   0.4959 |  0.5976 |
+### Italian
+| **model_name**           | **arc_challenge** | **hellaswag** | **belebele** | **mmlu** | **avg** |
+|--------------------------|-------------------|---------------|--------------|----------|---------|
+| Mistral-7B-v0.1          |            0.4979 |        0.6303 |       0.7356 |   0.5372 |  0.6002 |
+| leo-mistral-hessianai-7b |            0.4328 |        0.5580 |       0.5967 |   0.4311 |  0.5047 |
+| Occiglot-7B-EU5 |            0.5013 |        0.7008 |       0.6522 |   0.4949 |  0.5873 |
+## Acknowledgements
+The model training was supported by a compute grant at the [42 supercomputer](https://hessian.ai/)  which is a central component in the development of [hessian AI](https://hessian.ai/), the [AI Innovation Lab](https://hessian.ai/infrastructure/ai-innovationlab/) ([HMWK](https://wissenschaft.hessen.de) & [HMinD](https://innen.hessen.de)) and the [AI Service Centers](https://hessian.ai/infrastructure/ai-service-centre/) ([BMBF](https://www.bmbf.de/bmbf/en/home/home_node.html)).
+The curation of the training data is partially funded by the [German Federal Ministry for Economic Affairs and Climate Action (BMWK)](https://www.bmwk.de/Navigation/EN/Home/home.html)
+through the project [OpenGPT-X](https://opengpt-x.de/en/) (project no. 68GX21007D).
+## License
+[Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0.html)

config.json ADDED Viewed

	@@ -0,0 +1,26 @@

+{
+  "_name_or_path": "mistralai/Mistral-7B-v0.1",
+  "architectures": [
+    "MistralForCausalLM"
+  ],
+  "attention_dropout": 0.0,
+  "bos_token_id": 1,
+  "eos_token_id": 2,
+  "hidden_act": "silu",
+  "hidden_size": 4096,
+  "initializer_range": 0.02,
+  "intermediate_size": 14336,
+  "max_position_embeddings": 32768,
+  "model_type": "mistral",
+  "num_attention_heads": 32,
+  "num_hidden_layers": 32,
+  "num_key_value_heads": 8,
+  "rms_norm_eps": 1e-05,
+  "rope_theta": 10000.0,
+  "sliding_window": 4096,
+  "tie_word_embeddings": false,
+  "torch_dtype": "float32",
+  "transformers_version": "4.36.2",
+  "use_cache": true,
+  "vocab_size": 32000
+}

generation_config.json ADDED Viewed

	@@ -0,0 +1,6 @@

+{
+  "_from_model_config": true,
+  "bos_token_id": 1,
+  "eos_token_id": 2,
+  "transformers_version": "4.36.2"
+}

lm-datasets-config.yml ADDED Viewed

	@@ -0,0 +1,398 @@

+# Config file for https://github.com/malteos/lm-datasets
+#
+# EU top-5 (en,fr,es,de,it) + code
+# target size: 300B tokens (train first for 200B tokens)
+# a fixed random seed for shuffling etc.
+seed: 0
+# data split settings
+validation_ratio: 0.005  # number of documents in the split: len(dataset) * ratio
+validation_min_total_docs: 1_000  # to be used as validation set, the dataset must have at least n docs
+validation_max_split_docs: 1_000  # number of documents in validation split are capped at this numbers
+validation_min_split_docs: 10  # split must have at least this number of documents, otherwise it will be discarded
+tokenizer_train_ratio: 0.1
+selected_source_ids:
+  - starcoder
+selected_dataset_ids:
+  # english
+  - pes2o
+  - math_amps
+  - eurlex_en
+  - wikipedia_20231101_en
+  - wikibooks_en
+  - wikiquote_en
+  - wikinews_en
+  - wikisource_en
+  - wikivoyage_en
+  - colossal_oscar_2015-14_en
+  - colossal_oscar_2016-40_en
+  - colossal_oscar_2017-43_en
+  - colossal_oscar_2018-47_en
+  - colossal_oscar_2019-22_en
+  - colossal_oscar_2020-24_en
+  - colossal_oscar_2020-45_en
+  - colossal_oscar_2021-49_en
+  - colossal_oscar_2022-27_en
+  - colossal_oscar_2022-49_en
+  - colossal_oscar_2023-14_en
+  - colossal_oscar_2023-23_en
+  - pile_of_law_r_legaladvice
+  - pile_of_law_atticus_contracts
+  - pile_of_law_un_debates
+  - proof_pile2_open_web_math
+  - parlamint_gb
+  - redpajama_stackexchange
+  # french
+  - cabernet
+  - eurlex_fr
+  - legal_mc4_fr
+  - wikipedia_20231101_fr
+  - wikibooks_fr
+  - wikiquote_fr
+  - wikinews_fr
+  - wikisource_fr
+  - wikivoyage_fr
+  - colossal_oscar_2015-14_fr
+  - colossal_oscar_2016-40_fr
+  - colossal_oscar_2017-43_fr
+  - colossal_oscar_2018-47_fr
+  - colossal_oscar_2019-22_fr
+  - colossal_oscar_2020-24_fr
+  - colossal_oscar_2020-45_fr
+  - colossal_oscar_2021-49_fr
+  - colossal_oscar_2022-27_fr
+  - colossal_oscar_2022-49_fr
+  - colossal_oscar_2023-14_fr
+  - colossal_oscar_2023-23_fr
+  - opensubtitles_fr
+  - parlamint_fr
+  # spanish
+  - spanish_legal
+  - eurlex_es
+  - legal_mc4_es
+  - wikipedia_20231101_es
+  - wikibooks_es
+  - wikiquote_es
+  - wikinews_es
+  - wikisource_es
+  - wikivoyage_es
+  - colossal_oscar_2015-14_es
+  - colossal_oscar_2016-40_es
+  - colossal_oscar_2017-43_es
+  - colossal_oscar_2018-47_es
+  - colossal_oscar_2019-22_es
+  - colossal_oscar_2020-24_es
+  - colossal_oscar_2020-45_es
+  - colossal_oscar_2021-49_es
+  - colossal_oscar_2022-27_es
+  - colossal_oscar_2022-49_es
+  - colossal_oscar_2023-14_es
+  - colossal_oscar_2023-23_es
+  - opensubtitles_es
+  - parlamint_es
+  # german
+  - openlegaldata
+  - dewac
+  - eurlex_de
+  - legal_mc4_de
+  - wikipedia_20231101_de
+  - wikibooks_de
+  - wikiquote_de
+  - wikinews_de
+  - wikisource_de
+  - wikivoyage_de
+  - colossal_oscar_2015-14_de
+  - colossal_oscar_2016-40_de
+  - colossal_oscar_2017-43_de
+  - colossal_oscar_2018-47_de
+  - colossal_oscar_2019-22_de
+  - colossal_oscar_2020-24_de
+  - colossal_oscar_2020-45_de
+  - colossal_oscar_2021-49_de
+  - colossal_oscar_2022-27_de
+  - colossal_oscar_2022-49_de
+  - colossal_oscar_2023-14_de
+  - colossal_oscar_2023-23_de
+  - open_discourse_bundestag
+  - tagesschau_2018_2023
+  - opensubtitles_de
+  - parlamint_at
+  # italian
+  - itwac
+  - eurlex_it
+  - legal_mc4_it
+  - wikipedia_20231101_it
+  - wikibooks_it
+  - wikiquote_it
+  - wikinews_it
+  - wikisource_it
+  - wikivoyage_it
+  - colossal_oscar_2015-14_it
+  - colossal_oscar_2016-40_it
+  - colossal_oscar_2017-43_it
+  - colossal_oscar_2018-47_it
+  - colossal_oscar_2019-22_it
+  - colossal_oscar_2020-24_it
+  - colossal_oscar_2020-45_it
+  - colossal_oscar_2021-49_it
+  - colossal_oscar_2022-27_it
+  - colossal_oscar_2022-49_it
+  - colossal_oscar_2023-14_it
+  - colossal_oscar_2023-23_it
+  - opensubtitles_it
+  - parlamint_it
+  - tatoeba_translation_en_fr
+  - tatoeba_translation_en_es
+  - tatoeba_translation_en_it
+  - tatoeba_translation_fr_it
+  - tatoeba_translation_es_fr
+  - tatoeba_translation_es_it
+  - tatoeba_translation_de_en
+  - tatoeba_translation_de_fr
+  - tatoeba_translation_de_es
+  - tatoeba_translation_de_it
+  - opus100_translation_de_en
+  - opus100_translation_en_es
+  - opus100_translation_en_fr
+  - opus100_translation_en_it
+  - wmt19_translation_de_en
+  - wmt19_translation_fr_de
+sampling_factor_by_dataset_id:
+  redpajama_stackexchange: 0.1
+  pes2o: 0.1
+  math_amps: 0.1
+  openlegaldata: 0.75
+  dewac: 0.05
+  itwac: 1
+  cabernet: 1
+  spanish_legal: 0.1
+  eurlex_de: 0.5
+  eurlex_en: 0.5
+  eurlex_es: 1
+  eurlex_fr: 1
+  eurlex_it: 1
+  legal_mc4_de: 0.1
+  legal_mc4_es: 0.25
+  legal_mc4_fr: 0.25
+  legal_mc4_it: 1
+  wikipedia_20231101_de: 2
+  wikibooks_de: 1
+  wikiquote_de: 1
+  wikinews_de: 2
+  wikisource_de: 1
+  wikivoyage_de: 1
+  wikipedia_20231101_en: 1
+  wikibooks_en: 1
+  wikiquote_en: 0.25
+  wikinews_en: 1
+  wikisource_en: 1
+  wikivoyage_en: 1
+  wikipedia_20231101_es: 2
+  wikibooks_es: 1
+  wikiquote_es: 1
+  wikinews_es: 2
+  wikisource_es: 1
+  wikivoyage_es: 1
+  wikipedia_20231101_fr: 2
+  wikibooks_fr: 1
+  wikiquote_fr: 1
+  wikinews_fr: 2
+  wikisource_fr: 1
+  wikivoyage_fr: 1
+  wikipedia_20231101_it: 2
+  wikibooks_it: 1
+  wikiquote_it: 1
+  wikinews_it: 2
+  wikisource_it: 1
+  wikivoyage_it: 1
+  colossal_oscar_2015-14_de: 1
+  colossal_oscar_2016-40_de: 0.95
+  colossal_oscar_2017-43_de: 0.1
+  colossal_oscar_2018-47_de: 0.1
+  colossal_oscar_2019-22_de: 0.1
+  colossal_oscar_2020-24_de: 0.1
+  colossal_oscar_2020-45_de: 0.1
+  colossal_oscar_2021-49_de: 0.1
+  colossal_oscar_2022-27_de: 0.1
+  colossal_oscar_2022-49_de: 0.1
+  colossal_oscar_2023-14_de: 0.95
+  colossal_oscar_2023-23_de: 1
+  colossal_oscar_2015-14_en: 0.05
+  colossal_oscar_2016-40_en: 0.05
+  colossal_oscar_2017-43_en: 0.001
+  colossal_oscar_2018-47_en: 0.001
+  colossal_oscar_2019-22_en: 0.001
+  colossal_oscar_2020-24_en: 0.001
+  colossal_oscar_2020-45_en: 0.001
+  colossal_oscar_2021-49_en: 0.001
+  colossal_oscar_2022-27_en: 0.001
+  colossal_oscar_2022-49_en: 0.001
+  colossal_oscar_2023-14_en: 0.05
+  colossal_oscar_2023-23_en: 0.05
+  colossal_oscar_2015-14_es: 1
+  colossal_oscar_2016-40_es: 1
+  colossal_oscar_2017-43_es: 0.25
+  colossal_oscar_2018-47_es: 0.1
+  colossal_oscar_2019-22_es: 0.1
+  colossal_oscar_2020-24_es: 0.1
+  colossal_oscar_2020-45_es: 0.1
+  colossal_oscar_2021-49_es: 0.1
+  colossal_oscar_2022-27_es: 0.1
+  colossal_oscar_2022-49_es: 0.3
+  colossal_oscar_2023-14_es: 1
+  colossal_oscar_2023-23_es: 1
+  colossal_oscar_2015-14_fr: 1
+  colossal_oscar_2016-40_fr: 1
+  colossal_oscar_2017-43_fr: 0.25
+  colossal_oscar_2018-47_fr: 0.25
+  colossal_oscar_2019-22_fr: 0.1
+  colossal_oscar_2020-24_fr: 0.1
+  colossal_oscar_2020-45_fr: 0.1
+  colossal_oscar_2021-49_fr: 0.1
+  colossal_oscar_2022-27_fr: 0.1
+  colossal_oscar_2022-49_fr: 0.75
+  colossal_oscar_2023-14_fr: 1
+  colossal_oscar_2023-23_fr: 1
+  starcoder_emacs-lisp: 0.1
+  starcoder_literate-haskell: 0.1
+  starcoder_shell: 0.1
+  starcoder_ada: 0.1
+  starcoder_erlang: 0.1
+  starcoder_lua: 0.1
+  starcoder_smalltalk: 0.1
+  starcoder_agda: 0.1
+  starcoder_f-sharp: 0.1
+  starcoder_makefile: 0.1
+  starcoder_solidity: 0.1
+  starcoder_alloy: 0.1
+  starcoder_fortran: 0.1
+  starcoder_maple: 0.1
+  starcoder_sparql: 0.1
+  starcoder_antlr: 0.1
+  starcoder_git-commits-cleaned: 0.05
+  starcoder_markdown: 0.05
+  starcoder_sql: 0.1
+  starcoder_applescript: 0.1
+  starcoder_github-issues-filtered-structured: 0.075
+  starcoder_mathematica: 0.1
+  starcoder_stan: 0.1
+  starcoder_assembly: 0.1
+  starcoder_glsl: 0.1
+  starcoder_matlab: 0.1
+  starcoder_standard-ml: 0.1
+  starcoder_augeas: 0.1
+  starcoder_go: 0.05
+  starcoder_ocaml: 0.1
+  starcoder_stata: 0.1
+  starcoder_awk: 0.1
+  starcoder_groovy: 0.1
+  starcoder_pascal: 0.1
+  starcoder_systemverilog: 0.1
+  starcoder_batchfile: 0.1
+  starcoder_haskell: 0.1
+  starcoder_perl: 0.1
+  starcoder_tcl: 0.1
+  starcoder_bluespec: 0.1
+  starcoder_html: 0.05
+  starcoder_php: 0.05
+  starcoder_tcsh: 0.1
+  starcoder_c: 0.05
+  starcoder_idris: 0.1
+  starcoder_powershell: 0.1
+  starcoder_tex: 0.1
+  starcoder_c-sharp: 0.05
+  starcoder_isabelle: 0.1
+  starcoder_prolog: 0.1
+  starcoder_thrift: 0.1
+  starcoder_clojure: 0.1
+  starcoder_java: 0.05
+  starcoder_protocol-buffer: 0.1
+  starcoder_typescript: 0.05
+  starcoder_cmake: 0.1
+  starcoder_java-server-pages: 0.1
+  starcoder_python: 0.05
+  starcoder_verilog: 0.1
+  starcoder_coffeescript: 0.1
+  starcoder_javascript: 0.05
+  starcoder_r: 0.1
+  starcoder_vhdl: 0.1
+  starcoder_common-lisp: 0.1
+  starcoder_json: 0.1
+  starcoder_racket: 0.1
+  starcoder_visual-basic: 0.1
+  starcoder_cpp: 0.05
+  starcoder_julia: 0.1
+  starcoder_restructuredtext: 0.1
+  starcoder_xslt: 0.1
+  starcoder_css: 0.1
+  starcoder_jupyter-scripts-dedup-filtered: 0.1
+  starcoder_rmarkdown: 0.1
+  starcoder_yacc: 0.1
+  starcoder_cuda: 0.1
+  starcoder_jupyter-structured-clean-dedup: 0.1
+  starcoder_ruby: 0.1
+  starcoder_yaml: 0.1
+  starcoder_dart: 0.1
+  starcoder_kotlin: 0.1
+  starcoder_rust: 0.1
+  starcoder_zig: 0.1
+  starcoder_dockerfile: 0.1
+  starcoder_lean: 0.1
+  starcoder_sas: 0.1
+  starcoder_elixir: 0.1
+  starcoder_literate-agda: 0.1
+  starcoder_scala: 0.1
+  starcoder_elm: 0.1
+  starcoder_literate-coffeescript: 0.1
+  starcoder_scheme: 0.1
+  pile_of_law_r_legaladvice: 1
+  pile_of_law_atticus_contracts: 0.25
+  pile_of_law_un_debates: 1
+  open_discourse_bundestag: 0.5
+  tagesschau_2018_2023: 1
+  proof_pile2_open_web_math: 0.25
+  tatoeba_translation_en_fr: 1
+  tatoeba_translation_en_es: 1
+  tatoeba_translation_en_it: 1
+  tatoeba_translation_fr_it: 1
+  tatoeba_translation_es_fr: 1
+  tatoeba_translation_es_it: 1
+  tatoeba_translation_de_en: 1
+  tatoeba_translation_de_fr: 1
+  tatoeba_translation_de_es: 1
+  tatoeba_translation_de_it: 1
+  opus100_translation_de_en: 1
+  opus100_translation_en_es: 1
+  opus100_translation_en_fr: 1
+  opus100_translation_en_it: 1
+  wmt19_translation_de_en: 1
+  wmt19_translation_fr_de: 1
+  opensubtitles_es: 1
+  opensubtitles_fr: 1
+  opensubtitles_de: 1
+  opensubtitles_it: 1
+  parlamint_es: 1
+  parlamint_it: 1
+  parlamint_at: 1
+  parlamint_fr: 1
+  parlamint_gb: 1
+  colossal_oscar_2015-14_it: 1
+  colossal_oscar_2016-40_it: 1
+  colossal_oscar_2017-43_it: 0.75
+  colossal_oscar_2018-47_it: 0.75
+  colossal_oscar_2019-22_it: 0.75
+  colossal_oscar_2020-24_it: 0.75
+  colossal_oscar_2020-45_it: 0.75
+  colossal_oscar_2021-49_it: 0.75
+  colossal_oscar_2022-27_it: 0.75
+  colossal_oscar_2022-49_it: 0.75
+  colossal_oscar_2023-14_it: 0.9
+  colossal_oscar_2023-23_it: 1

model-00001-of-00006.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:3572c45423614e2b874ae7d421fd4f850e7cabc7e966980547b2cdce71e515f0
+size 4987196936

model-00002-of-00006.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:fdb2cd75b9f7dbb8ba0256ba6771bc7fbc54310ea9bb3f57d57eed297ea2d3f4
+size 4899116440

model-00003-of-00006.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:1c560c912b037e5af4c78bab14eacc2c3fc5f7724371fe68e6f0b0fe4ffafc19
+size 4999813120

model-00004-of-00006.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:729463205aaed76447ca86177110de7a825fe8e6d8b7db5dfd35caef42cd2531
+size 4999813128

model-00005-of-00006.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:897a386e7f7a04a78c0c130538f715c7d4c699ba471919b2dc3d381a67f358bb
+size 4832007496

model-00006-of-00006.safetensors ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:e8f37c13b68f2eda900c982f0dd6cae586d3df65dfa7bcf87dbb6717a8e434f6
+size 4249014896

model.safetensors.index.json ADDED Viewed

	@@ -0,0 +1,298 @@

+{
+  "metadata": {
+    "total_size": 28966928384
+  },
+  "weight_map": {
+    "lm_head.weight": "model-00006-of-00006.safetensors",
+    "model.embed_tokens.weight": "model-00001-of-00006.safetensors",
+    "model.layers.0.input_layernorm.weight": "model-00001-of-00006.safetensors",
+    "model.layers.0.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.0.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.0.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.0.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
+    "model.layers.0.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.0.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.0.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.0.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.1.input_layernorm.weight": "model-00001-of-00006.safetensors",
+    "model.layers.1.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.1.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.1.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.1.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
+    "model.layers.1.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.1.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.1.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.1.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.10.input_layernorm.weight": "model-00003-of-00006.safetensors",
+    "model.layers.10.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.10.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.10.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.10.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
+    "model.layers.10.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.10.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.10.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.10.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.11.input_layernorm.weight": "model-00003-of-00006.safetensors",
+    "model.layers.11.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.11.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.11.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.11.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
+    "model.layers.11.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.11.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.11.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.11.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.12.input_layernorm.weight": "model-00003-of-00006.safetensors",
+    "model.layers.12.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.12.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.12.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.12.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
+    "model.layers.12.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.12.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.12.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.12.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.13.input_layernorm.weight": "model-00003-of-00006.safetensors",
+    "model.layers.13.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.13.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.13.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.13.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
+    "model.layers.13.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.13.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.13.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.13.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.14.input_layernorm.weight": "model-00003-of-00006.safetensors",
+    "model.layers.14.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.14.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.14.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.14.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
+    "model.layers.14.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.14.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.14.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.14.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.15.input_layernorm.weight": "model-00003-of-00006.safetensors",
+    "model.layers.15.mlp.down_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.15.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.15.mlp.up_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.15.post_attention_layernorm.weight": "model-00003-of-00006.safetensors",
+    "model.layers.15.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.15.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.15.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.15.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.16.input_layernorm.weight": "model-00004-of-00006.safetensors",
+    "model.layers.16.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.16.mlp.gate_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.16.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.16.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
+    "model.layers.16.self_attn.k_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.16.self_attn.o_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.16.self_attn.q_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.16.self_attn.v_proj.weight": "model-00003-of-00006.safetensors",
+    "model.layers.17.input_layernorm.weight": "model-00004-of-00006.safetensors",
+    "model.layers.17.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.17.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.17.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.17.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
+    "model.layers.17.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.17.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.17.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.17.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.18.input_layernorm.weight": "model-00004-of-00006.safetensors",
+    "model.layers.18.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.18.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.18.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.18.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
+    "model.layers.18.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.18.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.18.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.18.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.19.input_layernorm.weight": "model-00004-of-00006.safetensors",
+    "model.layers.19.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.19.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.19.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.19.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
+    "model.layers.19.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.19.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.19.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.19.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.2.input_layernorm.weight": "model-00001-of-00006.safetensors",
+    "model.layers.2.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.2.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.2.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.2.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
+    "model.layers.2.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.2.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.2.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.2.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.20.input_layernorm.weight": "model-00004-of-00006.safetensors",
+    "model.layers.20.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.20.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.20.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.20.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
+    "model.layers.20.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.20.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.20.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.20.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.21.input_layernorm.weight": "model-00004-of-00006.safetensors",
+    "model.layers.21.mlp.down_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.21.mlp.gate_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.21.mlp.up_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.21.post_attention_layernorm.weight": "model-00004-of-00006.safetensors",
+    "model.layers.21.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.21.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.21.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.21.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.22.input_layernorm.weight": "model-00005-of-00006.safetensors",
+    "model.layers.22.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.22.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.22.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.22.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
+    "model.layers.22.self_attn.k_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.22.self_attn.o_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.22.self_attn.q_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.22.self_attn.v_proj.weight": "model-00004-of-00006.safetensors",
+    "model.layers.23.input_layernorm.weight": "model-00005-of-00006.safetensors",
+    "model.layers.23.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.23.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.23.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.23.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
+    "model.layers.23.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.23.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.23.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.23.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.24.input_layernorm.weight": "model-00005-of-00006.safetensors",
+    "model.layers.24.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.24.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.24.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.24.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
+    "model.layers.24.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.24.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.24.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.24.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.25.input_layernorm.weight": "model-00005-of-00006.safetensors",
+    "model.layers.25.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.25.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.25.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.25.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
+    "model.layers.25.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.25.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.25.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.25.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.26.input_layernorm.weight": "model-00005-of-00006.safetensors",
+    "model.layers.26.mlp.down_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.26.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.26.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.26.post_attention_layernorm.weight": "model-00005-of-00006.safetensors",
+    "model.layers.26.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.26.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.26.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.26.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.27.input_layernorm.weight": "model-00006-of-00006.safetensors",
+    "model.layers.27.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.27.mlp.gate_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.27.mlp.up_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.27.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
+    "model.layers.27.self_attn.k_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.27.self_attn.o_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.27.self_attn.q_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.27.self_attn.v_proj.weight": "model-00005-of-00006.safetensors",
+    "model.layers.28.input_layernorm.weight": "model-00006-of-00006.safetensors",
+    "model.layers.28.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.28.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.28.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.28.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
+    "model.layers.28.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.28.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.28.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.28.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.29.input_layernorm.weight": "model-00006-of-00006.safetensors",
+    "model.layers.29.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.29.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.29.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.29.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
+    "model.layers.29.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.29.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.29.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.29.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.3.input_layernorm.weight": "model-00001-of-00006.safetensors",
+    "model.layers.3.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.3.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.3.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.3.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
+    "model.layers.3.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.3.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.3.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.3.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.30.input_layernorm.weight": "model-00006-of-00006.safetensors",
+    "model.layers.30.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.30.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.30.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.30.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
+    "model.layers.30.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.30.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.30.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.30.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.31.input_layernorm.weight": "model-00006-of-00006.safetensors",
+    "model.layers.31.mlp.down_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.31.mlp.gate_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.31.mlp.up_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.31.post_attention_layernorm.weight": "model-00006-of-00006.safetensors",
+    "model.layers.31.self_attn.k_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.31.self_attn.o_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.31.self_attn.q_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.31.self_attn.v_proj.weight": "model-00006-of-00006.safetensors",
+    "model.layers.4.input_layernorm.weight": "model-00001-of-00006.safetensors",
+    "model.layers.4.mlp.down_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.4.mlp.gate_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.4.mlp.up_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.4.post_attention_layernorm.weight": "model-00001-of-00006.safetensors",
+    "model.layers.4.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.4.self_attn.o_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.4.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.4.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.5.input_layernorm.weight": "model-00002-of-00006.safetensors",
+    "model.layers.5.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.5.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.5.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.5.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
+    "model.layers.5.self_attn.k_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.5.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.5.self_attn.q_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.5.self_attn.v_proj.weight": "model-00001-of-00006.safetensors",
+    "model.layers.6.input_layernorm.weight": "model-00002-of-00006.safetensors",
+    "model.layers.6.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.6.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.6.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.6.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
+    "model.layers.6.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.6.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.6.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.6.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.7.input_layernorm.weight": "model-00002-of-00006.safetensors",
+    "model.layers.7.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.7.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.7.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.7.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
+    "model.layers.7.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.7.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.7.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.7.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.8.input_layernorm.weight": "model-00002-of-00006.safetensors",
+    "model.layers.8.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.8.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.8.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.8.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
+    "model.layers.8.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.8.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.8.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.8.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.9.input_layernorm.weight": "model-00002-of-00006.safetensors",
+    "model.layers.9.mlp.down_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.9.mlp.gate_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.9.mlp.up_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.9.post_attention_layernorm.weight": "model-00002-of-00006.safetensors",
+    "model.layers.9.self_attn.k_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.9.self_attn.o_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.9.self_attn.q_proj.weight": "model-00002-of-00006.safetensors",
+    "model.layers.9.self_attn.v_proj.weight": "model-00002-of-00006.safetensors",
+    "model.norm.weight": "model-00006-of-00006.safetensors"
+  }
+}

special_tokens_map.json ADDED Viewed

	@@ -0,0 +1,5 @@

+{
+  "bos_token": "<s>",
+  "eos_token": "</s>",
+  "unk_token": "<unk>"
+}

tokenizer.json ADDED Viewed

The diff for this file is too large to render. See raw diff

tokenizer.model ADDED Viewed

	@@ -0,0 +1,3 @@

+version https://git-lfs.github.com/spec/v1
+oid sha256:dadfd56d766715c61d2ef780a525ab43b8e6da4de6865bda3d95fdef5e134055
+size 493443

tokenizer_config.json ADDED Viewed

	@@ -0,0 +1,42 @@

+{
+  "add_bos_token": true,
+  "add_eos_token": false,
+  "added_tokens_decoder": {
+    "0": {
+      "content": "<unk>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "1": {
+      "content": "<s>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    },
+    "2": {
+      "content": "</s>",
+      "lstrip": false,
+      "normalized": false,
+      "rstrip": false,
+      "single_word": false,
+      "special": true
+    }
+  },
+  "additional_special_tokens": [],
+  "bos_token": "<s>",
+  "clean_up_tokenization_spaces": false,
+  "eos_token": "</s>",
+  "legacy": true,
+  "model_max_length": 1000000000000000019884624838656,
+  "pad_token": null,
+  "sp_model_kwargs": {},
+  "spaces_between_special_tokens": false,
+  "tokenizer_class": "LlamaTokenizer",
+  "unk_token": "<unk>",
+  "use_default_system_prompt": false
+}