3outeille's picture
3outeille HF staff
Upload llama-1B/16_GPUS/dp-1_tp-8_pp-2_mbz-256
0a85857 verified
raw
history blame
219 kB
========================
START TIME: Tue Jul 2 18:22:22 UTC 2024
python3 version = Python 3.10.14
========================
The token has not been saved to the git credentials helper. Pass `add_to_git_credential=True` in this function directly or `--add-to-git-credential` if using via `huggingface-cli` if you want to set the git credential as well.
Token is valid (permission: write).
Your token has been saved to /admin/home/ferdinand_mom/.cache/huggingface/token
Login successful
Already on 'bench_cluster'
M examples/config_tiny_llama.py
M examples/config_tiny_llama.yaml
M examples/train_tiny_llama.sh
M src/nanotron/models/llama.py
M src/nanotron/trainer.py
Your branch is up to date with 'origin/bench_cluster'.
Job status: RUNNING
W0702 18:22:24.801000 140244857890624 torch/distributed/run.py:757]
W0702 18:22:24.801000 140244857890624 torch/distributed/run.py:757] *****************************************
W0702 18:22:24.801000 140244857890624 torch/distributed/run.py:757] Setting OMP_NUM_THREADS environment variable for each process to be 1 in default, to avoid your system being overloaded, please further tune the variable for optimal performance in your application as needed.
W0702 18:22:24.801000 140244857890624 torch/distributed/run.py:757] *****************************************
W0702 18:22:24.798000 139628476266304 torch/distributed/run.py:757]
W0702 18:22:24.798000 139628476266304 torch/distributed/run.py:757] *****************************************
W0702 18:22:24.798000 139628476266304 torch/distributed/run.py:757] Setting OMP_NUM_THREADS environment variable for each process to be 1 in default, to avoid your system being overloaded, please further tune the variable for optimal performance in your application as needed.
W0702 18:22:24.798000 139628476266304 torch/distributed/run.py:757] *****************************************
[default0]:07/02/2024 18:22:43 [WARNING|DP=0|PP=0|TP=0|ip-26-0-163-147]: [Vocab Size Padding] Padded vocab (size: 50257) with 7 dummy tokens (new size: 50264)
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Config:
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Config(general=GeneralArgs(project='bench_cluster',
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: run='%date_%jobid',
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: seed=42,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: step=None,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: consumed_train_samples=None,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: benchmark_csv_path=None,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: ignore_sanity_checks=True),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: parallelism=ParallelismArgs(dp=1,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: pp=2,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tp=8,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: pp_engine=<nanotron.parallel.pipeline_parallel.engine.OneForwardOneBackwardPipelineEngine object at 0x7f822efb4910>,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tp_mode=<TensorParallelLinearMode.REDUCE_SCATTER: 2>,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tp_linear_async_communication=False,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: expert_parallel_size=1),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: model=ModelArgs(model_config=LlamaConfig(bos_token_id=1,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: eos_token_id=2,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: hidden_act='silu',
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: hidden_size=2048,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: initializer_range=0.02,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: intermediate_size=4096,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: is_llama_config=True,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: max_position_embeddings=4096,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_attention_heads=32,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_hidden_layers=24,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_key_value_heads=32,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: pad_token_id=None,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: pretraining_tp=1,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: rms_norm_eps=1e-05,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: rope_scaling=None,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: rope_theta=10000.0,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tie_word_embeddings=True,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: use_cache=True,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: vocab_size=50264),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: init_method=RandomInit(std=0.025),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: dtype=torch.bfloat16,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: make_vocab_size_divisible_by=1,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: ddp_bucket_cap_mb=25),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tokenizer=TokenizerArgs(tokenizer_name_or_path='openai-community/gpt2',
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tokenizer_revision=None,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tokenizer_max_length=None),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: checkpoints=CheckpointsArgs(checkpoints_path=Path('/dev/null'),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: checkpoint_interval=100000,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: save_initial_state=False,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: resume_checkpoint_path=None,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: checkpoints_path_is_shared_file_system=False),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: logging=LoggingArgs(log_level='info',
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: log_level_replica='info',
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: iteration_step_info_interval=1),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tokens=TokensArgs(sequence_length=4096,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: train_steps=20,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: micro_batch_size=256,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: batch_accumulation_per_replica=4,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: val_check_interval=-1,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: limit_val_batches=0,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: limit_test_batches=0),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: optimizer=OptimizerArgs(optimizer_factory=AdamWOptimizerArgs(adam_eps=1e-08,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: adam_beta1=0.9,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: adam_beta2=0.95,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: torch_adam_is_fused=True,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: name='adamW'),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: zero_stage=1,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: weight_decay=0.01,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: clip_grad=1.0,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: accumulate_grad_in_fp32=True,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: learning_rate_scheduler=LRSchedulerArgs(learning_rate=0.0001,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: lr_warmup_steps=1,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: lr_warmup_style='linear',
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: lr_decay_style='linear',
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: lr_decay_steps=19,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: lr_decay_starting_step=None,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: min_decay_lr=1e-05)),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: data_stages=[DatasetStageArgs(name='Training Stage',
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: start_training_step=1,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: data=DataArgs(dataset=PretrainDatasetsArgs(hf_dataset_or_datasets='roneneldan/TinyStories',
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: hf_dataset_splits='train',
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: hf_dataset_config_name=None,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: dataset_processing_num_proc_per_process=64,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: dataset_overwrite_cache=False,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: text_column_name='text'),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: seed=42,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_loading_workers=32))],
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: profiler=ProfilerArgs(profiler_export_path=Path('/fsx/ferdinandmom/ferdinand-hf/bench_cluster/results/llama-1B/16_GPUS/dp-1_tp-8_pp-2_mbz-256')),
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: lighteval=None)
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Model Config:
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: LlamaConfig(bos_token_id=1,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: eos_token_id=2,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: hidden_act='silu',
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: hidden_size=2048,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: initializer_range=0.02,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: intermediate_size=4096,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: is_llama_config=True,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: max_position_embeddings=4096,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_attention_heads=32,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_hidden_layers=24,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_key_value_heads=32,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: pad_token_id=None,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: pretraining_tp=1,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: rms_norm_eps=1e-05,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: rope_scaling=None,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: rope_theta=10000.0,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tie_word_embeddings=True,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: use_cache=True,
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: vocab_size=50264)
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Building model..
[default0]:07/02/2024 18:22:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Setting PP block ranks...
[default5]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=5|ip-26-0-163-226]: Local number of parameters: 65.3M (124.62MiB)
[default5]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=5|ip-26-0-163-226]: [After model building] Memory usage: 135.64MiB. Peak allocated: 137.67MiB Peak reserved: 150.00MiB
[default5]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=5|ip-26-0-163-226]: No checkpoint path provided.
[default2]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=2|ip-26-0-163-226]: Local number of parameters: 65.3M (124.62MiB)
[default2]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=2|ip-26-0-163-226]: [After model building] Memory usage: 135.64MiB. Peak allocated: 137.67MiB Peak reserved: 150.00MiB
[default2]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=2|ip-26-0-163-226]: No checkpoint path provided.
[default1]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=1|ip-26-0-163-226]: Local number of parameters: 65.3M (124.62MiB)
[default1]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=1|ip-26-0-163-226]: [After model building] Memory usage: 135.64MiB. Peak allocated: 137.67MiB Peak reserved: 150.00MiB
[default1]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=1|ip-26-0-163-226]: No checkpoint path provided.
[default0]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=0|ip-26-0-163-226]: Local number of parameters: 65.3M (124.62MiB)
[default0]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=0|ip-26-0-163-226]: [After model building] Memory usage: 135.64MiB. Peak allocated: 137.67MiB Peak reserved: 150.00MiB
[default0]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=0|ip-26-0-163-226]: No checkpoint path provided.
[default4]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=4|ip-26-0-163-226]: Local number of parameters: 65.3M (124.62MiB)
[default4]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=4|ip-26-0-163-226]: [After model building] Memory usage: 135.64MiB. Peak allocated: 137.67MiB Peak reserved: 150.00MiB
[default4]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=4|ip-26-0-163-226]: No checkpoint path provided.
[default3]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=3|ip-26-0-163-226]: Local number of parameters: 65.3M (124.62MiB)
[default3]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=3|ip-26-0-163-226]: [After model building] Memory usage: 135.64MiB. Peak allocated: 137.67MiB Peak reserved: 150.00MiB
[default3]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=3|ip-26-0-163-226]: No checkpoint path provided.
[default7]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=7|ip-26-0-163-226]: Local number of parameters: 65.3M (124.62MiB)
[default7]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=7|ip-26-0-163-226]: [After model building] Memory usage: 135.64MiB. Peak allocated: 137.67MiB Peak reserved: 150.00MiB
[default7]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=7|ip-26-0-163-226]: No checkpoint path provided.
[default6]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=6|ip-26-0-163-226]: Local number of parameters: 65.3M (124.62MiB)
[default6]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=6|ip-26-0-163-226]: [After model building] Memory usage: 135.64MiB. Peak allocated: 137.67MiB Peak reserved: 150.00MiB
[default6]:07/02/2024 18:22:58 [INFO|DP=0|PP=1|TP=6|ip-26-0-163-226]: No checkpoint path provided.
[default5]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=5|ip-26-0-163-147]: Local number of parameters: 86.3M (164.65MiB)
[default5]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=5|ip-26-0-163-147]: [After model building] Memory usage: 179.67MiB. Peak allocated: 181.70MiB Peak reserved: 198.00MiB
[default5]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=5|ip-26-0-163-147]: No checkpoint path provided.
[default3]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=3|ip-26-0-163-147]: Local number of parameters: 86.3M (164.65MiB)
[default3]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=3|ip-26-0-163-147]: [After model building] Memory usage: 179.67MiB. Peak allocated: 181.70MiB Peak reserved: 198.00MiB
[default3]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=3|ip-26-0-163-147]: No checkpoint path provided.
[default1]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=1|ip-26-0-163-147]: Local number of parameters: 86.3M (164.65MiB)
[default1]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=1|ip-26-0-163-147]: [After model building] Memory usage: 179.67MiB. Peak allocated: 181.70MiB Peak reserved: 198.00MiB
[default1]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=1|ip-26-0-163-147]: No checkpoint path provided.
[default4]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=4|ip-26-0-163-147]: Local number of parameters: 86.3M (164.65MiB)
[default4]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=4|ip-26-0-163-147]: [After model building] Memory usage: 179.67MiB. Peak allocated: 181.70MiB Peak reserved: 198.00MiB
[default4]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=4|ip-26-0-163-147]: No checkpoint path provided.
[default0]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Total number of parameters: 1.21G (2314.22MiB)
[default0]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Local number of parameters: 86.3M (164.65MiB)
[default0]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [After model building] Memory usage: 179.67MiB. Peak allocated: 181.70MiB Peak reserved: 198.00MiB
[default0]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: No checkpoint path provided.
[default0]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Parametrizing model parameters using StandardParametrizator
[default6]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=6|ip-26-0-163-147]: Local number of parameters: 86.3M (164.65MiB)
[default6]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=6|ip-26-0-163-147]: [After model building] Memory usage: 179.67MiB. Peak allocated: 181.70MiB Peak reserved: 198.00MiB
[default6]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=6|ip-26-0-163-147]: No checkpoint path provided.
[default2]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=2|ip-26-0-163-147]: Local number of parameters: 86.3M (164.65MiB)
[default2]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=2|ip-26-0-163-147]: [After model building] Memory usage: 179.67MiB. Peak allocated: 181.70MiB Peak reserved: 198.00MiB
[default2]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=2|ip-26-0-163-147]: No checkpoint path provided.
[default7]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=7|ip-26-0-163-147]: Local number of parameters: 86.3M (164.65MiB)
[default7]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=7|ip-26-0-163-147]: [After model building] Memory usage: 179.67MiB. Peak allocated: 181.70MiB Peak reserved: 198.00MiB
[default7]:07/02/2024 18:22:58 [INFO|DP=0|PP=0|TP=7|ip-26-0-163-147]: No checkpoint path provided.
[default0]:07/02/2024 18:22:59 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [Optimizer Building] Using LearningRateForSP as learning rate
[default0]:07/02/2024 18:22:59 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [ZeRO sharding] Size of optimizer params per rank:
[default0]:07/02/2024 18:22:59 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [ZeRO sharding] DP Rank 0 has 86.3M out of 86.3M (100.00%) params' optimizer states
[default0]:07/02/2024 18:23:01 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [Training Plan] Stage Training Stage has 19 remaining training steps and has consumed 0 samples
[default0]:07/02/2024 18:23:01 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Using `datasets` library
[default0]:07/02/2024 18:23:01 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Loading tokenizer from openai-community/gpt2 and transformers/hf_hub versions ('4.41.2', '0.23.4')
[default0]:07/02/2024 18:23:01 [WARNING|DP=0|PP=0|TP=0|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default0]:Repo card metadata block was not found. Setting CardData to empty.
[default0]:07/02/2024 18:23:02 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [Training Plan] There are 1 training stages
[default0]:07/02/2024 18:23:02 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [Stage Training Stage] start from step 1
[default0]:07/02/2024 18:23:02 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]:
[default0]:07/02/2024 18:23:02 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [Start training] datetime: 2024-07-02 18:23:02.061141 | mbs: 256 | grad_accum: 4 | global_batch_size: 1024 | sequence_length: 4096 | train_steps: 20 | start_iteration_step: 0 | consumed_train_samples: 0
[default0]:07/02/2024 18:23:02 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Resuming training from stage Training Stage, it has trained for 0 samples and has 19 remaining train steps
[default0]:07/02/2024 18:23:02 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Memory usage: 839.67MiB. Peak allocated 839.67MiB. Peak reserved: 858.00MiB
[default1]:07/02/2024 18:23:02 [WARNING|DP=0|PP=1|TP=1|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default1]:Repo card metadata block was not found. Setting CardData to empty.
[default0]:07/02/2024 18:23:02 [WARNING|DP=0|PP=1|TP=0|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default0]:Repo card metadata block was not found. Setting CardData to empty.
[default6]:Repo card metadata block was not found. Setting CardData to empty.
[default6]:07/02/2024 18:23:02 [WARNING|DP=0|PP=1|TP=6|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default3]:07/02/2024 18:23:02 [WARNING|DP=0|PP=0|TP=3|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default3]:Repo card metadata block was not found. Setting CardData to empty.
[default2]:07/02/2024 18:23:02 [WARNING|DP=0|PP=1|TP=2|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default3]:07/02/2024 18:23:02 [WARNING|DP=0|PP=1|TP=3|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default7]:07/02/2024 18:23:02 [WARNING|DP=0|PP=1|TP=7|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default2]:Repo card metadata block was not found. Setting CardData to empty.
[default3]:Repo card metadata block was not found. Setting CardData to empty.
[default7]:Repo card metadata block was not found. Setting CardData to empty.
[default5]:07/02/2024 18:23:02 [WARNING|DP=0|PP=0|TP=5|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default1]:07/02/2024 18:23:02 [WARNING|DP=0|PP=0|TP=1|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default1]:Repo card metadata block was not found. Setting CardData to empty.
[default6]:Repo card metadata block was not found. Setting CardData to empty.
[default5]:Repo card metadata block was not found. Setting CardData to empty.
[default2]:07/02/2024 18:23:02 [WARNING|DP=0|PP=0|TP=2|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default6]:07/02/2024 18:23:02 [WARNING|DP=0|PP=0|TP=6|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default2]:Repo card metadata block was not found. Setting CardData to empty.
[default4]:07/02/2024 18:23:02 [WARNING|DP=0|PP=1|TP=4|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default4]:Repo card metadata block was not found. Setting CardData to empty.
[default7]:Repo card metadata block was not found. Setting CardData to empty.
[default4]:Repo card metadata block was not found. Setting CardData to empty.
[default4]:07/02/2024 18:23:02 [WARNING|DP=0|PP=0|TP=4|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default7]:07/02/2024 18:23:02 [WARNING|DP=0|PP=0|TP=7|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default5]:07/02/2024 18:23:02 [WARNING|DP=0|PP=1|TP=5|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default5]:Repo card metadata block was not found. Setting CardData to empty.
[default7]:[rank7]: Traceback (most recent call last):
[default7]:[rank7]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default7]:[rank7]: trainer.train(dataloader)
[default7]:[rank7]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default7]:[rank7]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default7]:[rank7]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default7]:[rank7]: outputs = self.pipeline_engine.train_batch_iter(
[default7]:[rank7]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default7]:[rank7]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default7]:[rank7]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default7]:[rank7]: output = model(**micro_batch)
[default7]:[rank7]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default7]:[rank7]: return self._call_impl(*args, **kwargs)
[default7]:[rank7]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default7]:[rank7]: return forward_call(*args, **kwargs)
[default7]:[rank7]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default7]:[rank7]: sharded_logits = self.model(
[default7]:[rank7]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default7]:[rank7]: return self._call_impl(*args, **kwargs)
[default7]:[rank7]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default7]:[rank7]: return forward_call(*args, **kwargs)
[default7]:[rank7]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default7]:[rank7]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default7]:[rank7]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default7]:[rank7]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default7]:[rank7]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default7]:[rank7]: return self._call_impl(*args, **kwargs)
[default7]:[rank7]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default7]:[rank7]: return forward_call(*args, **kwargs)
[default7]:[rank7]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 151, in forward
[default7]:[rank7]: output = self.pp_block(**new_kwargs)
[default7]:[rank7]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default7]:[rank7]: return self._call_impl(*args, **kwargs)
[default4]:[rank4]: Traceback (most recent call last):
[default4]:[rank4]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default4]:[rank4]: trainer.train(dataloader)
[default4]:[rank4]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default4]:[rank4]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default4]:[rank4]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default4]:[rank4]: outputs = self.pipeline_engine.train_batch_iter(
[default4]:[rank4]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default4]:[rank4]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default4]:[rank4]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default4]:[rank4]: output = model(**micro_batch)
[default4]:[rank4]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default4]:[rank4]: return self._call_impl(*args, **kwargs)
[default4]:[rank4]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank4]: return forward_call(*args, **kwargs)
[default4]:[rank4]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default4]:[rank4]: sharded_logits = self.model(
[default4]:[rank4]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default4]:[rank4]: return self._call_impl(*args, **kwargs)
[default4]:[rank4]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank4]: return forward_call(*args, **kwargs)
[default4]:[rank4]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default4]:[rank4]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default4]:[rank4]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default4]:[rank4]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default4]:[rank4]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default4]:[rank4]: return self._call_impl(*args, **kwargs)
[default4]:[rank4]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank4]: return forward_call(*args, **kwargs)
[default4]:[rank4]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 151, in forward
[default4]:[rank4]: output = self.pp_block(**new_kwargs)
[default4]:[rank4]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default4]:[rank4]: return self._call_impl(*args, **kwargs)
[default4]:[rank4]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank4]: return forward_call(*args, **kwargs)
[default4]:[rank4]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 637, in forward
[default4]:[rank4]: hidden_states = self.mlp(hidden_states=hidden_states)["hidden_states"]
[default4]:[rank4]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default4]:[rank4]: return self._call_impl(*args, **kwargs)
[default4]:[rank4]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank4]: return forward_call(*args, **kwargs)
[default4]:[rank4]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 172, in forward
[default4]:[rank4]: hidden_states = self.down_proj(self.split_silu_mul(merged_states))
[default4]:[rank4]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default4]:[rank4]: return self._call_impl(*args, **kwargs)
[default4]:[rank4]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank4]: return forward_call(*args, **kwargs)
[default4]:[rank4]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/nn.py", line 159, in forward
[default4]:[rank4]: return row_linear(
[default4]:[rank4]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/functional.py", line 474, in row_linear
[default4]:[rank4]: out = F.linear(input, weight, bias)
[default4]:[rank4]: torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 4.00 GiB. GPU  has a total capacity of 79.33 GiB of which 1.89 GiB is free. Including non-PyTorch memory, this process has 77.42 GiB memory in use. Of the allocated memory 67.94 GiB is allocated by PyTorch, and 436.32 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
[default1]:[rank1]: Traceback (most recent call last):
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default1]:[rank1]: trainer.train(dataloader)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default1]:[rank1]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default1]:[rank1]: outputs = self.pipeline_engine.train_batch_iter(
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default1]:[rank1]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default1]:[rank1]: output = model(**micro_batch)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank1]: return self._call_impl(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank1]: return forward_call(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default1]:[rank1]: sharded_logits = self.model(
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank1]: return self._call_impl(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank1]: return forward_call(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default1]:[rank1]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default1]:[rank1]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank1]: return self._call_impl(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank1]: return forward_call(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 151, in forward
[default1]:[rank1]: output = self.pp_block(**new_kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank1]: return self._call_impl(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank1]: return forward_call(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 637, in forward
[default1]:[rank1]: hidden_states = self.mlp(hidden_states=hidden_states)["hidden_states"]
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank1]: return self._call_impl(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank1]: return forward_call(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 172, in forward
[default1]:[rank1]: hidden_states = self.down_proj(self.split_silu_mul(merged_states))
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank1]: return self._call_impl(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank1]: return forward_call(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/nn.py", line 159, in forward
[default1]:[rank1]: return row_linear(
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/functional.py", line 474, in row_linear
[default1]:[rank1]: out = F.linear(input, weight, bias)
[default1]:[rank1]: torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 4.00 GiB. GPU  has a total capacity of 79.33 GiB of which 2.16 GiB is free. Including non-PyTorch memory, this process has 77.16 GiB memory in use. Of the allocated memory 67.94 GiB is allocated by PyTorch, and 436.32 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
[default2]:[rank2]: Traceback (most recent call last):
[default2]:[rank2]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default2]:[rank2]: trainer.train(dataloader)
[default2]:[rank2]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default2]:[rank2]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default2]:[rank2]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default2]:[rank2]: outputs = self.pipeline_engine.train_batch_iter(
[default2]:[rank2]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default2]:[rank2]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default2]:[rank2]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default2]:[rank2]: output = model(**micro_batch)
[default2]:[rank2]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank2]: return self._call_impl(*args, **kwargs)
[default2]:[rank2]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default2]:[rank2]: return forward_call(*args, **kwargs)
[default2]:[rank2]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default2]:[rank2]: sharded_logits = self.model(
[default2]:[rank2]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank2]: return self._call_impl(*args, **kwargs)
[default2]:[rank2]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default2]:[rank2]: return forward_call(*args, **kwargs)
[default2]:[rank2]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default2]:[rank2]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default2]:[rank2]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default2]:[rank2]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default2]:[rank2]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank2]: return self._call_impl(*args, **kwargs)
[default2]:[rank2]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default2]:[rank2]: return forward_call(*args, **kwargs)
[default2]:[rank2]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 151, in forward
[default2]:[rank2]: output = self.pp_block(**new_kwargs)
[default2]:[rank2]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank2]: return self._call_impl(*args, **kwargs)
[default2]:[rank2]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default2]:[rank2]: return forward_call(*args, **kwargs)
[default2]:[rank2]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 637, in forward
[default2]:[rank2]: hidden_states = self.mlp(hidden_states=hidden_states)["hidden_states"]
[default2]:[rank2]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank2]: return self._call_impl(*args, **kwargs)
[default2]:[rank2]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default2]:[rank2]: return forward_call(*args, **kwargs)
[default2]:[rank2]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 172, in forward
[default2]:[rank2]: hidden_states = self.down_proj(self.split_silu_mul(merged_states))
[default2]:[rank2]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank2]: return self._call_impl(*args, **kwargs)
[default2]:[rank2]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default2]:[rank2]: return forward_call(*args, **kwargs)
[default2]:[rank2]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/nn.py", line 159, in forward
[default2]:[rank2]: return row_linear(
[default2]:[rank2]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/functional.py", line 474, in row_linear
[default2]:[rank2]: out = F.linear(input, weight, bias)
[default2]:[rank2]: torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 4.00 GiB. GPU  has a total capacity of 79.33 GiB of which 2.05 GiB is free. Including non-PyTorch memory, this process has 77.27 GiB memory in use. Of the allocated memory 67.94 GiB is allocated by PyTorch, and 436.32 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
[default5]:[rank5]: Traceback (most recent call last):
[default5]:[rank5]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default5]:[rank5]: trainer.train(dataloader)
[default5]:[rank5]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default5]:[rank5]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default5]:[rank5]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default5]:[rank5]: outputs = self.pipeline_engine.train_batch_iter(
[default5]:[rank5]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default5]:[rank5]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default5]:[rank5]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default5]:[rank5]: output = model(**micro_batch)
[default5]:[rank5]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default5]:[rank5]: return self._call_impl(*args, **kwargs)
[default5]:[rank5]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default5]:[rank5]: return forward_call(*args, **kwargs)
[default5]:[rank5]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default5]:[rank5]: sharded_logits = self.model(
[default5]:[rank5]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default5]:[rank5]: return self._call_impl(*args, **kwargs)
[default5]:[rank5]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default5]:[rank5]: return forward_call(*args, **kwargs)
[default5]:[rank5]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default5]:[rank5]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default5]:[rank5]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default5]:[rank5]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default5]:[rank5]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default5]:[rank5]: return self._call_impl(*args, **kwargs)
[default5]:[rank5]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default5]:[rank5]: return forward_call(*args, **kwargs)
[default5]:[rank5]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 151, in forward
[default5]:[rank5]: output = self.pp_block(**new_kwargs)
[default5]:[rank5]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default5]:[rank5]: return self._call_impl(*args, **kwargs)
[default5]:[rank5]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default5]:[rank5]: return forward_call(*args, **kwargs)
[default5]:[rank5]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 637, in forward
[default5]:[rank5]: hidden_states = self.mlp(hidden_states=hidden_states)["hidden_states"]
[default5]:[rank5]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default5]:[rank5]: return self._call_impl(*args, **kwargs)
[default5]:[rank5]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default5]:[rank5]: return forward_call(*args, **kwargs)
[default5]:[rank5]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 172, in forward
[default5]:[rank5]: hidden_states = self.down_proj(self.split_silu_mul(merged_states))
[default5]:[rank5]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default5]:[rank5]: return self._call_impl(*args, **kwargs)
[default5]:[rank5]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default5]:[rank5]: return forward_call(*args, **kwargs)
[default5]:[rank5]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/nn.py", line 159, in forward
[default5]:[rank5]: return row_linear(
[default5]:[rank5]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/functional.py", line 474, in row_linear
[default5]:[rank5]: out = F.linear(input, weight, bias)
[default5]:[rank5]: torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 4.00 GiB. GPU  has a total capacity of 79.33 GiB of which 1.89 GiB is free. Including non-PyTorch memory, this process has 77.42 GiB memory in use. Of the allocated memory 67.94 GiB is allocated by PyTorch, and 436.32 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
[default7]:[rank7]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default7]:[rank7]: return forward_call(*args, **kwargs)
[default7]:[rank7]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 637, in forward
[default7]:[rank7]: hidden_states = self.mlp(hidden_states=hidden_states)["hidden_states"]
[default7]:[rank7]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default7]:[rank7]: return self._call_impl(*args, **kwargs)
[default7]:[rank7]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default7]:[rank7]: return forward_call(*args, **kwargs)
[default7]:[rank7]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 172, in forward
[default7]:[rank7]: hidden_states = self.down_proj(self.split_silu_mul(merged_states))
[default7]:[rank7]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default7]:[rank7]: return self._call_impl(*args, **kwargs)
[default7]:[rank7]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default7]:[rank7]: return forward_call(*args, **kwargs)
[default7]:[rank7]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/nn.py", line 159, in forward
[default7]:[rank7]: return row_linear(
[default7]:[rank7]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/functional.py", line 474, in row_linear
[default7]:[rank7]: out = F.linear(input, weight, bias)
[default7]:[rank7]: torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 4.00 GiB. GPU  has a total capacity of 79.33 GiB of which 2.17 GiB is free. Including non-PyTorch memory, this process has 77.14 GiB memory in use. Of the allocated memory 67.94 GiB is allocated by PyTorch, and 436.32 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
[default0]:[rank0]: Traceback (most recent call last):
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default0]:[rank0]: trainer.train(dataloader)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default0]:[rank0]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default0]:[rank0]: outputs = self.pipeline_engine.train_batch_iter(
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default0]:[rank0]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default0]:[rank0]: output = model(**micro_batch)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank0]: return self._call_impl(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank0]: return forward_call(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default0]:[rank0]: sharded_logits = self.model(
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank0]: return self._call_impl(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank0]: return forward_call(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default0]:[rank0]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default0]:[rank0]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank0]: return self._call_impl(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank0]: return forward_call(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 151, in forward
[default0]:[rank0]: output = self.pp_block(**new_kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank0]: return self._call_impl(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank0]: return forward_call(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 637, in forward
[default0]:[rank0]: hidden_states = self.mlp(hidden_states=hidden_states)["hidden_states"]
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank0]: return self._call_impl(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank0]: return forward_call(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 172, in forward
[default0]:[rank0]: hidden_states = self.down_proj(self.split_silu_mul(merged_states))
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank0]: return self._call_impl(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank0]: return forward_call(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/nn.py", line 159, in forward
[default0]:[rank0]: return row_linear(
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/functional.py", line 474, in row_linear
[default0]:[rank0]: out = F.linear(input, weight, bias)
[default0]:[rank0]: torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 4.00 GiB. GPU
[default3]:[rank3]: Traceback (most recent call last):
[default3]:[rank3]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default3]:[rank3]: trainer.train(dataloader)
[default3]:[rank3]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default3]:[rank3]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default3]:[rank3]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default3]:[rank3]: outputs = self.pipeline_engine.train_batch_iter(
[default3]:[rank3]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default3]:[rank3]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default3]:[rank3]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default3]:[rank3]: output = model(**micro_batch)
[default3]:[rank3]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default3]:[rank3]: return self._call_impl(*args, **kwargs)
[default3]:[rank3]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default3]:[rank3]: return forward_call(*args, **kwargs)
[default3]:[rank3]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default3]:[rank3]: sharded_logits = self.model(
[default3]:[rank3]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default3]:[rank3]: return self._call_impl(*args, **kwargs)
[default3]:[rank3]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default3]:[rank3]: return forward_call(*args, **kwargs)
[default3]:[rank3]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default3]:[rank3]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default3]:[rank3]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default3]:[rank3]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default3]:[rank3]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default3]:[rank3]: return self._call_impl(*args, **kwargs)
[default3]:[rank3]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default3]:[rank3]: return forward_call(*args, **kwargs)
[default3]:[rank3]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 151, in forward
[default3]:[rank3]: output = self.pp_block(**new_kwargs)
[default3]:[rank3]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default3]:[rank3]: return self._call_impl(*args, **kwargs)
[default3]:[rank3]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default3]:[rank3]: return forward_call(*args, **kwargs)
[default3]:[rank3]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 637, in forward
[default3]:[rank3]: hidden_states = self.mlp(hidden_states=hidden_states)["hidden_states"]
[default3]:[rank3]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default3]:[rank3]: return self._call_impl(*args, **kwargs)
[default3]:[rank3]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default3]:[rank3]: return forward_call(*args, **kwargs)
[default3]:[rank3]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 172, in forward
[default3]:[rank3]: hidden_states = self.down_proj(self.split_silu_mul(merged_states))
[default3]:[rank3]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default3]:[rank3]: return self._call_impl(*args, **kwargs)
[default3]:[rank3]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default3]:[rank3]: return forward_call(*args, **kwargs)
[default3]:[rank3]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/nn.py", line 159, in forward
[default3]:[rank3]: return row_linear(
[default3]:[rank3]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/functional.py", line 474, in row_linear
[default3]:[rank3]: out = F.linear(input, weight, bias)
[default3]:[rank3]: torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 4.00 GiB. GPU  has a total capacity of 79.33 GiB of which 2.02 GiB is free. Including non-PyTorch memory, this process has 77.30 GiB memory in use. Of the allocated memory 67.94 GiB is allocated by PyTorch, and 436.32 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
[default6]:[rank6]: Traceback (most recent call last):
[default6]:[rank6]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default6]:[rank6]: trainer.train(dataloader)
[default6]:[rank6]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default6]:[rank6]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default6]:[rank6]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default6]:[rank6]: outputs = self.pipeline_engine.train_batch_iter(
[default6]:[rank6]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default6]:[rank6]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default6]:[rank6]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default6]:[rank6]: output = model(**micro_batch)
[default6]:[rank6]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default6]:[rank6]: return self._call_impl(*args, **kwargs)
[default6]:[rank6]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default6]:[rank6]: return forward_call(*args, **kwargs)
[default6]:[rank6]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default6]:[rank6]: sharded_logits = self.model(
[default6]:[rank6]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default6]:[rank6]: return self._call_impl(*args, **kwargs)
[default6]:[rank6]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default6]:[rank6]: return forward_call(*args, **kwargs)
[default6]:[rank6]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default6]:[rank6]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default6]:[rank6]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default6]:[rank6]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default6]:[rank6]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default6]:[rank6]: return self._call_impl(*args, **kwargs)
[default6]:[rank6]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default6]:[rank6]: return forward_call(*args, **kwargs)
[default6]:[rank6]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 151, in forward
[default6]:[rank6]: output = self.pp_block(**new_kwargs)
[default6]:[rank6]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default6]:[rank6]: return self._call_impl(*args, **kwargs)
[default6]:[rank6]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default6]:[rank6]: return forward_call(*args, **kwargs)
[default6]:[rank6]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 637, in forward
[default6]:[rank6]: hidden_states = self.mlp(hidden_states=hidden_states)["hidden_states"]
[default6]:[rank6]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default6]:[rank6]: return self._call_impl(*args, **kwargs)
[default6]:[rank6]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default6]:[rank6]: return forward_call(*args, **kwargs)
[default6]:[rank6]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 172, in forward
[default6]:[rank6]: hidden_states = self.down_proj(self.split_silu_mul(merged_states))
[default6]:[rank6]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default6]:[rank6]: return self._call_impl(*args, **kwargs)
[default6]:[rank6]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default6]:[rank6]: return forward_call(*args, **kwargs)
[default6]:[rank6]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/nn.py", line 159, in forward
[default6]:[rank6]: return row_linear(
[default6]:[rank6]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/functional.py", line 474, in row_linear
[default6]:[rank6]: out = F.linear(input, weight, bias)
[default6]:[rank6]: torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 4.00 GiB. GPU  has a total capacity of 79.33 GiB of which 1.89 GiB is free. Including non-PyTorch memory, this process has 77.42 GiB memory in use. Of the allocated memory 67.94 GiB is allocated by PyTorch, and 436.32 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
[default6]:[rank14]: Traceback (most recent call last):
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default6]:[rank14]: trainer.train(dataloader)
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default6]:[rank14]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default6]:[rank14]: outputs = self.pipeline_engine.train_batch_iter(
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter
[default6]:[rank14]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default6]:[rank14]: output = model(**micro_batch)
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default6]:[rank14]: return self._call_impl(*args, **kwargs)
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default6]:[rank14]: return forward_call(*args, **kwargs)
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default6]:[rank14]: sharded_logits = self.model(
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default6]:[rank14]: return self._call_impl(*args, **kwargs)
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default6]:[rank14]: return forward_call(*args, **kwargs)
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default6]:[rank14]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default6]:[rank14]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default6]:[rank14]: return self._call_impl(*args, **kwargs)
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default6]:[rank14]: return forward_call(*args, **kwargs)
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default6]:[rank14]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default6]:[rank14]: pipeline_state.run_communication()
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default6]:[rank14]: recv_activation_tensor = recv_activation()
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default6]:[rank14]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default6]:[rank14]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default6]:[rank14]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default6]:[rank14]: dist.recv(
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default6]:[rank14]: return func(*args, **kwargs)
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default6]:[rank14]: pg.recv([tensor], group_src_rank, tag).wait()
[default6]:[rank14]: torch.distributed.DistBackendError: [1] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '0:1', but store->get('0:1') got error: Connection reset by peer
[default6]:[rank14]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default6]:[rank14]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7fa2d8e4b897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default6]:[rank14]: frame #1: <unknown function> + 0x5b3a23e (0x7fa31296823e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7fa312962c87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7fa312962f82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7fa312963fd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fa312918371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fa312918371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fa312918371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fa312918371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7fa2da125189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default6]:[rank14]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7fa2da12c610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default6]:[rank14]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7fa2da14b978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default6]:[rank14]: frame #12: <unknown function> + 0x5adc309 (0x7fa31290a309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #13: <unknown function> + 0x5ae6f10 (0x7fa312914f10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #14: <unknown function> + 0x5ae6fa5 (0x7fa312914fa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #15: <unknown function> + 0x5124446 (0x7fa311f52446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #16: <unknown function> + 0x1acf4b8 (0x7fa30e8fd4b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #17: <unknown function> + 0x5aee004 (0x7fa31291c004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #18: <unknown function> + 0x5af36b5 (0x7fa3129216b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #19: <unknown function> + 0xd2631e (0x7fa32550b31e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default6]:[rank14]: frame #20: <unknown function> + 0x47def4 (0x7fa324c62ef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default6]:[rank14]: frame #21: <unknown function> + 0x1445a6 (0x55fc016c45a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #22: _PyObject_MakeTpCall + 0x26b (0x55fc016bda6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #23: <unknown function> + 0x150866 (0x55fc016d0866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x55fc016b9142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #25: _PyFunction_Vectorcall + 0x6c (0x55fc016c4a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #26: PyObject_Call + 0xbc (0x55fc016d0f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x55fc016b72b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #28: _PyFunction_Vectorcall + 0x6c (0x55fc016c4a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x55fc016b58fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #30: <unknown function> + 0x150582 (0x55fc016d0582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x55fc016b58fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #32: <unknown function> + 0x150582 (0x55fc016d0582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x55fc016b58fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #34: <unknown function> + 0x150582 (0x55fc016d0582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x55fc016b58fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x55fc016bcf50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #37: _PyObject_Call_Prepend + 0x69 (0x55fc016cec39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #38: <unknown function> + 0x211239 (0x55fc01791239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #39: _PyObject_MakeTpCall + 0x26b (0x55fc016bda6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x55fc016b93e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #41: _PyFunction_Vectorcall + 0x6c (0x55fc016c4a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x55fc016b4c5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #43: _PyFunction_Vectorcall + 0x6c (0x55fc016c4a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x55fc016b58fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #45: <unknown function> + 0x150582 (0x55fc016d0582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #46: PyObject_Call + 0xbc (0x55fc016d0f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x55fc016b72b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #48: <unknown function> + 0x150582 (0x55fc016d0582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #49: PyObject_Call + 0xbc (0x55fc016d0f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x55fc016b72b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #51: _PyFunction_Vectorcall + 0x6c (0x55fc016c4a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x55fc016bd007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #53: _PyObject_Call_Prepend + 0x69 (0x55fc016cec39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #54: <unknown function> + 0x211239 (0x55fc01791239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #55: PyObject_Call + 0x207 (0x55fc016d1067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x55fc016b72b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #57: <unknown function> + 0x150582 (0x55fc016d0582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x55fc016b58fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #59: <unknown function> + 0x150582 (0x55fc016d0582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #60: PyObject_Call + 0xbc (0x55fc016d0f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x55fc016b72b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #62: <unknown function> + 0x150582 (0x55fc016d0582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #63: PyObject_Call + 0xbc (0x55fc016d0f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: . This may indicate a possible application crash on rank 0 or a network set up issue.
[default5]:[rank13]: Traceback (most recent call last):
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default5]:[rank13]: trainer.train(dataloader)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default5]:[rank13]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default5]:[rank13]: outputs = self.pipeline_engine.train_batch_iter(
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter
[default5]:[rank13]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default5]:[rank13]: output = model(**micro_batch)
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default5]:[rank13]: return self._call_impl(*args, **kwargs)
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default5]:[rank13]: return forward_call(*args, **kwargs)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default5]:[rank13]: sharded_logits = self.model(
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default5]:[rank13]: return self._call_impl(*args, **kwargs)
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default5]:[rank13]: return forward_call(*args, **kwargs)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default5]:[rank13]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default5]:[rank13]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default5]:[rank13]: return self._call_impl(*args, **kwargs)
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default5]:[rank13]: return forward_call(*args, **kwargs)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default5]:[rank13]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default5]:[rank13]: pipeline_state.run_communication()
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default5]:[rank13]: recv_activation_tensor = recv_activation()
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default5]:[rank13]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default5]:[rank13]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default5]:[rank13]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default5]:[rank13]: dist.recv(
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default5]:[rank13]: return func(*args, **kwargs)
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default5]:[rank13]: pg.recv([tensor], group_src_rank, tag).wait()
[default5]:[rank13]: torch.distributed.DistBackendError: [1] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '0:1', but store->get('0:1') got error: Connection reset by peer
[default5]:[rank13]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default5]:[rank13]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7f5656262897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default5]:[rank13]: frame #1: <unknown function> + 0x5b3a23e (0x7f568fd7f23e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7f568fd79c87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7f568fd79f82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7f568fd7afd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f568fd2f371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f568fd2f371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f568fd2f371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f568fd2f371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7f565753c189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default5]:[rank13]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7f5657543610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default5]:[rank13]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7f5657562978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default5]:[rank13]: frame #12: <unknown function> + 0x5adc309 (0x7f568fd21309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #13: <unknown function> + 0x5ae6f10 (0x7f568fd2bf10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #14: <unknown function> + 0x5ae6fa5 (0x7f568fd2bfa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #15: <unknown function> + 0x5124446 (0x7f568f369446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #16: <unknown function> + 0x1acf4b8 (0x7f568bd144b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #17: <unknown function> + 0x5aee004 (0x7f568fd33004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #18: <unknown function> + 0x5af36b5 (0x7f568fd386b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #19: <unknown function> + 0xd2631e (0x7f56a292231e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default5]:[rank13]: frame #20: <unknown function> + 0x47def4 (0x7f56a2079ef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default5]:[rank13]: frame #21: <unknown function> + 0x1445a6 (0x5655184bf5a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #22: _PyObject_MakeTpCall + 0x26b (0x5655184b8a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #23: <unknown function> + 0x150866 (0x5655184cb866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x5655184b4142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #25: _PyFunction_Vectorcall + 0x6c (0x5655184bfa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #26: PyObject_Call + 0xbc (0x5655184cbf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x5655184b22b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #28: _PyFunction_Vectorcall + 0x6c (0x5655184bfa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x5655184b08fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #30: <unknown function> + 0x150582 (0x5655184cb582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x5655184b08fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #32: <unknown function> + 0x150582 (0x5655184cb582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x5655184b08fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #34: <unknown function> + 0x150582 (0x5655184cb582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x5655184b08fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x5655184b7f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #37: _PyObject_Call_Prepend + 0x69 (0x5655184c9c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #38: <unknown function> + 0x211239 (0x56551858c239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #39: _PyObject_MakeTpCall + 0x26b (0x5655184b8a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x5655184b43e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #41: _PyFunction_Vectorcall + 0x6c (0x5655184bfa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x5655184afc5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #43: _PyFunction_Vectorcall + 0x6c (0x5655184bfa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x5655184b08fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #45: <unknown function> + 0x150582 (0x5655184cb582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #46: PyObject_Call + 0xbc (0x5655184cbf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x5655184b22b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #48: <unknown function> + 0x150582 (0x5655184cb582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #49: PyObject_Call + 0xbc (0x5655184cbf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x5655184b22b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #51: _PyFunction_Vectorcall + 0x6c (0x5655184bfa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x5655184b8007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #53: _PyObject_Call_Prepend + 0x69 (0x5655184c9c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #54: <unknown function> + 0x211239 (0x56551858c239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #55: PyObject_Call + 0x207 (0x5655184cc067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x5655184b22b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #57: <unknown function> + 0x150582 (0x5655184cb582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x5655184b08fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #59: <unknown function> + 0x150582 (0x5655184cb582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #60: PyObject_Call + 0xbc (0x5655184cbf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x5655184b22b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #62: <unknown function> + 0x150582 (0x5655184cb582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #63: PyObject_Call + 0xbc (0x5655184cbf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: . This may indicate a possible application crash on rank 0 or a network set up issue.
[default1]:[rank9]: Traceback (most recent call last):
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default1]:[rank9]: trainer.train(dataloader)
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default1]:[rank9]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default1]:[rank9]: outputs = self.pipeline_engine.train_batch_iter(
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter
[default1]:[rank9]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default1]:[rank9]: output = model(**micro_batch)
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank9]: return self._call_impl(*args, **kwargs)
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank9]: return forward_call(*args, **kwargs)
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default1]:[rank9]: sharded_logits = self.model(
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank9]: return self._call_impl(*args, **kwargs)
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank9]: return forward_call(*args, **kwargs)
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default1]:[rank9]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default1]:[rank9]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank9]: return self._call_impl(*args, **kwargs)
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank9]: return forward_call(*args, **kwargs)
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default1]:[rank9]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default1]:[rank9]: pipeline_state.run_communication()
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default1]:[rank9]: recv_activation_tensor = recv_activation()
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default1]:[rank9]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default1]:[rank9]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default1]:[rank9]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default1]:[rank9]: dist.recv(
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default1]:[rank9]: return func(*args, **kwargs)
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default1]:[rank9]: pg.recv([tensor], group_src_rank, tag).wait()
[default1]:[rank9]: torch.distributed.DistBackendError: [1] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '0:1', but store->get('0:1') got error: Connection reset by peer
[default1]:[rank9]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default1]:[rank9]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7fafba5f5897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default1]:[rank9]: frame #1: <unknown function> + 0x5b3a23e (0x7faff411223e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7faff410cc87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7faff410cf82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7faff410dfd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7faff40c2371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7faff40c2371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7faff40c2371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7faff40c2371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7fafbb8cf189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default1]:[rank9]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7fafbb8d6610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default1]:[rank9]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7fafbb8f5978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default1]:[rank9]: frame #12: <unknown function> + 0x5adc309 (0x7faff40b4309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #13: <unknown function> + 0x5ae6f10 (0x7faff40bef10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #14: <unknown function> + 0x5ae6fa5 (0x7faff40befa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #15: <unknown function> + 0x5124446 (0x7faff36fc446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #16: <unknown function> + 0x1acf4b8 (0x7faff00a74b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #17: <unknown function> + 0x5aee004 (0x7faff40c6004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #18: <unknown function> + 0x5af36b5 (0x7faff40cb6b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default1]:[rank9]: frame #19: <unknown function> + 0xd2631e (0x7fb006cb531e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default1]:[rank9]: frame #20: <unknown function> + 0x47def4 (0x7fb00640cef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default1]:[rank9]: frame #21: <unknown function> + 0x1445a6 (0x558c40f3f5a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #22: _PyObject_MakeTpCall + 0x26b (0x558c40f38a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #23: <unknown function> + 0x150866 (0x558c40f4b866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x558c40f34142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #25: _PyFunction_Vectorcall + 0x6c (0x558c40f3fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #26: PyObject_Call + 0xbc (0x558c40f4bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x558c40f322b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #28: _PyFunction_Vectorcall + 0x6c (0x558c40f3fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x558c40f308fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #30: <unknown function> + 0x150582 (0x558c40f4b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x558c40f308fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #32: <unknown function> + 0x150582 (0x558c40f4b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x558c40f308fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #34: <unknown function> + 0x150582 (0x558c40f4b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x558c40f308fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x558c40f37f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #37: _PyObject_Call_Prepend + 0x69 (0x558c40f49c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #38: <unknown function> + 0x211239 (0x558c4100c239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #39: _PyObject_MakeTpCall + 0x26b (0x558c40f38a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x558c40f343e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #41: _PyFunction_Vectorcall + 0x6c (0x558c40f3fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x558c40f2fc5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #43: _PyFunction_Vectorcall + 0x6c (0x558c40f3fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x558c40f308fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #45: <unknown function> + 0x150582 (0x558c40f4b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #46: PyObject_Call + 0xbc (0x558c40f4bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x558c40f322b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #48: <unknown function> + 0x150582 (0x558c40f4b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #49: PyObject_Call + 0xbc (0x558c40f4bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x558c40f322b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #51: _PyFunction_Vectorcall + 0x6c (0x558c40f3fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x558c40f38007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #53: _PyObject_Call_Prepend + 0x69 (0x558c40f49c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #54: <unknown function> + 0x211239 (0x558c4100c239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #55: PyObject_Call + 0x207 (0x558c40f4c067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x558c40f322b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #57: <unknown function> + 0x150582 (0x558c40f4b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x558c40f308fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #59: <unknown function> + 0x150582 (0x558c40f4b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #60: PyObject_Call + 0xbc (0x558c40f4bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x558c40f322b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #62: <unknown function> + 0x150582 (0x558c40f4b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: frame #63: PyObject_Call + 0xbc (0x558c40f4bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default1]:[rank9]: . This may indicate a possible application crash on rank 0 or a network set up issue.
[default0]:[rank8]: Traceback (most recent call last):
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default0]:[rank8]: trainer.train(dataloader)
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default0]:[rank8]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default0]:[rank8]: outputs = self.pipeline_engine.train_batch_iter(
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter
[default0]:[rank8]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default0]:[rank8]: output = model(**micro_batch)
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank8]: return self._call_impl(*args, **kwargs)
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank8]: return forward_call(*args, **kwargs)
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default0]:[rank8]: sharded_logits = self.model(
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank8]: return self._call_impl(*args, **kwargs)
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank8]: return forward_call(*args, **kwargs)
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default0]:[rank8]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default0]:[rank8]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank8]: return self._call_impl(*args, **kwargs)
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank8]: return forward_call(*args, **kwargs)
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default0]:[rank8]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default0]:[rank8]: pipeline_state.run_communication()
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default0]:[rank8]: recv_activation_tensor = recv_activation()
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default0]:[rank8]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default0]:[rank8]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default0]:[rank8]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default0]:[rank8]: dist.recv(
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default0]:[rank8]: return func(*args, **kwargs)
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default0]:[rank8]: pg.recv([tensor], group_src_rank, tag).wait()
[default0]:[rank8]: torch.distributed.DistBackendError: [1] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '0:1', but store->get('0:1') got error: Connection reset by peer
[default0]:[rank8]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default0]:[rank8]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7f405424e897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default0]:[rank8]: frame #1: <unknown function> + 0x5b3a23e (0x7f408dd6b23e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7f408dd65c87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7f408dd65f82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7f408dd66fd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f408dd1b371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f408dd1b371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f408dd1b371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f408dd1b371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7f4055528189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default0]:[rank8]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7f405552f610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default0]:[rank8]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7f405554e978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default0]:[rank8]: frame #12: <unknown function> + 0x5adc309 (0x7f408dd0d309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #13: <unknown function> + 0x5ae6f10 (0x7f408dd17f10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #14: <unknown function> + 0x5ae6fa5 (0x7f408dd17fa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #15: <unknown function> + 0x5124446 (0x7f408d355446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #16: <unknown function> + 0x1acf4b8 (0x7f4089d004b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #17: <unknown function> + 0x5aee004 (0x7f408dd1f004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #18: <unknown function> + 0x5af36b5 (0x7f408dd246b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default0]:[rank8]: frame #19: <unknown function> + 0xd2631e (0x7f40a090e31e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default0]:[rank8]: frame #20: <unknown function> + 0x47def4 (0x7f40a0065ef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default0]:[rank8]: frame #21: <unknown function> + 0x1445a6 (0x55c7768495a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #22: _PyObject_MakeTpCall + 0x26b (0x55c776842a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #23: <unknown function> + 0x150866 (0x55c776855866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x55c77683e142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #25: _PyFunction_Vectorcall + 0x6c (0x55c776849a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #26: PyObject_Call + 0xbc (0x55c776855f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x55c77683c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #28: _PyFunction_Vectorcall + 0x6c (0x55c776849a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x55c77683a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #30: <unknown function> + 0x150582 (0x55c776855582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x55c77683a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #32: <unknown function> + 0x150582 (0x55c776855582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x55c77683a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #34: <unknown function> + 0x150582 (0x55c776855582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x55c77683a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x55c776841f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #37: _PyObject_Call_Prepend + 0x69 (0x55c776853c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #38: <unknown function> + 0x211239 (0x55c776916239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #39: _PyObject_MakeTpCall + 0x26b (0x55c776842a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x55c77683e3e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #41: _PyFunction_Vectorcall + 0x6c (0x55c776849a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x55c776839c5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #43: _PyFunction_Vectorcall + 0x6c (0x55c776849a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x55c77683a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #45: <unknown function> + 0x150582 (0x55c776855582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #46: PyObject_Call + 0xbc (0x55c776855f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x55c77683c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #48: <unknown function> + 0x150582 (0x55c776855582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #49: PyObject_Call + 0xbc (0x55c776855f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x55c77683c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #51: _PyFunction_Vectorcall + 0x6c (0x55c776849a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x55c776842007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #53: _PyObject_Call_Prepend + 0x69 (0x55c776853c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #54: <unknown function> + 0x211239 (0x55c776916239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #55: PyObject_Call + 0x207 (0x55c776856067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x55c77683c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #57: <unknown function> + 0x150582 (0x55c776855582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x55c77683a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #59: <unknown function> + 0x150582 (0x55c776855582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #60: PyObject_Call + 0xbc (0x55c776855f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x55c77683c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #62: <unknown function> + 0x150582 (0x55c776855582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: frame #63: PyObject_Call + 0xbc (0x55c776855f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default0]:[rank8]: . This may indicate a possible application crash on rank 0 or a network set up issue.
[default2]:[rank10]: Traceback (most recent call last):
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default2]:[rank10]: trainer.train(dataloader)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default2]:[rank10]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default2]:[rank10]: outputs = self.pipeline_engine.train_batch_iter(
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter
[default2]:[rank10]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default2]:[rank10]: output = model(**micro_batch)
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank10]: return self._call_impl(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default2]:[rank10]: return forward_call(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default2]:[rank10]: sharded_logits = self.model(
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank10]: return self._call_impl(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default2]:[rank10]: return forward_call(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default2]:[rank10]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default2]:[rank10]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank10]: return self._call_impl(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default2]:[rank10]: return forward_call(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default2]:[rank10]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default2]:[rank10]: pipeline_state.run_communication()
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default2]:[rank10]: recv_activation_tensor = recv_activation()
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default2]:[rank10]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default2]:[rank10]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default2]:[rank10]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default2]:[rank10]: dist.recv(
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default2]:[rank10]: return func(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default2]:[rank10]: pg.recv([tensor], group_src_rank, tag).wait()
[default2]:[rank10]: torch.distributed.DistBackendError: [1] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '0:1', but store->get('0:1') got error: Connection reset by peer
[default2]:[rank10]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default2]:[rank10]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7f280f4f3897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default2]:[rank10]: frame #1: <unknown function> + 0x5b3a23e (0x7f284901023e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7f284900ac87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7f284900af82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7f284900bfd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f2848fc0371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f2848fc0371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f2848fc0371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f2848fc0371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7f28107cd189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default2]:[rank10]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7f28107d4610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default2]:[rank10]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7f28107f3978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default2]:[rank10]: frame #12: <unknown function> + 0x5adc309 (0x7f2848fb2309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #13: <unknown function> + 0x5ae6f10 (0x7f2848fbcf10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #14: <unknown function> + 0x5ae6fa5 (0x7f2848fbcfa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #15: <unknown function> + 0x5124446 (0x7f28485fa446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #16: <unknown function> + 0x1acf4b8 (0x7f2844fa54b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #17: <unknown function> + 0x5aee004 (0x7f2848fc4004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #18: <unknown function> + 0x5af36b5 (0x7f2848fc96b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #19: <unknown function> + 0xd2631e (0x7f285bbb331e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default2]:[rank10]: frame #20: <unknown function> + 0x47def4 (0x7f285b30aef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default2]:[rank10]: frame #21: <unknown function> + 0x1445a6 (0x55ec71a8c5a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #22: _PyObject_MakeTpCall + 0x26b (0x55ec71a85a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #23: <unknown function> + 0x150866 (0x55ec71a98866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x55ec71a81142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #25: _PyFunction_Vectorcall + 0x6c (0x55ec71a8ca2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #26: PyObject_Call + 0xbc (0x55ec71a98f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x55ec71a7f2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #28: _PyFunction_Vectorcall + 0x6c (0x55ec71a8ca2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x55ec71a7d8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #30: <unknown function> + 0x150582 (0x55ec71a98582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x55ec71a7d8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #32: <unknown function> + 0x150582 (0x55ec71a98582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x55ec71a7d8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #34: <unknown function> + 0x150582 (0x55ec71a98582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x55ec71a7d8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x55ec71a84f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #37: _PyObject_Call_Prepend + 0x69 (0x55ec71a96c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #38: <unknown function> + 0x211239 (0x55ec71b59239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #39: _PyObject_MakeTpCall + 0x26b (0x55ec71a85a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x55ec71a813e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #41: _PyFunction_Vectorcall + 0x6c (0x55ec71a8ca2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x55ec71a7cc5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #43: _PyFunction_Vectorcall + 0x6c (0x55ec71a8ca2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x55ec71a7d8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #45: <unknown function> + 0x150582 (0x55ec71a98582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #46: PyObject_Call + 0xbc (0x55ec71a98f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x55ec71a7f2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #48: <unknown function> + 0x150582 (0x55ec71a98582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #49: PyObject_Call + 0xbc (0x55ec71a98f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x55ec71a7f2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #51: _PyFunction_Vectorcall + 0x6c (0x55ec71a8ca2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x55ec71a85007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #53: _PyObject_Call_Prepend + 0x69 (0x55ec71a96c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #54: <unknown function> + 0x211239 (0x55ec71b59239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #55: PyObject_Call + 0x207 (0x55ec71a99067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x55ec71a7f2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #57: <unknown function> + 0x150582 (0x55ec71a98582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x55ec71a7d8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #59: <unknown function> + 0x150582 (0x55ec71a98582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #60: PyObject_Call + 0xbc (0x55ec71a98f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x55ec71a7f2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #62: <unknown function> + 0x150582 (0x55ec71a98582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #63: PyObject_Call + 0xbc (0x55ec71a98f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: . This may indicate a possible application crash on rank 0 or a network set up issue.
[default7]:[rank15]: Traceback (most recent call last):
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default7]:[rank15]: trainer.train(dataloader)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default7]:[rank15]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default7]:[rank15]: outputs = self.pipeline_engine.train_batch_iter(
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter
[default7]:[rank15]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default7]:[rank15]: output = model(**micro_batch)
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default7]:[rank15]: return self._call_impl(*args, **kwargs)
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default7]:[rank15]: return forward_call(*args, **kwargs)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default7]:[rank15]: sharded_logits = self.model(
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default7]:[rank15]: return self._call_impl(*args, **kwargs)
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default7]:[rank15]: return forward_call(*args, **kwargs)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default7]:[rank15]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default7]:[rank15]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default7]:[rank15]: return self._call_impl(*args, **kwargs)
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default7]:[rank15]: return forward_call(*args, **kwargs)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default7]:[rank15]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default7]:[rank15]: pipeline_state.run_communication()
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default7]:[rank15]: recv_activation_tensor = recv_activation()
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default7]:[rank15]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default7]:[rank15]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default7]:[rank15]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default7]:[rank15]: dist.recv(
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default7]:[rank15]: return func(*args, **kwargs)
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default7]:[rank15]: pg.recv([tensor], group_src_rank, tag).wait()
[default7]:[rank15]: torch.distributed.DistBackendError: [1] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '0:1', but store->get('0:1') got error: Connection reset by peer
[default7]:[rank15]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default7]:[rank15]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7f67cbb38897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default7]:[rank15]: frame #1: <unknown function> + 0x5b3a23e (0x7f680565523e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7f680564fc87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7f680564ff82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7f6805650fd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f6805605371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f6805605371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f6805605371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f6805605371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7f67cce12189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default7]:[rank15]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7f67cce19610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default7]:[rank15]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7f67cce38978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default7]:[rank15]: frame #12: <unknown function> + 0x5adc309 (0x7f68055f7309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #13: <unknown function> + 0x5ae6f10 (0x7f6805601f10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #14: <unknown function> + 0x5ae6fa5 (0x7f6805601fa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #15: <unknown function> + 0x5124446 (0x7f6804c3f446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #16: <unknown function> + 0x1acf4b8 (0x7f68015ea4b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #17: <unknown function> + 0x5aee004 (0x7f6805609004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #18: <unknown function> + 0x5af36b5 (0x7f680560e6b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #19: <unknown function> + 0xd2631e (0x7f68181f831e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default7]:[rank15]: frame #20: <unknown function> + 0x47def4 (0x7f681794fef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default7]:[rank15]: frame #21: <unknown function> + 0x1445a6 (0x55fd549895a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #22: _PyObject_MakeTpCall + 0x26b (0x55fd54982a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #23: <unknown function> + 0x150866 (0x55fd54995866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x55fd5497e142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #25: _PyFunction_Vectorcall + 0x6c (0x55fd54989a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #26: PyObject_Call + 0xbc (0x55fd54995f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x55fd5497c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #28: _PyFunction_Vectorcall + 0x6c (0x55fd54989a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x55fd5497a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #30: <unknown function> + 0x150582 (0x55fd54995582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x55fd5497a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #32: <unknown function> + 0x150582 (0x55fd54995582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x55fd5497a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #34: <unknown function> + 0x150582 (0x55fd54995582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x55fd5497a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x55fd54981f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #37: _PyObject_Call_Prepend + 0x69 (0x55fd54993c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #38: <unknown function> + 0x211239 (0x55fd54a56239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #39: _PyObject_MakeTpCall + 0x26b (0x55fd54982a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x55fd5497e3e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #41: _PyFunction_Vectorcall + 0x6c (0x55fd54989a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x55fd54979c5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #43: _PyFunction_Vectorcall + 0x6c (0x55fd54989a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x55fd5497a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #45: <unknown function> + 0x150582 (0x55fd54995582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #46: PyObject_Call + 0xbc (0x55fd54995f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x55fd5497c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #48: <unknown function> + 0x150582 (0x55fd54995582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #49: PyObject_Call + 0xbc (0x55fd54995f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x55fd5497c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #51: _PyFunction_Vectorcall + 0x6c (0x55fd54989a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x55fd54982007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #53: _PyObject_Call_Prepend + 0x69 (0x55fd54993c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #54: <unknown function> + 0x211239 (0x55fd54a56239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #55: PyObject_Call + 0x207 (0x55fd54996067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x55fd5497c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #57: <unknown function> + 0x150582 (0x55fd54995582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x55fd5497a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #59: <unknown function> + 0x150582 (0x55fd54995582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #60: PyObject_Call + 0xbc (0x55fd54995f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x55fd5497c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #62: <unknown function> + 0x150582 (0x55fd54995582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #63: PyObject_Call + 0xbc (0x55fd54995f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: . This may indicate a possible application crash on rank 0 or a network set up issue.
[default3]:[rank11]: Traceback (most recent call last):
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default3]:[rank11]: trainer.train(dataloader)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default3]:[rank11]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default3]:[rank11]: outputs = self.pipeline_engine.train_batch_iter(
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter
[default3]:[rank11]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default3]:[rank11]: output = model(**micro_batch)
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default3]:[rank11]: return self._call_impl(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default3]:[rank11]: return forward_call(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default3]:[rank11]: sharded_logits = self.model(
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default3]:[rank11]: return self._call_impl(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default3]:[rank11]: return forward_call(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default3]:[rank11]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default3]:[rank11]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default3]:[rank11]: return self._call_impl(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default3]:[rank11]: return forward_call(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default3]:[rank11]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default3]:[rank11]: pipeline_state.run_communication()
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default3]:[rank11]: recv_activation_tensor = recv_activation()
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default3]:[rank11]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default3]:[rank11]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default3]:[rank11]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default3]:[rank11]: dist.recv(
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default3]:[rank11]: return func(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default3]:[rank11]: pg.recv([tensor], group_src_rank, tag).wait()
[default3]:[rank11]: torch.distributed.DistBackendError: [1] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '0:1', but store->get('0:1') got error: Connection reset by peer
[default3]:[rank11]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default3]:[rank11]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7fd42658c897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default3]:[rank11]: frame #1: <unknown function> + 0x5b3a23e (0x7fd4600a923e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7fd4600a3c87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7fd4600a3f82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7fd4600a4fd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fd460059371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fd460059371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fd460059371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fd460059371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7fd427866189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default3]:[rank11]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7fd42786d610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default3]:[rank11]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7fd42788c978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default3]:[rank11]: frame #12: <unknown function> + 0x5adc309 (0x7fd46004b309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #13: <unknown function> + 0x5ae6f10 (0x7fd460055f10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #14: <unknown function> + 0x5ae6fa5 (0x7fd460055fa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #15: <unknown function> + 0x5124446 (0x7fd45f693446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #16: <unknown function> + 0x1acf4b8 (0x7fd45c03e4b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #17: <unknown function> + 0x5aee004 (0x7fd46005d004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #18: <unknown function> + 0x5af36b5 (0x7fd4600626b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #19: <unknown function> + 0xd2631e (0x7fd472c4c31e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default3]:[rank11]: frame #20: <unknown function> + 0x47def4 (0x7fd4723a3ef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default3]:[rank11]: frame #21: <unknown function> + 0x1445a6 (0x55855f8995a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #22: _PyObject_MakeTpCall + 0x26b (0x55855f892a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #23: <unknown function> + 0x150866 (0x55855f8a5866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x55855f88e142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #25: _PyFunction_Vectorcall + 0x6c (0x55855f899a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #26: PyObject_Call + 0xbc (0x55855f8a5f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x55855f88c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #28: _PyFunction_Vectorcall + 0x6c (0x55855f899a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x55855f88a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #30: <unknown function> + 0x150582 (0x55855f8a5582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x55855f88a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #32: <unknown function> + 0x150582 (0x55855f8a5582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x55855f88a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #34: <unknown function> + 0x150582 (0x55855f8a5582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x55855f88a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x55855f891f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #37: _PyObject_Call_Prepend + 0x69 (0x55855f8a3c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #38: <unknown function> + 0x211239 (0x55855f966239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #39: _PyObject_MakeTpCall + 0x26b (0x55855f892a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x55855f88e3e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #41: _PyFunction_Vectorcall + 0x6c (0x55855f899a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x55855f889c5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #43: _PyFunction_Vectorcall + 0x6c (0x55855f899a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x55855f88a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #45: <unknown function> + 0x150582 (0x55855f8a5582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #46: PyObject_Call + 0xbc (0x55855f8a5f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x55855f88c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #48: <unknown function> + 0x150582 (0x55855f8a5582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #49: PyObject_Call + 0xbc (0x55855f8a5f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x55855f88c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #51: _PyFunction_Vectorcall + 0x6c (0x55855f899a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x55855f892007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #53: _PyObject_Call_Prepend + 0x69 (0x55855f8a3c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #54: <unknown function> + 0x211239 (0x55855f966239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #55: PyObject_Call + 0x207 (0x55855f8a6067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x55855f88c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #57: <unknown function> + 0x150582 (0x55855f8a5582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x55855f88a8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #59: <unknown function> + 0x150582 (0x55855f8a5582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #60: PyObject_Call + 0xbc (0x55855f8a5f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x55855f88c2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #62: <unknown function> + 0x150582 (0x55855f8a5582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #63: PyObject_Call + 0xbc (0x55855f8a5f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: . This may indicate a possible application crash on rank 0 or a network set up issue.
[default4]:[rank12]: Traceback (most recent call last):
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default4]:[rank12]: trainer.train(dataloader)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default4]:[rank12]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default4]:[rank12]: outputs = self.pipeline_engine.train_batch_iter(
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter
[default4]:[rank12]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default4]:[rank12]: output = model(**micro_batch)
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default4]:[rank12]: return self._call_impl(*args, **kwargs)
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank12]: return forward_call(*args, **kwargs)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default4]:[rank12]: sharded_logits = self.model(
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default4]:[rank12]: return self._call_impl(*args, **kwargs)
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank12]: return forward_call(*args, **kwargs)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default4]:[rank12]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default4]:[rank12]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default4]:[rank12]: return self._call_impl(*args, **kwargs)
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank12]: return forward_call(*args, **kwargs)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default4]:[rank12]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default4]:[rank12]: pipeline_state.run_communication()
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default4]:[rank12]: recv_activation_tensor = recv_activation()
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default4]:[rank12]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default4]:[rank12]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default4]:[rank12]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default4]:[rank12]: dist.recv(
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default4]:[rank12]: return func(*args, **kwargs)
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default4]:[rank12]: pg.recv([tensor], group_src_rank, tag).wait()
[default4]:[rank12]: torch.distributed.DistBackendError: [1] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '0:1', but store->get('0:1') got error: Connection reset by peer
[default4]:[rank12]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default4]:[rank12]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7effe6c65897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default4]:[rank12]: frame #1: <unknown function> + 0x5b3a23e (0x7f002078223e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7f002077cc87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7f002077cf82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7f002077dfd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f0020732371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f0020732371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f0020732371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f0020732371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7effe7f3f189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default4]:[rank12]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7effe7f46610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default4]:[rank12]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7effe7f65978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default4]:[rank12]: frame #12: <unknown function> + 0x5adc309 (0x7f0020724309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #13: <unknown function> + 0x5ae6f10 (0x7f002072ef10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #14: <unknown function> + 0x5ae6fa5 (0x7f002072efa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #15: <unknown function> + 0x5124446 (0x7f001fd6c446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #16: <unknown function> + 0x1acf4b8 (0x7f001c7174b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #17: <unknown function> + 0x5aee004 (0x7f0020736004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #18: <unknown function> + 0x5af36b5 (0x7f002073b6b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #19: <unknown function> + 0xd2631e (0x7f003332531e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default4]:[rank12]: frame #20: <unknown function> + 0x47def4 (0x7f0032a7cef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default4]:[rank12]: frame #21: <unknown function> + 0x1445a6 (0x5572698325a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #22: _PyObject_MakeTpCall + 0x26b (0x55726982ba6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #23: <unknown function> + 0x150866 (0x55726983e866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x557269827142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #25: _PyFunction_Vectorcall + 0x6c (0x557269832a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #26: PyObject_Call + 0xbc (0x55726983ef1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x5572698252b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #28: _PyFunction_Vectorcall + 0x6c (0x557269832a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x5572698238fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #30: <unknown function> + 0x150582 (0x55726983e582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x5572698238fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #32: <unknown function> + 0x150582 (0x55726983e582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x5572698238fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #34: <unknown function> + 0x150582 (0x55726983e582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x5572698238fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x55726982af50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #37: _PyObject_Call_Prepend + 0x69 (0x55726983cc39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #38: <unknown function> + 0x211239 (0x5572698ff239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #39: _PyObject_MakeTpCall + 0x26b (0x55726982ba6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x5572698273e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #41: _PyFunction_Vectorcall + 0x6c (0x557269832a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x557269822c5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #43: _PyFunction_Vectorcall + 0x6c (0x557269832a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x5572698238fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #45: <unknown function> + 0x150582 (0x55726983e582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #46: PyObject_Call + 0xbc (0x55726983ef1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x5572698252b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #48: <unknown function> + 0x150582 (0x55726983e582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #49: PyObject_Call + 0xbc (0x55726983ef1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x5572698252b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #51: _PyFunction_Vectorcall + 0x6c (0x557269832a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x55726982b007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #53: _PyObject_Call_Prepend + 0x69 (0x55726983cc39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #54: <unknown function> + 0x211239 (0x5572698ff239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #55: PyObject_Call + 0x207 (0x55726983f067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x5572698252b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #57: <unknown function> + 0x150582 (0x55726983e582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x5572698238fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #59: <unknown function> + 0x150582 (0x55726983e582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #60: PyObject_Call + 0xbc (0x55726983ef1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x5572698252b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #62: <unknown function> + 0x150582 (0x55726983e582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #63: PyObject_Call + 0xbc (0x55726983ef1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: . This may indicate a possible application crash on rank 0 or a network set up issue.
E0702 18:23:21.277000 139628476266304 torch/distributed/elastic/multiprocessing/api.py:826] failed (exitcode: 1) local_rank: 0 (pid: 579870) of binary: /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10
Traceback (most recent call last):
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/torchrun", line 8, in <module>
sys.exit(main())
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 347, in wrapper
return f(*args, **kwargs)
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/run.py", line 879, in main
run(args)
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/run.py", line 870, in run
elastic_launch(
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 132, in __call__
return launch_agent(self._config, self._entrypoint, list(args))
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 263, in launch_agent
raise ChildFailedError(
torch.distributed.elastic.multiprocessing.errors.ChildFailedError:
============================================================
/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py FAILED
------------------------------------------------------------
Failures:
[1]:
time : 2024-07-02_18:23:21
host : ip-26-0-163-147.ec2.internal
rank : 1 (local_rank: 1)
exitcode : 1 (pid: 579871)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[2]:
time : 2024-07-02_18:23:21
host : ip-26-0-163-147.ec2.internal
rank : 2 (local_rank: 2)
exitcode : 1 (pid: 579872)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[3]:
time : 2024-07-02_18:23:21
host : ip-26-0-163-147.ec2.internal
rank : 3 (local_rank: 3)
exitcode : 1 (pid: 579873)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[4]:
time : 2024-07-02_18:23:21
host : ip-26-0-163-147.ec2.internal
rank : 4 (local_rank: 4)
exitcode : 1 (pid: 579874)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[5]:
time : 2024-07-02_18:23:21
host : ip-26-0-163-147.ec2.internal
rank : 5 (local_rank: 5)
exitcode : 1 (pid: 579875)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[6]:
time : 2024-07-02_18:23:21
host : ip-26-0-163-147.ec2.internal
rank : 6 (local_rank: 6)
exitcode : 1 (pid: 579876)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[7]:
time : 2024-07-02_18:23:21
host : ip-26-0-163-147.ec2.internal
rank : 7 (local_rank: 7)
exitcode : 1 (pid: 579877)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
------------------------------------------------------------
Root Cause (first observed failure):
[0]:
time : 2024-07-02_18:23:21
host : ip-26-0-163-147.ec2.internal
rank : 0 (local_rank: 0)
exitcode : 1 (pid: 579870)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
============================================================
srun: error: ip-26-0-163-147: task 0: Exited with exit code 1
W0702 18:23:25.138000 140239191070464 torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1252] The node 'ip-26-0-163-226.ec2.internal_2940821_0' has failed to send a keep-alive heartbeat to the rendezvous 'none' due to an error of type RendezvousConnectionError.
E0702 18:23:26.283000 140244857890624 torch/distributed/elastic/multiprocessing/api.py:826] failed (exitcode: 1) local_rank: 0 (pid: 2940889) of binary: /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10
W0702 18:23:26.290000 140244857890624 torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1203] The node 'ip-26-0-163-226.ec2.internal_2940821_0' has failed to shutdown the rendezvous 'none' due to an error of type RendezvousConnectionError.
W0702 18:23:26.316000 140244857890624 torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1203] The node 'ip-26-0-163-226.ec2.internal_2940821_0' has failed to shutdown the rendezvous 'none' due to an error of type RendezvousConnectionError.
W0702 18:23:26.354000 140244857890624 torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1203] The node 'ip-26-0-163-226.ec2.internal_2940821_0' has failed to shutdown the rendezvous 'none' due to an error of type RendezvousConnectionError.
Traceback (most recent call last):
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/torchrun", line 8, in <module>
sys.exit(main())
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 347, in wrapper
return f(*args, **kwargs)
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/run.py", line 879, in main
run(args)
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/run.py", line 870, in run
elastic_launch(
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 132, in __call__
return launch_agent(self._config, self._entrypoint, list(args))
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 263, in launch_agent
raise ChildFailedError(
torch.distributed.elastic.multiprocessing.errors.ChildFailedError:
============================================================
/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py FAILED
------------------------------------------------------------
Failures:
[1]:
time : 2024-07-02_18:23:26
host : ip-26-0-163-226.ec2.internal
rank : 9 (local_rank: 1)
exitcode : 1 (pid: 2940890)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[2]:
time : 2024-07-02_18:23:26
host : ip-26-0-163-226.ec2.internal
rank : 10 (local_rank: 2)
exitcode : 1 (pid: 2940891)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[3]:
time : 2024-07-02_18:23:26
host : ip-26-0-163-226.ec2.internal
rank : 11 (local_rank: 3)
exitcode : 1 (pid: 2940892)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[4]:
time : 2024-07-02_18:23:26
host : ip-26-0-163-226.ec2.internal
rank : 12 (local_rank: 4)
exitcode : 1 (pid: 2940893)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[5]:
time : 2024-07-02_18:23:26
host : ip-26-0-163-226.ec2.internal
rank : 13 (local_rank: 5)
exitcode : 1 (pid: 2940894)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[6]:
time : 2024-07-02_18:23:26
host : ip-26-0-163-226.ec2.internal
rank : 14 (local_rank: 6)
exitcode : 1 (pid: 2940895)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[7]:
time : 2024-07-02_18:23:26
host : ip-26-0-163-226.ec2.internal
rank : 15 (local_rank: 7)
exitcode : 1 (pid: 2940896)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
------------------------------------------------------------
Root Cause (first observed failure):
[0]:
time : 2024-07-02_18:23:26
host : ip-26-0-163-226.ec2.internal
rank : 8 (local_rank: 0)
exitcode : 1 (pid: 2940889)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
============================================================
srun: error: ip-26-0-163-226: task 1: Exited with exit code 1
Consider using `hf_transfer` for faster uploads. This solution comes with some limitations. See https://huggingface.co/docs/huggingface_hub/hf_transfer for more details.