3outeille's picture
3outeille HF staff
Upload llama-1B/16_GPUS/dp-1_tp-2_pp-8_mbz-32
22d6c3e verified
raw
history blame
147 kB
========================
START TIME: Tue Jul 2 19:46:03 UTC 2024
python3 version = Python 3.10.14
========================
The token has not been saved to the git credentials helper. Pass `add_to_git_credential=True` in this function directly or `--add-to-git-credential` if using via `huggingface-cli` if you want to set the git credential as well.
Token is valid (permission: write).
Your token has been saved to /admin/home/ferdinand_mom/.cache/huggingface/token
Login successful
Already on 'bench_cluster'
M examples/config_tiny_llama.py
M examples/config_tiny_llama.yaml
M examples/train_tiny_llama.sh
M src/nanotron/models/llama.py
M src/nanotron/trainer.py
Your branch is up to date with 'origin/bench_cluster'.
Job status: RUNNING
W0702 19:46:06.477000 139631175944000 torch/distributed/run.py:757]
W0702 19:46:06.477000 139631175944000 torch/distributed/run.py:757] *****************************************
W0702 19:46:06.477000 139631175944000 torch/distributed/run.py:757] Setting OMP_NUM_THREADS environment variable for each process to be 1 in default, to avoid your system being overloaded, please further tune the variable for optimal performance in your application as needed.
W0702 19:46:06.477000 139631175944000 torch/distributed/run.py:757] *****************************************
W0702 19:46:07.102000 139998354859840 torch/distributed/run.py:757]
W0702 19:46:07.102000 139998354859840 torch/distributed/run.py:757] *****************************************
W0702 19:46:07.102000 139998354859840 torch/distributed/run.py:757] Setting OMP_NUM_THREADS environment variable for each process to be 1 in default, to avoid your system being overloaded, please further tune the variable for optimal performance in your application as needed.
W0702 19:46:07.102000 139998354859840 torch/distributed/run.py:757] *****************************************
[default0]:07/02/2024 19:46:26 [WARNING|DP=0|PP=0|TP=0|ip-26-0-163-147]: [Vocab Size Padding] Padded vocab (size: 50257) with 1 dummy tokens (new size: 50258)
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Config:
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Config(general=GeneralArgs(project='bench_cluster',
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: run='%date_%jobid',
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: seed=42,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: step=None,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: consumed_train_samples=None,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: benchmark_csv_path=None,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: ignore_sanity_checks=True),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: parallelism=ParallelismArgs(dp=1,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: pp=8,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tp=2,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: pp_engine=<nanotron.parallel.pipeline_parallel.engine.OneForwardOneBackwardPipelineEngine object at 0x7f42fb6b8910>,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tp_mode=<TensorParallelLinearMode.REDUCE_SCATTER: 2>,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tp_linear_async_communication=False,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: expert_parallel_size=1),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: model=ModelArgs(model_config=LlamaConfig(bos_token_id=1,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: eos_token_id=2,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: hidden_act='silu',
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: hidden_size=2048,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: initializer_range=0.02,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: intermediate_size=4096,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: is_llama_config=True,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: max_position_embeddings=4096,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_attention_heads=32,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_hidden_layers=24,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_key_value_heads=32,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: pad_token_id=None,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: pretraining_tp=1,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: rms_norm_eps=1e-05,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: rope_scaling=None,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: rope_theta=10000.0,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tie_word_embeddings=True,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: use_cache=True,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: vocab_size=50258),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: init_method=RandomInit(std=0.025),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: dtype=torch.bfloat16,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: make_vocab_size_divisible_by=1,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: ddp_bucket_cap_mb=25),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tokenizer=TokenizerArgs(tokenizer_name_or_path='openai-community/gpt2',
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tokenizer_revision=None,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tokenizer_max_length=None),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: checkpoints=CheckpointsArgs(checkpoints_path=Path('/dev/null'),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: checkpoint_interval=100000,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: save_initial_state=False,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: resume_checkpoint_path=None,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: checkpoints_path_is_shared_file_system=False),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: logging=LoggingArgs(log_level='info',
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: log_level_replica='info',
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: iteration_step_info_interval=1),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tokens=TokensArgs(sequence_length=4096,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: train_steps=20,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: micro_batch_size=32,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: batch_accumulation_per_replica=32,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: val_check_interval=-1,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: limit_val_batches=0,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: limit_test_batches=0),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: optimizer=OptimizerArgs(optimizer_factory=AdamWOptimizerArgs(adam_eps=1e-08,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: adam_beta1=0.9,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: adam_beta2=0.95,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: torch_adam_is_fused=True,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: name='adamW'),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: zero_stage=1,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: weight_decay=0.01,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: clip_grad=1.0,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: accumulate_grad_in_fp32=True,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: learning_rate_scheduler=LRSchedulerArgs(learning_rate=0.0001,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: lr_warmup_steps=1,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: lr_warmup_style='linear',
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: lr_decay_style='linear',
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: lr_decay_steps=19,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: lr_decay_starting_step=None,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: min_decay_lr=1e-05)),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: data_stages=[DatasetStageArgs(name='Training Stage',
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: start_training_step=1,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: data=DataArgs(dataset=PretrainDatasetsArgs(hf_dataset_or_datasets='roneneldan/TinyStories',
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: hf_dataset_splits='train',
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: hf_dataset_config_name=None,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: dataset_processing_num_proc_per_process=64,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: dataset_overwrite_cache=False,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: text_column_name='text'),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: seed=42,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_loading_workers=32))],
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: profiler=ProfilerArgs(profiler_export_path=Path('/fsx/ferdinandmom/ferdinand-hf/bench_cluster/results/llama-1B/16_GPUS/dp-1_tp-2_pp-8_mbz-32')),
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: lighteval=None)
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Model Config:
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: LlamaConfig(bos_token_id=1,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: eos_token_id=2,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: hidden_act='silu',
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: hidden_size=2048,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: initializer_range=0.02,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: intermediate_size=4096,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: is_llama_config=True,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: max_position_embeddings=4096,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_attention_heads=32,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_hidden_layers=24,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: num_key_value_heads=32,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: pad_token_id=None,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: pretraining_tp=1,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: rms_norm_eps=1e-05,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: rope_scaling=None,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: rope_theta=10000.0,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: tie_word_embeddings=True,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: use_cache=True,
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: vocab_size=50258)
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Building model..
[default0]:07/02/2024 19:46:26 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Setting PP block ranks...
[default0]:07/02/2024 19:46:41 [INFO|DP=0|PP=4|TP=0|ip-26-0-163-226]: Local number of parameters: 62.9M (120.02MiB)
[default0]:07/02/2024 19:46:41 [INFO|DP=0|PP=4|TP=0|ip-26-0-163-226]: [After model building] Memory usage: 123.03MiB. Peak allocated: 125.06MiB Peak reserved: 138.00MiB
[default0]:07/02/2024 19:46:41 [INFO|DP=0|PP=4|TP=0|ip-26-0-163-226]: No checkpoint path provided.
[default1]:07/02/2024 19:46:41 [INFO|DP=0|PP=4|TP=1|ip-26-0-163-226]: Local number of parameters: 62.9M (120.02MiB)
[default1]:07/02/2024 19:46:41 [INFO|DP=0|PP=4|TP=1|ip-26-0-163-226]: [After model building] Memory usage: 123.03MiB. Peak allocated: 125.06MiB Peak reserved: 138.00MiB
[default1]:07/02/2024 19:46:41 [INFO|DP=0|PP=4|TP=1|ip-26-0-163-226]: No checkpoint path provided.
[default2]:07/02/2024 19:46:41 [INFO|DP=0|PP=5|TP=0|ip-26-0-163-226]: Local number of parameters: 62.9M (120.02MiB)
[default2]:07/02/2024 19:46:41 [INFO|DP=0|PP=5|TP=0|ip-26-0-163-226]: [After model building] Memory usage: 123.03MiB. Peak allocated: 125.06MiB Peak reserved: 138.00MiB
[default2]:07/02/2024 19:46:41 [INFO|DP=0|PP=5|TP=0|ip-26-0-163-226]: No checkpoint path provided.
[default3]:07/02/2024 19:46:41 [INFO|DP=0|PP=5|TP=1|ip-26-0-163-226]: Local number of parameters: 62.9M (120.02MiB)
[default3]:07/02/2024 19:46:41 [INFO|DP=0|PP=5|TP=1|ip-26-0-163-226]: [After model building] Memory usage: 123.03MiB. Peak allocated: 125.06MiB Peak reserved: 138.00MiB
[default3]:07/02/2024 19:46:41 [INFO|DP=0|PP=5|TP=1|ip-26-0-163-226]: No checkpoint path provided.
[default5]:07/02/2024 19:46:41 [INFO|DP=0|PP=6|TP=1|ip-26-0-163-226]: Local number of parameters: 83.9M (160.03MiB)
[default5]:07/02/2024 19:46:41 [INFO|DP=0|PP=6|TP=1|ip-26-0-163-226]: [After model building] Memory usage: 164.04MiB. Peak allocated: 166.07MiB Peak reserved: 180.00MiB
[default5]:07/02/2024 19:46:41 [INFO|DP=0|PP=6|TP=1|ip-26-0-163-226]: No checkpoint path provided.
[default4]:07/02/2024 19:46:41 [INFO|DP=0|PP=6|TP=0|ip-26-0-163-226]: Local number of parameters: 83.9M (160.03MiB)
[default4]:07/02/2024 19:46:41 [INFO|DP=0|PP=6|TP=0|ip-26-0-163-226]: [After model building] Memory usage: 164.04MiB. Peak allocated: 166.07MiB Peak reserved: 180.00MiB
[default4]:07/02/2024 19:46:41 [INFO|DP=0|PP=6|TP=0|ip-26-0-163-226]: No checkpoint path provided.
[default6]:07/02/2024 19:46:41 [INFO|DP=0|PP=7|TP=0|ip-26-0-163-226]: Local number of parameters: 51.5M (98.16MiB)
[default6]:07/02/2024 19:46:41 [INFO|DP=0|PP=7|TP=0|ip-26-0-163-226]: [After model building] Memory usage: 98.17MiB. Peak allocated: 98.18MiB Peak reserved: 102.00MiB
[default6]:07/02/2024 19:46:41 [INFO|DP=0|PP=7|TP=0|ip-26-0-163-226]: No checkpoint path provided.
[default0]:07/02/2024 19:46:41 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Total number of parameters: 1.21G (2313.02MiB)
[default0]:07/02/2024 19:46:41 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Local number of parameters: 135M (258.19MiB)
[default0]:07/02/2024 19:46:41 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [After model building] Memory usage: 262.20MiB. Peak allocated: 264.23MiB Peak reserved: 280.00MiB
[default0]:07/02/2024 19:46:41 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: No checkpoint path provided.
[default0]:07/02/2024 19:46:41 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Parametrizing model parameters using StandardParametrizator
[default3]:07/02/2024 19:46:41 [INFO|DP=0|PP=1|TP=1|ip-26-0-163-147]: Local number of parameters: 62.9M (120.02MiB)
[default3]:07/02/2024 19:46:41 [INFO|DP=0|PP=1|TP=1|ip-26-0-163-147]: [After model building] Memory usage: 123.03MiB. Peak allocated: 125.06MiB Peak reserved: 138.00MiB
[default4]:07/02/2024 19:46:41 [INFO|DP=0|PP=2|TP=0|ip-26-0-163-147]: Local number of parameters: 62.9M (120.02MiB)
[default4]:07/02/2024 19:46:41 [INFO|DP=0|PP=2|TP=0|ip-26-0-163-147]: [After model building] Memory usage: 123.03MiB. Peak allocated: 125.06MiB Peak reserved: 138.00MiB
[default4]:07/02/2024 19:46:41 [INFO|DP=0|PP=2|TP=0|ip-26-0-163-147]: No checkpoint path provided.
[default3]:07/02/2024 19:46:41 [INFO|DP=0|PP=1|TP=1|ip-26-0-163-147]: No checkpoint path provided.
[default7]:07/02/2024 19:46:41 [INFO|DP=0|PP=7|TP=1|ip-26-0-163-226]: Local number of parameters: 51.5M (98.16MiB)
[default7]:07/02/2024 19:46:41 [INFO|DP=0|PP=7|TP=1|ip-26-0-163-226]: [After model building] Memory usage: 98.17MiB. Peak allocated: 98.18MiB Peak reserved: 102.00MiB
[default7]:07/02/2024 19:46:41 [INFO|DP=0|PP=7|TP=1|ip-26-0-163-226]: No checkpoint path provided.
[default2]:07/02/2024 19:46:41 [INFO|DP=0|PP=1|TP=0|ip-26-0-163-147]: Local number of parameters: 62.9M (120.02MiB)
[default2]:07/02/2024 19:46:41 [INFO|DP=0|PP=1|TP=0|ip-26-0-163-147]: [After model building] Memory usage: 123.03MiB. Peak allocated: 125.06MiB Peak reserved: 138.00MiB
[default2]:07/02/2024 19:46:41 [INFO|DP=0|PP=1|TP=0|ip-26-0-163-147]: No checkpoint path provided.
[default6]:07/02/2024 19:46:41 [INFO|DP=0|PP=3|TP=0|ip-26-0-163-147]: Local number of parameters: 83.9M (160.03MiB)
[default6]:07/02/2024 19:46:41 [INFO|DP=0|PP=3|TP=0|ip-26-0-163-147]: [After model building] Memory usage: 164.04MiB. Peak allocated: 166.07MiB Peak reserved: 180.00MiB
[default6]:07/02/2024 19:46:41 [INFO|DP=0|PP=3|TP=0|ip-26-0-163-147]: No checkpoint path provided.
[default1]:07/02/2024 19:46:41 [INFO|DP=0|PP=0|TP=1|ip-26-0-163-147]: Local number of parameters: 135M (258.19MiB)
[default1]:07/02/2024 19:46:41 [INFO|DP=0|PP=0|TP=1|ip-26-0-163-147]: [After model building] Memory usage: 262.20MiB. Peak allocated: 264.23MiB Peak reserved: 280.00MiB
[default7]:07/02/2024 19:46:41 [INFO|DP=0|PP=3|TP=1|ip-26-0-163-147]: Local number of parameters: 83.9M (160.03MiB)
[default7]:07/02/2024 19:46:41 [INFO|DP=0|PP=3|TP=1|ip-26-0-163-147]: [After model building] Memory usage: 164.04MiB. Peak allocated: 166.07MiB Peak reserved: 180.00MiB
[default7]:07/02/2024 19:46:41 [INFO|DP=0|PP=3|TP=1|ip-26-0-163-147]: No checkpoint path provided.
[default1]:07/02/2024 19:46:41 [INFO|DP=0|PP=0|TP=1|ip-26-0-163-147]: No checkpoint path provided.
[default5]:07/02/2024 19:46:41 [INFO|DP=0|PP=2|TP=1|ip-26-0-163-147]: Local number of parameters: 62.9M (120.02MiB)
[default5]:07/02/2024 19:46:41 [INFO|DP=0|PP=2|TP=1|ip-26-0-163-147]: [After model building] Memory usage: 123.03MiB. Peak allocated: 125.06MiB Peak reserved: 138.00MiB
[default5]:07/02/2024 19:46:41 [INFO|DP=0|PP=2|TP=1|ip-26-0-163-147]: No checkpoint path provided.
[default0]:07/02/2024 19:46:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [Optimizer Building] Using LearningRateForSP as learning rate
[default0]:07/02/2024 19:46:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [ZeRO sharding] Size of optimizer params per rank:
[default0]:07/02/2024 19:46:43 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [ZeRO sharding] DP Rank 0 has 135M out of 135M (100.00%) params' optimizer states
[default0]:07/02/2024 19:46:44 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [Training Plan] Stage Training Stage has 19 remaining training steps and has consumed 0 samples
[default0]:07/02/2024 19:46:44 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Using `datasets` library
[default0]:07/02/2024 19:46:44 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Loading tokenizer from openai-community/gpt2 and transformers/hf_hub versions ('4.41.2', '0.23.4')
[default0]:07/02/2024 19:46:44 [WARNING|DP=0|PP=0|TP=0|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default0]:Repo card metadata block was not found. Setting CardData to empty.
[default0]:07/02/2024 19:46:46 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [Training Plan] There are 1 training stages
[default0]:07/02/2024 19:46:46 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [Stage Training Stage] start from step 1
[default0]:07/02/2024 19:46:46 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]:
[default0]:07/02/2024 19:46:46 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: [Start training] datetime: 2024-07-02 19:46:46.385273 | mbs: 32 | grad_accum: 32 | global_batch_size: 1024 | sequence_length: 4096 | train_steps: 20 | start_iteration_step: 0 | consumed_train_samples: 0
[default0]:07/02/2024 19:46:46 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Resuming training from stage Training Stage, it has trained for 0 samples and has 19 remaining train steps
[default0]:07/02/2024 19:46:46 [INFO|DP=0|PP=0|TP=0|ip-26-0-163-147]: Memory usage: 1294.97MiB. Peak allocated 1294.97MiB. Peak reserved: 1316.00MiB
[default5]:07/02/2024 19:46:46 [WARNING|DP=0|PP=6|TP=1|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default4]:07/02/2024 19:46:46 [WARNING|DP=0|PP=6|TP=0|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default5]:Repo card metadata block was not found. Setting CardData to empty.
[default4]:Repo card metadata block was not found. Setting CardData to empty.
[default2]:07/02/2024 19:46:46 [WARNING|DP=0|PP=1|TP=0|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default6]:07/02/2024 19:46:46 [WARNING|DP=0|PP=3|TP=0|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default1]:07/02/2024 19:46:46 [WARNING|DP=0|PP=0|TP=1|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default5]:07/02/2024 19:46:46 [WARNING|DP=0|PP=2|TP=1|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default7]:07/02/2024 19:46:46 [WARNING|DP=0|PP=3|TP=1|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default1]:Repo card metadata block was not found. Setting CardData to empty.
[default5]:Repo card metadata block was not found. Setting CardData to empty.
[default4]:Repo card metadata block was not found. Setting CardData to empty.
[default2]:Repo card metadata block was not found. Setting CardData to empty.
[default6]:Repo card metadata block was not found. Setting CardData to empty.
[default7]:Repo card metadata block was not found. Setting CardData to empty.
[default1]:07/02/2024 19:46:46 [WARNING|DP=0|PP=4|TP=1|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default2]:07/02/2024 19:46:46 [WARNING|DP=0|PP=5|TP=0|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default3]:Repo card metadata block was not found. Setting CardData to empty.
[default7]:Repo card metadata block was not found. Setting CardData to empty.
[default1]:Repo card metadata block was not found. Setting CardData to empty.
[default2]:Repo card metadata block was not found. Setting CardData to empty.
[default3]:07/02/2024 19:46:46 [WARNING|DP=0|PP=5|TP=1|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default7]:07/02/2024 19:46:46 [WARNING|DP=0|PP=7|TP=1|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default4]:07/02/2024 19:46:46 [WARNING|DP=0|PP=2|TP=0|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default3]:07/02/2024 19:46:46 [WARNING|DP=0|PP=1|TP=1|ip-26-0-163-147]: Repo card metadata block was not found. Setting CardData to empty.
[default3]:Repo card metadata block was not found. Setting CardData to empty.
[default0]:Repo card metadata block was not found. Setting CardData to empty.
[default0]:07/02/2024 19:46:46 [WARNING|DP=0|PP=4|TP=0|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default6]:Repo card metadata block was not found. Setting CardData to empty.
[default6]:07/02/2024 19:46:46 [WARNING|DP=0|PP=7|TP=0|ip-26-0-163-226]: Repo card metadata block was not found. Setting CardData to empty.
[default1]:[rank1]: Traceback (most recent call last):
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default1]:[rank1]: trainer.train(dataloader)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default1]:[rank1]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default1]:[rank1]: outputs = self.pipeline_engine.train_batch_iter(
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default1]:[rank1]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default1]:[rank1]: output = model(**micro_batch)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank1]: return self._call_impl(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank1]: return forward_call(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default1]:[rank1]: sharded_logits = self.model(
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank1]: return self._call_impl(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank1]: return forward_call(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default1]:[rank1]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default1]:[rank1]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank1]: return self._call_impl(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank1]: return forward_call(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 151, in forward
[default1]:[rank1]: output = self.pp_block(**new_kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank1]: return self._call_impl(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank1]: return forward_call(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 637, in forward
[default1]:[rank1]: hidden_states = self.mlp(hidden_states=hidden_states)["hidden_states"]
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank1]: return self._call_impl(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank1]: return forward_call(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 172, in forward
[default1]:[rank1]: hidden_states = self.down_proj(self.split_silu_mul(merged_states))
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default1]:[rank1]: return self._call_impl(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default1]:[rank1]: return forward_call(*args, **kwargs)
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/nn.py", line 159, in forward
[default1]:[rank1]: return row_linear(
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/functional.py", line 474, in row_linear
[default1]:[rank1]: out = F.linear(input, weight, bias)
[default1]:[rank1]: torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 512.00 MiB. GPU  has a total capacity of 79.33 GiB of which 433.94 MiB is free. Including non-PyTorch memory, this process has 78.89 GiB memory in use. Of the allocated memory 70.16 GiB is allocated by PyTorch, and 682.76 MiB is reserved by PyTorch but unallocated. If reserved but unallocated memory is large try setting PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True to avoid fragmentation. See documentation for Memory Management (https://pytorch.org/docs/stable/notes/cuda.html#environment-variables)
[default0]:[rank0]: Traceback (most recent call last):
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default0]:[rank0]: trainer.train(dataloader)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default0]:[rank0]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default0]:[rank0]: outputs = self.pipeline_engine.train_batch_iter(
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default0]:[rank0]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default0]:[rank0]: output = model(**micro_batch)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank0]: return self._call_impl(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank0]: return forward_call(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default0]:[rank0]: sharded_logits = self.model(
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank0]: return self._call_impl(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank0]: return forward_call(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default0]:[rank0]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default0]:[rank0]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank0]: return self._call_impl(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank0]: return forward_call(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 151, in forward
[default0]:[rank0]: output = self.pp_block(**new_kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank0]: return self._call_impl(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank0]: return forward_call(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 631, in forward
[default0]:[rank0]: output = self.attn(hidden_states=hidden_states, sequence_mask=sequence_mask)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank0]: return self._call_impl(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank0]: return forward_call(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 360, in forward
[default0]:[rank0]: qkv_states = self.qkv_proj(
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default0]:[rank0]: return self._call_impl(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default0]:[rank0]: return forward_call(*args, **kwargs)
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/nn.py", line 87, in forward
[default0]:[rank0]: return column_linear(
[default0]:[rank0]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/tensor_parallel/functional.py", line 359, in column_linear
[default0]:[rank0]: return F.linear(input, weight, bias)
[default0]:[rank0]: torch.cuda.OutOfMemoryError: CUDA out of memory. Tried to allocate 768.00 MiB. GPU
W0702 19:47:02.862000 139631175944000 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 674634 closing signal SIGTERM
W0702 19:47:02.863000 139631175944000 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 674636 closing signal SIGTERM
W0702 19:47:02.864000 139631175944000 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 674637 closing signal SIGTERM
W0702 19:47:02.864000 139631175944000 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 674638 closing signal SIGTERM
W0702 19:47:02.865000 139631175944000 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 674639 closing signal SIGTERM
W0702 19:47:02.865000 139631175944000 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 674640 closing signal SIGTERM
W0702 19:47:02.866000 139631175944000 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 674641 closing signal SIGTERM
[default3]:[rank11]: Traceback (most recent call last):
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default3]:[rank11]: trainer.train(dataloader)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default3]:[rank11]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default3]:[rank11]: outputs = self.pipeline_engine.train_batch_iter(
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default3]:[rank11]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default3]:[rank11]: output = model(**micro_batch)
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default3]:[rank11]: return self._call_impl(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default3]:[rank11]: return forward_call(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default3]:[rank11]: sharded_logits = self.model(
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default3]:[rank11]: return self._call_impl(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default3]:[rank11]: return forward_call(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default3]:[rank11]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default3]:[rank11]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default3]:[rank11]: return self._call_impl(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default3]:[rank11]: return forward_call(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default3]:[rank11]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default3]:[rank11]: pipeline_state.run_communication()
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default3]:[rank11]: recv_activation_tensor = recv_activation()
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default3]:[rank11]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default3]:[rank11]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default3]:[rank11]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default3]:[rank11]: dist.recv(
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default3]:[rank11]: return func(*args, **kwargs)
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default3]:[rank11]: pg.recv([tensor], group_src_rank, tag).wait()
[default3]:[rank11]: torch.distributed.DistBackendError: [5] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '4:5', but store->get('4:5') got error: Connection reset by peer
[default3]:[rank11]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default3]:[rank11]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7f11f2176897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default3]:[rank11]: frame #1: <unknown function> + 0x5b3a23e (0x7f122bc9323e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7f122bc8dc87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7f122bc8df82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7f122bc8efd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f122bc43371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: Traceback (most recent call last):
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default7]:[rank15]: Traceback (most recent call last):
[default3]:[rank11]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f122bc43371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f122bc43371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: Traceback (most recent call last):
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default3]:[rank11]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f122bc43371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7f11f3450189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default4]:[rank12]: Traceback (most recent call last):
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default3]:[rank11]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7f11f3457610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default3]:[rank11]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7f11f3476978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default5]:[rank13]: trainer.train(dataloader)
[default2]:[rank10]: Traceback (most recent call last):
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default3]:[rank11]: frame #12: <unknown function> + 0x5adc309 (0x7f122bc35309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #13: <unknown function> + 0x5ae6f10 (0x7f122bc3ff10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module>
[default7]:[rank15]: trainer.train(dataloader)
[default6]:[rank14]: trainer.train(dataloader)
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default3]:[rank11]: frame #14: <unknown function> + 0x5ae6fa5 (0x7f122bc3ffa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default4]:[rank12]: trainer.train(dataloader)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default2]:[rank10]: trainer.train(dataloader)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default2]:[rank10]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default2]:[rank10]: outputs = self.pipeline_engine.train_batch_iter(
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default2]:[rank10]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default2]:[rank10]: output = model(**micro_batch)
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank10]: return self._call_impl(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default2]:[rank10]: return forward_call(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default2]:[rank10]: sharded_logits = self.model(
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank10]: return self._call_impl(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default2]:[rank10]: return forward_call(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default2]:[rank10]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default6]:[rank14]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default3]:[rank11]: frame #15: <unknown function> + 0x5124446 (0x7f122b27d446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train
[default4]:[rank12]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default4]:[rank12]: outputs = self.pipeline_engine.train_batch_iter(
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default4]:[rank12]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default4]:[rank12]: output = model(**micro_batch)
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default4]:[rank12]: return self._call_impl(*args, **kwargs)
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank12]: return forward_call(*args, **kwargs)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default4]:[rank12]: sharded_logits = self.model(
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default4]:[rank12]: return self._call_impl(*args, **kwargs)
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank12]: return forward_call(*args, **kwargs)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default4]:[rank12]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default4]:[rank12]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default4]:[rank12]: return self._call_impl(*args, **kwargs)
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank12]: return forward_call(*args, **kwargs)
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default4]:[rank12]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default4]:[rank12]: pipeline_state.run_communication()
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default4]:[rank12]: recv_activation_tensor = recv_activation()
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default4]:[rank12]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default7]:[rank15]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default4]:[rank12]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default5]:[rank13]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step
[default6]:[rank14]: outputs = self.pipeline_engine.train_batch_iter(
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default5]:[rank13]: outputs = self.pipeline_engine.train_batch_iter(
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter
[default5]:[rank13]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default7]:[rank15]: outputs = self.pipeline_engine.train_batch_iter(
[default4]:[rank12]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default2]:[rank10]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default3]:[rank11]: frame #16: <unknown function> + 0x1acf4b8 (0x7f1227c284b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #17: <unknown function> + 0x5aee004 (0x7f122bc47004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #18: <unknown function> + 0x5af36b5 (0x7f122bc4c6b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default3]:[rank11]: frame #19: <unknown function> + 0xd2631e (0x7f123e83631e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default3]:[rank11]: frame #20: <unknown function> + 0x47def4 (0x7f123df8def4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default3]:[rank11]: frame #21: <unknown function> + 0x1445a6 (0x55563b2035a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #22: _PyObject_MakeTpCall + 0x26b (0x55563b1fca6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #23: <unknown function> + 0x150866 (0x55563b20f866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x55563b1f8142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #25: _PyFunction_Vectorcall + 0x6c (0x55563b203a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #26: PyObject_Call + 0xbc (0x55563b20ff1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x55563b1f62b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #28: _PyFunction_Vectorcall + 0x6c (0x55563b203a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x55563b1f48fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #30: <unknown function> + 0x150582 (0x55563b20f582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x55563b1f48fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #32: <unknown function> + 0x150582 (0x55563b20f582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x55563b1f48fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #34: <unknown function> + 0x150582 (0x55563b20f582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x55563b1f48fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x55563b1fbf50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #37: _PyObject_Call_Prepend + 0x69 (0x55563b20dc39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #38: <unknown function> + 0x211239 (0x55563b2d0239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #39: _PyObject_MakeTpCall + 0x26b (0x55563b1fca6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x55563b1f83e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #41: _PyFunction_Vectorcall + 0x6c (0x55563b203a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x55563b1f3c5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #43: _PyFunction_Vectorcall + 0x6c (0x55563b203a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x55563b1f48fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #45: <unknown function> + 0x150582 (0x55563b20f582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #46: PyObject_Call + 0xbc (0x55563b20ff1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x55563b1f62b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #48: <unknown function> + 0x150582 (0x55563b20f582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #49: PyObject_Call + 0xbc (0x55563b20ff1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x55563b1f62b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #51: _PyFunction_Vectorcall + 0x6c (0x55563b203a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x55563b1fc007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #53: _PyObject_Call_Prepend + 0x69 (0x55563b20dc39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #54: <unknown function> + 0x211239 (0x55563b2d0239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #55: PyObject_Call + 0x207 (0x55563b210067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x55563b1f62b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #57: <unknown function> + 0x150582 (0x55563b20f582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x55563b1f48fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #59: <unknown function> + 0x150582 (0x55563b20f582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #60: PyObject_Call + 0xbc (0x55563b20ff1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x55563b1f62b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #62: <unknown function> + 0x150582 (0x55563b20f582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: frame #63: PyObject_Call + 0xbc (0x55563b20ff1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default3]:[rank11]: . This may indicate a possible application crash on rank 0 or a network set up issue.
[default6]:[rank14]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default6]:[rank14]: output = model(**micro_batch)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter
[default7]:[rank15]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward
[default5]:[rank13]: output = model(**micro_batch)
[default2]:[rank10]: return self._call_impl(*args, **kwargs)
[default7]:[rank15]: output = model(**micro_batch)
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default6]:[rank14]: return self._call_impl(*args, **kwargs)
[default7]:[rank15]: return self._call_impl(*args, **kwargs)
[default5]:[rank13]: return self._call_impl(*args, **kwargs)
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank12]: dist.recv(
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default2]:[rank10]: return forward_call(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default4]:[rank12]: return func(*args, **kwargs)
[default5]:[rank13]: return forward_call(*args, **kwargs)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default6]:[rank14]: return forward_call(*args, **kwargs)
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default5]:[rank13]: sharded_logits = self.model(
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default5]:[rank13]: return self._call_impl(*args, **kwargs)
[default7]:[rank15]: return forward_call(*args, **kwargs)
[default2]:[rank10]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default2]:[rank10]: pipeline_state.run_communication()
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default6]:[rank14]: sharded_logits = self.model(
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default2]:[rank10]: recv_activation_tensor = recv_activation()
[default4]:[rank12]: pg.recv([tensor], group_src_rank, tag).wait()
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default7]:[rank15]: sharded_logits = self.model(
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default6]:[rank14]: return self._call_impl(*args, **kwargs)
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default4]:[rank12]: torch.distributed.DistBackendError: [6] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '5:6', but store->get('5:6') got error: Connection reset by peer
[default7]:[rank15]: return self._call_impl(*args, **kwargs)
[default2]:[rank10]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default4]:[rank12]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default4]:[rank12]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7f01ff48b897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default5]:[rank13]: return forward_call(*args, **kwargs)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default7]:[rank15]: return forward_call(*args, **kwargs)
[default5]:[rank13]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default2]:[rank10]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states
[default6]:[rank14]: return forward_call(*args, **kwargs)
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default5]:[rank13]: hidden_encoder_states = encoder_block(**hidden_encoder_states)
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward
[default6]:[rank14]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 782, in forward_with_hidden_states
[default7]:[rank15]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0]
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 782, in forward_with_hidden_states
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default4]:[rank12]: frame #1: <unknown function> + 0x5b3a23e (0x7f0238fa823e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: hidden_states = self.final_layer_norm(input=hidden_encoder_states["hidden_states"])["hidden_states"]
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default6]:[rank14]: hidden_states = self.final_layer_norm(input=hidden_encoder_states["hidden_states"])["hidden_states"]
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl
[default2]:[rank10]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default4]:[rank12]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7f0238fa2c87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: return self._call_impl(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default4]:[rank12]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7f0238fa2f82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7f0238fa3fd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f0238f58371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: dist.recv(
[default7]:[rank15]: return self._call_impl(*args, **kwargs)
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default7]:[rank15]: return forward_call(*args, **kwargs)
[default4]:[rank12]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f0238f58371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f0238f58371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default6]:[rank14]: return forward_call(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default2]:[rank10]: return func(*args, **kwargs)
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default4]:[rank12]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f0238f58371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: return self._call_impl(*args, **kwargs)
[default2]:[rank10]: pg.recv([tensor], group_src_rank, tag).wait()
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default7]:[rank15]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl
[default2]:[rank10]: torch.distributed.DistBackendError: [5] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '4:5', but store->get('4:5') got error: Connection reset by peer
[default2]:[rank10]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default5]:[rank13]: return forward_call(*args, **kwargs)
[default6]:[rank14]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default7]:[rank15]: pipeline_state.run_communication()
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default4]:[rank12]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7f0200765189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default2]:[rank10]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7fed32b1f897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default2]:[rank10]: frame #1: <unknown function> + 0x5b3a23e (0x7fed6c63c23e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7f020076c610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward
[default6]:[rank14]: pipeline_state.run_communication()
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default2]:[rank10]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7fed6c636c87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7f020078b978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default4]:[rank12]: frame #12: <unknown function> + 0x5adc309 (0x7f0238f4a309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #13: <unknown function> + 0x5ae6f10 (0x7f0238f54f10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: recv_activation_tensor = recv_activation()
[default5]:[rank13]: new_kwargs[name] = recv_from_pipeline_state_buffer(
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer
[default5]:[rank13]: pipeline_state.run_communication()
[default6]:[rank14]: recv_activation_tensor = recv_activation()
[default2]:[rank10]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7fed6c636f82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7fed6c637fd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default2]:[rank10]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fed6c5ec371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fed6c5ec371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fed6c5ec371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: recv_activation_tensor = recv_activation()
[default6]:[rank14]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default2]:[rank10]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fed6c5ec371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default4]:[rank12]: frame #14: <unknown function> + 0x5ae6fa5 (0x7f0238f54fa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #15: <unknown function> + 0x5124446 (0x7f0238592446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #16: <unknown function> + 0x1acf4b8 (0x7f0234f3d4b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7fed33df9189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default2]:[rank10]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7fed33e00610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__
[default6]:[rank14]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default2]:[rank10]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7fed33e1f978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default5]:[rank13]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default6]:[rank14]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default2]:[rank10]: frame #12: <unknown function> + 0x5adc309 (0x7fed6c5de309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #13: <unknown function> + 0x5ae6f10 (0x7fed6c5e8f10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #17: <unknown function> + 0x5aee004 (0x7f0238f5c004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #18: <unknown function> + 0x5af36b5 (0x7f0238f616b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #19: <unknown function> + 0xd2631e (0x7f024bb4b31e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default4]:[rank12]: frame #20: <unknown function> + 0x47def4 (0x7f024b2a2ef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default7]:[rank15]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0]
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors
[default2]:[rank10]: frame #14: <unknown function> + 0x5ae6fa5 (0x7fed6c5e8fa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #15: <unknown function> + 0x5124446 (0x7fed6bc26446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default6]:[rank14]: dist.recv(
[default4]:[rank12]: frame #21: <unknown function> + 0x1445a6 (0x5568abee75a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default7]:[rank15]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default6]:[rank14]: return func(*args, **kwargs)
[default2]:[rank10]: frame #16: <unknown function> + 0x1acf4b8 (0x7fed685d14b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default2]:[rank10]: frame #17: <unknown function> + 0x5aee004 (0x7fed6c5f0004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #18: <unknown function> + 0x5af36b5 (0x7fed6c5f56b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: dist.recv(
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors
[default6]:[rank14]: pg.recv([tensor], group_src_rank, tag).wait()
[default4]:[rank12]: frame #22: _PyObject_MakeTpCall + 0x26b (0x5568abee0a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #23: <unknown function> + 0x150866 (0x5568abef3866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #19: <unknown function> + 0xd2631e (0x7fed7f1df31e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default6]:[rank14]: torch.distributed.DistBackendError: [7] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '6:7', but store->get('6:7') got error: Connection reset by peer
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default4]:[rank12]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x5568abedc142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #25: _PyFunction_Vectorcall + 0x6c (0x5568abee7a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default5]:[rank13]: meta = self._recv_meta(from_rank=from_rank, tag=tag)
[default6]:[rank14]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7f14318ea897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default7]:[rank15]: return func(*args, **kwargs)
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default6]:[rank14]: frame #1: <unknown function> + 0x5b3a23e (0x7f146b40723e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7f146b401c87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #20: <unknown function> + 0x47def4 (0x7fed7e936ef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default7]:[rank15]: pg.recv([tensor], group_src_rank, tag).wait()
[default4]:[rank12]: frame #26: PyObject_Call + 0xbc (0x5568abef3f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #21: <unknown function> + 0x1445a6 (0x561a62e985a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7f146b401f82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7f146b402fd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta
[default5]:[rank13]: dist.recv(
[default7]:[rank15]: torch.distributed.DistBackendError: [7] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '6:7', but store->get('6:7') got error: Connection reset by peer
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper
[default2]:[rank10]: frame #22: _PyObject_MakeTpCall + 0x26b (0x561a62e91a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default5]:[rank13]: return func(*args, **kwargs)
[default4]:[rank12]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x5568abeda2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f146b3b7371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv
[default5]:[rank13]: pg.recv([tensor], group_src_rank, tag).wait()
[default7]:[rank15]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7fe065370897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default4]:[rank12]: frame #28: _PyFunction_Vectorcall + 0x6c (0x5568abee7a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f146b3b7371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x5568abed88fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #1: <unknown function> + 0x5b3a23e (0x7fe09ee8d23e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #30: <unknown function> + 0x150582 (0x5568abef3582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f146b3b7371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f146b3b7371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7fe09ee87c87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7fe09ee87f82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x5568abed88fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #23: <unknown function> + 0x150866 (0x561a62ea4866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x561a62e8d142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7fe09ee88fd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fe09ee3d371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7f1432bc4189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default6]:[rank14]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7f1432bcb610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default2]:[rank10]: frame #25: _PyFunction_Vectorcall + 0x6c (0x561a62e98a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: torch.distributed.DistBackendError: [6] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '5:6', but store->get('5:6') got error: Connection reset by peer
[default5]:[rank13]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first):
[default7]:[rank15]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fe09ee3d371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fe09ee3d371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #26: PyObject_Call + 0xbc (0x561a62ea4f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x561a62e8b2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #32: <unknown function> + 0x150582 (0x5568abef3582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fe09ee3d371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7f3d58106897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so)
[default5]:[rank13]: frame #1: <unknown function> + 0x5b3a23e (0x7f3d91c2323e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7f1432bea978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default4]:[rank12]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x5568abed88fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7fe06664a189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default7]:[rank15]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7fe066651610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default5]:[rank13]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7f3d91c1dc87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #34: <unknown function> + 0x150582 (0x5568abef3582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #28: _PyFunction_Vectorcall + 0x6c (0x561a62e98a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x561a62e898fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #30: <unknown function> + 0x150582 (0x561a62ea4582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x5568abed88fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x5568abedff50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7fe066670978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default4]:[rank12]: frame #37: _PyObject_Call_Prepend + 0x69 (0x5568abef1c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #38: <unknown function> + 0x211239 (0x5568abfb4239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #12: <unknown function> + 0x5adc309 (0x7f146b3a9309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #13: <unknown function> + 0x5ae6f10 (0x7f146b3b3f10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #12: <unknown function> + 0x5adc309 (0x7fe09ee2f309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #13: <unknown function> + 0x5ae6f10 (0x7fe09ee39f10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7f3d91c1df82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7f3d91c1efd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f3d91bd3371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #14: <unknown function> + 0x5ae6fa5 (0x7f146b3b3fa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #15: <unknown function> + 0x5124446 (0x7f146a9f1446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f3d91bd3371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f3d91bd3371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #39: _PyObject_MakeTpCall + 0x26b (0x5568abee0a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x561a62e898fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #16: <unknown function> + 0x1acf4b8 (0x7f146739c4b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f3d91bd3371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #32: <unknown function> + 0x150582 (0x561a62ea4582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x5568abedc3e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #41: _PyFunction_Vectorcall + 0x6c (0x5568abee7a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #14: <unknown function> + 0x5ae6fa5 (0x7fe09ee39fa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #15: <unknown function> + 0x5124446 (0x7fe09e477446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7f3d593e0189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default6]:[rank14]: frame #17: <unknown function> + 0x5aee004 (0x7f146b3bb004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #18: <unknown function> + 0x5af36b5 (0x7f146b3c06b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7f3d593e7610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default5]:[rank13]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7f3d59406978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so)
[default2]:[rank10]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x561a62e898fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x5568abed7c5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #43: _PyFunction_Vectorcall + 0x6c (0x5568abee7a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #19: <unknown function> + 0xd2631e (0x7f147dfaa31e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default6]:[rank14]: frame #20: <unknown function> + 0x47def4 (0x7f147d701ef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default7]:[rank15]: frame #16: <unknown function> + 0x1acf4b8 (0x7fe09ae224b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #17: <unknown function> + 0x5aee004 (0x7fe09ee41004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #34: <unknown function> + 0x150582 (0x561a62ea4582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #12: <unknown function> + 0x5adc309 (0x7f3d91bc5309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #13: <unknown function> + 0x5ae6f10 (0x7f3d91bcff10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default6]:[rank14]: frame #21: <unknown function> + 0x1445a6 (0x55be896875a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #22: _PyObject_MakeTpCall + 0x26b (0x55be89680a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x5568abed88fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #45: <unknown function> + 0x150582 (0x5568abef3582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #18: <unknown function> + 0x5af36b5 (0x7fe09ee466b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default7]:[rank15]: frame #19: <unknown function> + 0xd2631e (0x7fe0b1a3031e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default6]:[rank14]: frame #23: <unknown function> + 0x150866 (0x55be89693866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #20: <unknown function> + 0x47def4 (0x7fe0b1187ef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default7]:[rank15]: frame #21: <unknown function> + 0x1445a6 (0x5649a87b65a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #22: _PyObject_MakeTpCall + 0x26b (0x5649a87afa6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #23: <unknown function> + 0x150866 (0x5649a87c2866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x5649a87ab142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #25: _PyFunction_Vectorcall + 0x6c (0x5649a87b6a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #26: PyObject_Call + 0xbc (0x5649a87c2f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x5649a87a92b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #28: _PyFunction_Vectorcall + 0x6c (0x5649a87b6a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x5649a87a78fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #30: <unknown function> + 0x150582 (0x5649a87c2582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x5649a87a78fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #32: <unknown function> + 0x150582 (0x5649a87c2582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x5649a87a78fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #34: <unknown function> + 0x150582 (0x5649a87c2582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x5649a87a78fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x5649a87aef50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #37: _PyObject_Call_Prepend + 0x69 (0x5649a87c0c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #38: <unknown function> + 0x211239 (0x5649a8883239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #39: _PyObject_MakeTpCall + 0x26b (0x5649a87afa6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x5649a87ab3e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #41: _PyFunction_Vectorcall + 0x6c (0x5649a87b6a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x5649a87a6c5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #43: _PyFunction_Vectorcall + 0x6c (0x5649a87b6a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x5649a87a78fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #45: <unknown function> + 0x150582 (0x5649a87c2582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #46: PyObject_Call + 0xbc (0x5649a87c2f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x5649a87a92b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #48: <unknown function> + 0x150582 (0x5649a87c2582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #49: PyObject_Call + 0xbc (0x5649a87c2f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x5649a87a92b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #51: _PyFunction_Vectorcall + 0x6c (0x5649a87b6a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x5649a87af007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #53: _PyObject_Call_Prepend + 0x69 (0x5649a87c0c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #54: <unknown function> + 0x211239 (0x5649a8883239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #55: _PyObject_MakeTpCall + 0x26b (0x5649a87afa6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #56: _PyEval_EvalFrameDefault + 0x5723 (0x5649a87abc53 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #57: <unknown function> + 0x150582 (0x5649a87c2582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x5649a87a78fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #59: <unknown function> + 0x150582 (0x5649a87c2582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #60: PyObject_Call + 0xbc (0x5649a87c2f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x5649a87a92b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #62: <unknown function> + 0x150582 (0x5649a87c2582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: frame #63: PyObject_Call + 0xbc (0x5649a87c2f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default7]:[rank15]: . This may indicate a possible application crash on rank 0 or a network set up issue.
[default6]:[rank14]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x55be8967c142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #25: _PyFunction_Vectorcall + 0x6c (0x55be89687a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #26: PyObject_Call + 0xbc (0x55be89693f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #46: PyObject_Call + 0xbc (0x5568abef3f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x561a62e898fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #14: <unknown function> + 0x5ae6fa5 (0x7f3d91bcffa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #15: <unknown function> + 0x5124446 (0x7f3d9120d446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x561a62e90f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x5568abeda2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #37: _PyObject_Call_Prepend + 0x69 (0x561a62ea2c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #38: <unknown function> + 0x211239 (0x561a62f65239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #16: <unknown function> + 0x1acf4b8 (0x7f3d8dbb84b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default5]:[rank13]: frame #17: <unknown function> + 0x5aee004 (0x7f3d91bd7004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default2]:[rank10]: frame #39: _PyObject_MakeTpCall + 0x26b (0x561a62e91a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x561a62e8d3e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #18: <unknown function> + 0x5af36b5 (0x7f3d91bdc6b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so)
[default4]:[rank12]: frame #48: <unknown function> + 0x150582 (0x5568abef3582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #41: _PyFunction_Vectorcall + 0x6c (0x561a62e98a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x561a62e88c5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #19: <unknown function> + 0xd2631e (0x7f3da47c631e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default4]:[rank12]: frame #49: PyObject_Call + 0xbc (0x5568abef3f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x5568abeda2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #43: _PyFunction_Vectorcall + 0x6c (0x561a62e98a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x561a62e898fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #45: <unknown function> + 0x150582 (0x561a62ea4582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x55be8967a2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #28: _PyFunction_Vectorcall + 0x6c (0x55be89687a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x55be896788fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #46: PyObject_Call + 0xbc (0x561a62ea4f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #51: _PyFunction_Vectorcall + 0x6c (0x5568abee7a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #30: <unknown function> + 0x150582 (0x55be89693582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x55be896788fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #32: <unknown function> + 0x150582 (0x55be89693582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x5568abee0007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #20: <unknown function> + 0x47def4 (0x7f3da3f1def4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so)
[default6]:[rank14]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x55be896788fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x561a62e8b2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #21: <unknown function> + 0x1445a6 (0x56291d36c5a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #22: _PyObject_MakeTpCall + 0x26b (0x56291d365a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #34: <unknown function> + 0x150582 (0x55be89693582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x55be896788fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #23: <unknown function> + 0x150866 (0x56291d378866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #48: <unknown function> + 0x150582 (0x561a62ea4582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #49: PyObject_Call + 0xbc (0x561a62ea4f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x55be8967ff50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #53: _PyObject_Call_Prepend + 0x69 (0x5568abef1c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x56291d361142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #25: _PyFunction_Vectorcall + 0x6c (0x56291d36ca2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #54: <unknown function> + 0x211239 (0x5568abfb4239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #37: _PyObject_Call_Prepend + 0x69 (0x55be89691c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #38: <unknown function> + 0x211239 (0x55be89754239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #55: PyObject_Call + 0x207 (0x5568abef4067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x5568abeda2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #57: <unknown function> + 0x150582 (0x5568abef3582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x561a62e8b2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #51: _PyFunction_Vectorcall + 0x6c (0x561a62e98a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #39: _PyObject_MakeTpCall + 0x26b (0x55be89680a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x5568abed88fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x561a62e91007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #53: _PyObject_Call_Prepend + 0x69 (0x561a62ea2c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #59: <unknown function> + 0x150582 (0x5568abef3582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #60: PyObject_Call + 0xbc (0x5568abef3f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x55be8967c3e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #41: _PyFunction_Vectorcall + 0x6c (0x55be89687a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x55be89677c5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x5568abeda2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #54: <unknown function> + 0x211239 (0x561a62f65239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #55: PyObject_Call + 0x207 (0x561a62ea5067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #62: <unknown function> + 0x150582 (0x5568abef3582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #26: PyObject_Call + 0xbc (0x56291d378f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x561a62e8b2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #57: <unknown function> + 0x150582 (0x561a62ea4582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x56291d35f2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #28: _PyFunction_Vectorcall + 0x6c (0x56291d36ca2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #43: _PyFunction_Vectorcall + 0x6c (0x55be89687a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: frame #63: PyObject_Call + 0xbc (0x5568abef3f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default4]:[rank12]: . This may indicate a possible application crash on rank 0 or a network set up issue.
[default2]:[rank10]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x561a62e898fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #59: <unknown function> + 0x150582 (0x561a62ea4582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x55be896788fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #60: PyObject_Call + 0xbc (0x561a62ea4f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #45: <unknown function> + 0x150582 (0x55be89693582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x561a62e8b2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #46: PyObject_Call + 0xbc (0x55be89693f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x55be8967a2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x56291d35d8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #30: <unknown function> + 0x150582 (0x56291d378582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x56291d35d8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #48: <unknown function> + 0x150582 (0x55be89693582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #62: <unknown function> + 0x150582 (0x561a62ea4582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #32: <unknown function> + 0x150582 (0x56291d378582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x56291d35d8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #49: PyObject_Call + 0xbc (0x55be89693f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x55be8967a2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #34: <unknown function> + 0x150582 (0x56291d378582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: frame #63: PyObject_Call + 0xbc (0x561a62ea4f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default2]:[rank10]: . This may indicate a possible application crash on rank 0 or a network set up issue.
[default6]:[rank14]: frame #51: _PyFunction_Vectorcall + 0x6c (0x55be89687a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x56291d35d8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x55be89680007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #53: _PyObject_Call_Prepend + 0x69 (0x55be89691c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #54: <unknown function> + 0x211239 (0x55be89754239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #55: _PyObject_MakeTpCall + 0x26b (0x55be89680a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x56291d364f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #56: _PyEval_EvalFrameDefault + 0x5723 (0x55be8967cc53 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #57: <unknown function> + 0x150582 (0x55be89693582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x55be896788fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #59: <unknown function> + 0x150582 (0x55be89693582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #60: PyObject_Call + 0xbc (0x55be89693f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x55be8967a2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #37: _PyObject_Call_Prepend + 0x69 (0x56291d376c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #62: <unknown function> + 0x150582 (0x55be89693582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: frame #63: PyObject_Call + 0xbc (0x55be89693f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default6]:[rank14]: . This may indicate a possible application crash on rank 0 or a network set up issue.
[default5]:[rank13]: frame #38: <unknown function> + 0x211239 (0x56291d439239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #39: _PyObject_MakeTpCall + 0x26b (0x56291d365a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x56291d3613e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #41: _PyFunction_Vectorcall + 0x6c (0x56291d36ca2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x56291d35cc5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #43: _PyFunction_Vectorcall + 0x6c (0x56291d36ca2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x56291d35d8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #45: <unknown function> + 0x150582 (0x56291d378582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #46: PyObject_Call + 0xbc (0x56291d378f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x56291d35f2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #48: <unknown function> + 0x150582 (0x56291d378582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #49: PyObject_Call + 0xbc (0x56291d378f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x56291d35f2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #51: _PyFunction_Vectorcall + 0x6c (0x56291d36ca2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x56291d365007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #53: _PyObject_Call_Prepend + 0x69 (0x56291d376c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #54: <unknown function> + 0x211239 (0x56291d439239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #55: PyObject_Call + 0x207 (0x56291d379067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x56291d35f2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #57: <unknown function> + 0x150582 (0x56291d378582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x56291d35d8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #59: <unknown function> + 0x150582 (0x56291d378582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #60: PyObject_Call + 0xbc (0x56291d378f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x56291d35f2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #62: <unknown function> + 0x150582 (0x56291d378582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: frame #63: PyObject_Call + 0xbc (0x56291d378f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10)
[default5]:[rank13]: . This may indicate a possible application crash on rank 0 or a network set up issue.
E0702 19:47:05.492000 139631175944000 torch/distributed/elastic/multiprocessing/api.py:826] failed (exitcode: 1) local_rank: 1 (pid: 674635) of binary: /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10
Traceback (most recent call last):
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/torchrun", line 8, in <module>
sys.exit(main())
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 347, in wrapper
return f(*args, **kwargs)
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/run.py", line 879, in main
run(args)
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/run.py", line 870, in run
elastic_launch(
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 132, in __call__
return launch_agent(self._config, self._entrypoint, list(args))
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 263, in launch_agent
raise ChildFailedError(
torch.distributed.elastic.multiprocessing.errors.ChildFailedError:
============================================================
/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py FAILED
------------------------------------------------------------
Failures:
<NO_OTHER_FAILURES>
------------------------------------------------------------
Root Cause (first observed failure):
[0]:
time : 2024-07-02_19:47:02
host : ip-26-0-163-147.ec2.internal
rank : 1 (local_rank: 1)
exitcode : 1 (pid: 674635)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
============================================================
srun: error: ip-26-0-163-147: task 0: Exited with exit code 1
W0702 19:47:07.289000 139992688039680 torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1252] The node 'ip-26-0-163-226.ec2.internal_3080467_0' has failed to send a keep-alive heartbeat to the rendezvous 'none' due to an error of type RendezvousConnectionError.
W0702 19:47:07.866000 139998354859840 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 3080536 closing signal SIGTERM
W0702 19:47:07.867000 139998354859840 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 3080537 closing signal SIGTERM
W0702 19:47:07.868000 139998354859840 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 3080542 closing signal SIGTERM
W0702 19:47:07.869000 139998354859840 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 3080543 closing signal SIGTERM
E0702 19:47:09.494000 139998354859840 torch/distributed/elastic/multiprocessing/api.py:826] failed (exitcode: 1) local_rank: 2 (pid: 3080538) of binary: /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10
W0702 19:47:09.500000 139998354859840 torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1203] The node 'ip-26-0-163-226.ec2.internal_3080467_0' has failed to shutdown the rendezvous 'none' due to an error of type RendezvousConnectionError.
W0702 19:47:09.530000 139998354859840 torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1203] The node 'ip-26-0-163-226.ec2.internal_3080467_0' has failed to shutdown the rendezvous 'none' due to an error of type RendezvousConnectionError.
W0702 19:47:09.547000 139998354859840 torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1203] The node 'ip-26-0-163-226.ec2.internal_3080467_0' has failed to shutdown the rendezvous 'none' due to an error of type RendezvousConnectionError.
Traceback (most recent call last):
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/torchrun", line 8, in <module>
sys.exit(main())
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 347, in wrapper
return f(*args, **kwargs)
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/run.py", line 879, in main
run(args)
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/run.py", line 870, in run
elastic_launch(
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 132, in __call__
return launch_agent(self._config, self._entrypoint, list(args))
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 263, in launch_agent
raise ChildFailedError(
torch.distributed.elastic.multiprocessing.errors.ChildFailedError:
============================================================
/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py FAILED
------------------------------------------------------------
Failures:
[1]:
time : 2024-07-02_19:47:07
host : ip-26-0-163-226.ec2.internal
rank : 11 (local_rank: 3)
exitcode : 1 (pid: 3080539)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[2]:
time : 2024-07-02_19:47:07
host : ip-26-0-163-226.ec2.internal
rank : 12 (local_rank: 4)
exitcode : 1 (pid: 3080540)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
[3]:
time : 2024-07-02_19:47:07
host : ip-26-0-163-226.ec2.internal
rank : 13 (local_rank: 5)
exitcode : 1 (pid: 3080541)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
------------------------------------------------------------
Root Cause (first observed failure):
[0]:
time : 2024-07-02_19:47:07
host : ip-26-0-163-226.ec2.internal
rank : 10 (local_rank: 2)
exitcode : 1 (pid: 3080538)
error_file: <N/A>
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html
============================================================
srun: error: ip-26-0-163-226: task 1: Exited with exit code 1
Consider using `hf_transfer` for faster uploads. This solution comes with some limitations. See https://huggingface.co/docs/huggingface_hub/hf_transfer for more details.