|
======================== |
|
START TIME: Tue Jul 2 16:30:02 UTC 2024 |
|
python3 version = Python 3.10.14 |
|
======================== |
|
The token has not been saved to the git credentials helper. Pass `add_to_git_credential=True` in this function directly or `--add-to-git-credential` if using via `huggingface-cli` if you want to set the git credential as well. |
|
Token is valid (permission: write). |
|
OSError: [Errno 122] Disk quota exceeded |
|
|
|
During handling of the above exception, another exception occurred: |
|
|
|
Traceback (most recent call last): |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/huggingface-cli", line 8, in <module> |
|
sys.exit(main()) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/huggingface_hub/commands/huggingface_cli.py", line 51, in main |
|
service.run() |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/huggingface_hub/commands/user.py", line 98, in run |
|
login(token=self.args.token, add_to_git_credential=self.args.add_to_git_credential) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/huggingface_hub/_login.py", line 111, in login |
|
_login(token, add_to_git_credential=add_to_git_credential, write_permission=write_permission) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/huggingface_hub/_login.py", line 328, in _login |
|
path.write_text(token) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/pathlib.py", line 1154, in write_text |
|
with self.open(mode='w', encoding=encoding, errors=errors, newline=newline) as f: |
|
OSError: [Errno 122] Disk quota exceeded |
|
Already on 'bench_cluster' |
|
M examples/config_tiny_llama.py |
|
M examples/config_tiny_llama.yaml |
|
M examples/train_tiny_llama.sh |
|
M src/nanotron/models/llama.py |
|
M src/nanotron/trainer.py |
|
Your branch is up to date with 'origin/bench_cluster'. |
|
Job status: RUNNING |
|
W0702 16:30:05.241000 140670700869440 torch/distributed/run.py:757] |
|
W0702 16:30:05.241000 140670700869440 torch/distributed/run.py:757] ***************************************** |
|
W0702 16:30:05.241000 140670700869440 torch/distributed/run.py:757] Setting OMP_NUM_THREADS environment variable for each process to be 1 in default, to avoid your system being overloaded, please further tune the variable for optimal performance in your application as needed. |
|
W0702 16:30:05.241000 140670700869440 torch/distributed/run.py:757] ***************************************** |
|
W0702 16:30:05.281000 140356314117952 torch/distributed/run.py:757] |
|
W0702 16:30:05.281000 140356314117952 torch/distributed/run.py:757] ***************************************** |
|
W0702 16:30:05.281000 140356314117952 torch/distributed/run.py:757] Setting OMP_NUM_THREADS environment variable for each process to be 1 in default, to avoid your system being overloaded, please further tune the variable for optimal performance in your application as needed. |
|
W0702 16:30:05.281000 140356314117952 torch/distributed/run.py:757] ***************************************** |
|
[default0]:07/02/2024 16:30:22 [WARNING|DP=0|PP=0|TP=0|ip-26-0-160-225]: [Vocab Size Padding] Padded vocab (size: 50257) with 3 dummy tokens (new size: 50260) |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: Config: |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: Config(general=GeneralArgs(project='bench_cluster', |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: run='%date_%jobid', |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: seed=42, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: step=None, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: consumed_train_samples=None, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: benchmark_csv_path=None, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: ignore_sanity_checks=True), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: parallelism=ParallelismArgs(dp=1, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: pp=4, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: tp=4, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: pp_engine=<nanotron.parallel.pipeline_parallel.engine.OneForwardOneBackwardPipelineEngine object at 0x7f7ea463c730>, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: tp_mode=<TensorParallelLinearMode.REDUCE_SCATTER: 2>, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: tp_linear_async_communication=False, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: expert_parallel_size=1), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: model=ModelArgs(model_config=LlamaConfig(bos_token_id=1, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: eos_token_id=2, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: hidden_act='silu', |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: hidden_size=2048, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: initializer_range=0.02, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: intermediate_size=4096, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: is_llama_config=True, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: max_position_embeddings=4096, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: num_attention_heads=32, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: num_hidden_layers=24, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: num_key_value_heads=32, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: pad_token_id=None, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: pretraining_tp=1, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: rms_norm_eps=1e-05, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: rope_scaling=None, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: rope_theta=10000.0, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: tie_word_embeddings=True, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: use_cache=True, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: vocab_size=50260), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: init_method=RandomInit(std=0.025), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: dtype=torch.bfloat16, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: make_vocab_size_divisible_by=1, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: ddp_bucket_cap_mb=25), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: tokenizer=TokenizerArgs(tokenizer_name_or_path='openai-community/gpt2', |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: tokenizer_revision=None, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: tokenizer_max_length=None), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: checkpoints=CheckpointsArgs(checkpoints_path=Path('/dev/null'), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: checkpoint_interval=100000, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: save_initial_state=False, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: resume_checkpoint_path=None, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: checkpoints_path_is_shared_file_system=False), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: logging=LoggingArgs(log_level='info', |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: log_level_replica='info', |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: iteration_step_info_interval=1), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: tokens=TokensArgs(sequence_length=4096, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: train_steps=20, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: micro_batch_size=256, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: batch_accumulation_per_replica=4, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: val_check_interval=-1, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: limit_val_batches=0, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: limit_test_batches=0), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: optimizer=OptimizerArgs(optimizer_factory=AdamWOptimizerArgs(adam_eps=1e-08, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: adam_beta1=0.9, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: adam_beta2=0.95, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: torch_adam_is_fused=True, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: name='adamW'), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: zero_stage=1, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: weight_decay=0.01, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: clip_grad=1.0, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: accumulate_grad_in_fp32=True, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: learning_rate_scheduler=LRSchedulerArgs(learning_rate=0.0001, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: lr_warmup_steps=1, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: lr_warmup_style='linear', |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: lr_decay_style='linear', |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: lr_decay_steps=19, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: lr_decay_starting_step=None, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: min_decay_lr=1e-05)), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: data_stages=[DatasetStageArgs(name='Training Stage', |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: start_training_step=1, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: data=DataArgs(dataset=PretrainDatasetsArgs(hf_dataset_or_datasets='roneneldan/TinyStories', |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: hf_dataset_splits='train', |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: hf_dataset_config_name=None, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: dataset_processing_num_proc_per_process=64, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: dataset_overwrite_cache=False, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: text_column_name='text'), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: seed=42, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: num_loading_workers=32))], |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: profiler=ProfilerArgs(profiler_export_path=Path('/fsx/ferdinandmom/ferdinand-hf/bench_cluster/results/llama-1B/16_GPUS/dp-1_tp-4_pp-4_mbz-256')), |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: lighteval=None) |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: Model Config: |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: LlamaConfig(bos_token_id=1, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: eos_token_id=2, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: hidden_act='silu', |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: hidden_size=2048, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: initializer_range=0.02, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: intermediate_size=4096, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: is_llama_config=True, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: max_position_embeddings=4096, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: num_attention_heads=32, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: num_hidden_layers=24, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: num_key_value_heads=32, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: pad_token_id=None, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: pretraining_tp=1, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: rms_norm_eps=1e-05, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: rope_scaling=None, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: rope_theta=10000.0, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: tie_word_embeddings=True, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: use_cache=True, |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: vocab_size=50260) |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: Building model.. |
|
[default0]:07/02/2024 16:30:22 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: Setting PP block ranks... |
|
[default2]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=2|ip-26-0-160-225]: Local number of parameters: 99.2M (189.14MiB) |
|
[default2]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=2|ip-26-0-160-225]: [After model building] Memory usage: 197.07MiB. Peak allocated: 199.10MiB Peak reserved: 200.00MiB |
|
[default1]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=1|ip-26-0-160-225]: Local number of parameters: 99.2M (189.14MiB) |
|
[default1]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=1|ip-26-0-160-225]: [After model building] Memory usage: 197.07MiB. Peak allocated: 199.10MiB Peak reserved: 200.00MiB |
|
[default1]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=1|ip-26-0-160-225]: No checkpoint path provided. |
|
[default2]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=2|ip-26-0-160-225]: No checkpoint path provided. |
|
[default5]:07/02/2024 16:30:37 [INFO|DP=0|PP=3|TP=1|ip-26-0-171-56]: Local number of parameters: 67.7M (129.12MiB) |
|
[default5]:07/02/2024 16:30:37 [INFO|DP=0|PP=3|TP=1|ip-26-0-171-56]: [After model building] Memory usage: 134.05MiB. Peak allocated: 136.08MiB Peak reserved: 138.00MiB |
|
[default5]:07/02/2024 16:30:37 [INFO|DP=0|PP=3|TP=1|ip-26-0-171-56]: No checkpoint path provided. |
|
[default7]:07/02/2024 16:30:37 [INFO|DP=0|PP=3|TP=3|ip-26-0-171-56]: Local number of parameters: 67.7M (129.12MiB) |
|
[default7]:07/02/2024 16:30:37 [INFO|DP=0|PP=3|TP=3|ip-26-0-171-56]: [After model building] Memory usage: 134.05MiB. Peak allocated: 136.08MiB Peak reserved: 138.00MiB |
|
[default7]:07/02/2024 16:30:37 [INFO|DP=0|PP=3|TP=3|ip-26-0-171-56]: No checkpoint path provided. |
|
[default7]:07/02/2024 16:30:37 [INFO|DP=0|PP=1|TP=3|ip-26-0-160-225]: Local number of parameters: 73.4M (140.05MiB) |
|
[default7]:07/02/2024 16:30:37 [INFO|DP=0|PP=1|TP=3|ip-26-0-160-225]: [After model building] Memory usage: 147.07MiB. Peak allocated: 149.10MiB Peak reserved: 150.00MiB |
|
[default7]:07/02/2024 16:30:37 [INFO|DP=0|PP=1|TP=3|ip-26-0-160-225]: No checkpoint path provided. |
|
[default4]:07/02/2024 16:30:37 [INFO|DP=0|PP=3|TP=0|ip-26-0-171-56]: Local number of parameters: 67.7M (129.12MiB) |
|
[default4]:07/02/2024 16:30:37 [INFO|DP=0|PP=3|TP=0|ip-26-0-171-56]: [After model building] Memory usage: 134.05MiB. Peak allocated: 136.08MiB Peak reserved: 138.00MiB |
|
[default3]:07/02/2024 16:30:37 [INFO|DP=0|PP=2|TP=3|ip-26-0-171-56]: Local number of parameters: 62.9M (120.05MiB) |
|
[default3]:07/02/2024 16:30:37 [INFO|DP=0|PP=2|TP=3|ip-26-0-171-56]: [After model building] Memory usage: 126.06MiB. Peak allocated: 128.09MiB Peak reserved: 130.00MiB |
|
[default4]:07/02/2024 16:30:37 [INFO|DP=0|PP=3|TP=0|ip-26-0-171-56]: No checkpoint path provided. |
|
[default3]:07/02/2024 16:30:37 [INFO|DP=0|PP=2|TP=3|ip-26-0-171-56]: No checkpoint path provided. |
|
[default0]:07/02/2024 16:30:37 [INFO|DP=0|PP=2|TP=0|ip-26-0-171-56]: Local number of parameters: 62.9M (120.05MiB) |
|
[default0]:07/02/2024 16:30:37 [INFO|DP=0|PP=2|TP=0|ip-26-0-171-56]: [After model building] Memory usage: 126.06MiB. Peak allocated: 128.09MiB Peak reserved: 130.00MiB |
|
[default0]:07/02/2024 16:30:37 [INFO|DP=0|PP=2|TP=0|ip-26-0-171-56]: No checkpoint path provided. |
|
[default6]:07/02/2024 16:30:37 [INFO|DP=0|PP=3|TP=2|ip-26-0-171-56]: Local number of parameters: 67.7M (129.12MiB) |
|
[default2]:07/02/2024 16:30:37 [INFO|DP=0|PP=2|TP=2|ip-26-0-171-56]: Local number of parameters: 62.9M (120.05MiB) |
|
[default2]:07/02/2024 16:30:37 [INFO|DP=0|PP=2|TP=2|ip-26-0-171-56]: [After model building] Memory usage: 126.06MiB. Peak allocated: 128.09MiB Peak reserved: 130.00MiB |
|
[default6]:07/02/2024 16:30:37 [INFO|DP=0|PP=3|TP=2|ip-26-0-171-56]: [After model building] Memory usage: 134.05MiB. Peak allocated: 136.08MiB Peak reserved: 138.00MiB |
|
[default6]:07/02/2024 16:30:37 [INFO|DP=0|PP=3|TP=2|ip-26-0-171-56]: No checkpoint path provided. |
|
[default2]:07/02/2024 16:30:37 [INFO|DP=0|PP=2|TP=2|ip-26-0-171-56]: No checkpoint path provided. |
|
[default1]:07/02/2024 16:30:37 [INFO|DP=0|PP=2|TP=1|ip-26-0-171-56]: Local number of parameters: 62.9M (120.05MiB) |
|
[default1]:07/02/2024 16:30:37 [INFO|DP=0|PP=2|TP=1|ip-26-0-171-56]: [After model building] Memory usage: 126.06MiB. Peak allocated: 128.09MiB Peak reserved: 130.00MiB |
|
[default1]:07/02/2024 16:30:37 [INFO|DP=0|PP=2|TP=1|ip-26-0-171-56]: No checkpoint path provided. |
|
[default4]:07/02/2024 16:30:37 [INFO|DP=0|PP=1|TP=0|ip-26-0-160-225]: Local number of parameters: 73.4M (140.05MiB) |
|
[default4]:07/02/2024 16:30:37 [INFO|DP=0|PP=1|TP=0|ip-26-0-160-225]: [After model building] Memory usage: 147.07MiB. Peak allocated: 149.10MiB Peak reserved: 150.00MiB |
|
[default4]:07/02/2024 16:30:37 [INFO|DP=0|PP=1|TP=0|ip-26-0-160-225]: No checkpoint path provided. |
|
[default0]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: Total number of parameters: 1.21G (2313.42MiB) |
|
[default0]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: Local number of parameters: 99.2M (189.14MiB) |
|
[default0]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: [After model building] Memory usage: 197.07MiB. Peak allocated: 199.10MiB Peak reserved: 200.00MiB |
|
[default0]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: No checkpoint path provided. |
|
[default0]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: Parametrizing model parameters using StandardParametrizator |
|
[default5]:07/02/2024 16:30:37 [INFO|DP=0|PP=1|TP=1|ip-26-0-160-225]: Local number of parameters: 73.4M (140.05MiB) |
|
[default5]:07/02/2024 16:30:37 [INFO|DP=0|PP=1|TP=1|ip-26-0-160-225]: [After model building] Memory usage: 147.07MiB. Peak allocated: 149.10MiB Peak reserved: 150.00MiB |
|
[default5]:07/02/2024 16:30:37 [INFO|DP=0|PP=1|TP=1|ip-26-0-160-225]: No checkpoint path provided. |
|
[default3]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=3|ip-26-0-160-225]: Local number of parameters: 99.2M (189.14MiB) |
|
[default3]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=3|ip-26-0-160-225]: [After model building] Memory usage: 197.07MiB. Peak allocated: 199.10MiB Peak reserved: 200.00MiB |
|
[default6]:07/02/2024 16:30:37 [INFO|DP=0|PP=1|TP=2|ip-26-0-160-225]: Local number of parameters: 73.4M (140.05MiB) |
|
[default6]:07/02/2024 16:30:37 [INFO|DP=0|PP=1|TP=2|ip-26-0-160-225]: [After model building] Memory usage: 147.07MiB. Peak allocated: 149.10MiB Peak reserved: 150.00MiB |
|
[default6]:07/02/2024 16:30:37 [INFO|DP=0|PP=1|TP=2|ip-26-0-160-225]: No checkpoint path provided. |
|
[default3]:07/02/2024 16:30:37 [INFO|DP=0|PP=0|TP=3|ip-26-0-160-225]: No checkpoint path provided. |
|
[default0]:07/02/2024 16:30:38 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: [Optimizer Building] Using LearningRateForSP as learning rate |
|
[default0]:07/02/2024 16:30:38 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: [ZeRO sharding] Size of optimizer params per rank: |
|
[default0]:07/02/2024 16:30:38 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: [ZeRO sharding] DP Rank 0 has 99.2M out of 99.2M (100.00%) params' optimizer states |
|
[default0]:07/02/2024 16:30:39 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: [Training Plan] Stage Training Stage has 19 remaining training steps and has consumed 0 samples |
|
[default0]:07/02/2024 16:30:39 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: Using `datasets` library |
|
[default0]:07/02/2024 16:30:39 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: Loading tokenizer from openai-community/gpt2 and transformers/hf_hub versions ('4.41.2', '0.23.4') |
|
[default0]:07/02/2024 16:30:39 [WARNING|DP=0|PP=0|TP=0|ip-26-0-160-225]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default0]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default0]:07/02/2024 16:30:40 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: [Training Plan] There are 1 training stages |
|
[default0]:07/02/2024 16:30:40 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: [Stage Training Stage] start from step 1 |
|
[default0]:07/02/2024 16:30:40 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: |
|
[default0]:07/02/2024 16:30:40 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: [Start training] datetime: 2024-07-02 16:30:40.470610 | mbs: 256 | grad_accum: 4 | global_batch_size: 1024 | sequence_length: 4096 | train_steps: 20 | start_iteration_step: 0 | consumed_train_samples: 0 |
|
[default0]:07/02/2024 16:30:40 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: Resuming training from stage Training Stage, it has trained for 0 samples and has 19 remaining train steps |
|
[default0]:07/02/2024 16:30:40 [INFO|DP=0|PP=0|TP=0|ip-26-0-160-225]: Memory usage: 953.61MiB. Peak allocated 953.61MiB. Peak reserved: 960.00MiB |
|
[default5]:07/02/2024 16:30:40 [WARNING|DP=0|PP=3|TP=1|ip-26-0-171-56]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default7]:07/02/2024 16:30:40 [WARNING|DP=0|PP=3|TP=3|ip-26-0-171-56]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default0]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default7]:07/02/2024 16:30:40 [WARNING|DP=0|PP=1|TP=3|ip-26-0-160-225]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default1]:07/02/2024 16:30:40 [WARNING|DP=0|PP=0|TP=1|ip-26-0-160-225]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default2]:07/02/2024 16:30:40 [WARNING|DP=0|PP=0|TP=2|ip-26-0-160-225]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default4]:07/02/2024 16:30:40 [WARNING|DP=0|PP=3|TP=0|ip-26-0-171-56]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default7]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default3]:07/02/2024 16:30:40 [WARNING|DP=0|PP=2|TP=3|ip-26-0-171-56]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default0]:07/02/2024 16:30:40 [WARNING|DP=0|PP=2|TP=0|ip-26-0-171-56]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default6]:07/02/2024 16:30:40 [WARNING|DP=0|PP=3|TP=2|ip-26-0-171-56]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default1]:07/02/2024 16:30:40 [WARNING|DP=0|PP=2|TP=1|ip-26-0-171-56]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default2]:07/02/2024 16:30:40 [WARNING|DP=0|PP=2|TP=2|ip-26-0-171-56]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default1]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default3]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default6]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default5]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default2]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default6]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default4]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default5]:07/02/2024 16:30:40 [WARNING|DP=0|PP=1|TP=1|ip-26-0-160-225]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default5]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default7]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default6]:07/02/2024 16:30:40 [WARNING|DP=0|PP=1|TP=2|ip-26-0-160-225]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default1]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default2]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default4]:07/02/2024 16:30:40 [WARNING|DP=0|PP=1|TP=0|ip-26-0-160-225]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default3]:07/02/2024 16:30:40 [WARNING|DP=0|PP=0|TP=3|ip-26-0-160-225]: Repo card metadata block was not found. Setting CardData to empty. |
|
[default4]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default3]:Repo card metadata block was not found. Setting CardData to empty. |
|
[default1]:[rank1]: OSError: [Errno 122] Disk quota exceeded |
|
[default1]: |
|
[default1]:[rank1]: During handling of the above exception, another exception occurred: |
|
[default1]: |
|
[default1]:[rank1]: Traceback (most recent call last): |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module> |
|
[default1]:[rank1]: trainer.train(dataloader) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train |
|
[default1]:[rank1]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step |
|
[default1]:[rank1]: outputs = self.pipeline_engine.train_batch_iter( |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter |
|
[default1]:[rank1]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward |
|
[default1]:[rank1]: output = model(**micro_batch) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default1]:[rank1]: return self._call_impl(*args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default1]:[rank1]: return forward_call(*args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward |
|
[default1]:[rank1]: sharded_logits = self.model( |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default1]:[rank1]: return self._call_impl(*args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default1]:[rank1]: return forward_call(*args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward |
|
[default1]:[rank1]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0] |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states |
|
[default1]:[rank1]: hidden_encoder_states = encoder_block(**hidden_encoder_states) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default1]:[rank1]: return self._call_impl(*args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default1]:[rank1]: return forward_call(*args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 151, in forward |
|
[default1]:[rank1]: output = self.pp_block(**new_kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default1]:[rank1]: return self._call_impl(*args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default1]:[rank1]: return forward_call(*args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 629, in forward |
|
[default1]:[rank1]: hidden_states = self.input_layernorm(hidden_states) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default1]:[rank1]: return self._call_impl(*args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default1]:[rank1]: return forward_call(*args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/nn/layer_norm.py", line 42, in forward |
|
[default1]:[rank1]: return layer_norm_fn( |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/flash_attn/ops/triton/layer_norm.py", line 875, in layer_norm_fn |
|
[default1]:[rank1]: return LayerNormFn.apply( |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/autograd/function.py", line 598, in apply |
|
[default1]:[rank1]: return super().apply(*args, **kwargs) # type: ignore[misc] |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/flash_attn/ops/triton/layer_norm.py", line 748, in forward |
|
[default1]:[rank1]: y, y1, mean, rstd, residual_out, seeds, dropout_mask, dropout_mask1 = _layer_norm_fwd( |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/flash_attn/ops/triton/layer_norm.py", line 335, in _layer_norm_fwd |
|
[default1]:[rank1]: _layer_norm_fwd_1pass_kernel[(M,)]( |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/triton/runtime/jit.py", line 167, in <lambda> |
|
[default1]:[rank1]: return lambda *args, **kwargs: self.run(grid=grid, warmup=False, *args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/triton/runtime/autotuner.py", line 143, in run |
|
[default1]:[rank1]: timings = {config: self._bench(*args, config=config, **kwargs) for config in pruned_configs} |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/triton/runtime/autotuner.py", line 143, in <dictcomp> |
|
[default1]:[rank1]: timings = {config: self._bench(*args, config=config, **kwargs) for config in pruned_configs} |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/triton/runtime/autotuner.py", line 122, in _bench |
|
[default1]:[rank1]: return do_bench(kernel_call, warmup=self.warmup, rep=self.rep, quantiles=(0.5, 0.2, 0.8)) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/triton/testing.py", line 102, in do_bench |
|
[default1]:[rank1]: fn() |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/triton/runtime/autotuner.py", line 110, in kernel_call |
|
[default1]:[rank1]: self.fn.run( |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/triton/runtime/autotuner.py", line 305, in run |
|
[default1]:[rank1]: return self.fn.run(*args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/triton/runtime/autotuner.py", line 305, in run |
|
[default1]:[rank1]: return self.fn.run(*args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/triton/runtime/autotuner.py", line 305, in run |
|
[default1]:[rank1]: return self.fn.run(*args, **kwargs) |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/triton/runtime/jit.py", line 416, in run |
|
[default1]:[rank1]: self.cache[device][key] = compile( |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/triton/compiler/compiler.py", line 194, in compile |
|
[default1]:[rank1]: metadata_group[f"{src.name}.{ext}"] = fn_cache_manager.put(next_module, f"{src.name}.{ext}") |
|
[default1]:[rank1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/triton/runtime/cache.py", line 123, in put |
|
[default1]:[rank1]: with open(temp_path, mode) as f: |
|
[default1]:[rank1]: OSError: [Errno 122] Disk quota exceeded |
|
[default1]:Exception in thread Thread-2 (_pin_memory_loop): |
|
[default1]:Traceback (most recent call last): |
|
[default1]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/threading.py", line 1016, in _bootstrap_inner |
|
W0702 16:30:51.522000 140670700869440 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 1371214 closing signal SIGTERM |
|
W0702 16:30:51.525000 140670700869440 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 1371216 closing signal SIGTERM |
|
W0702 16:30:51.530000 140670700869440 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 1371217 closing signal SIGTERM |
|
W0702 16:30:51.533000 140670700869440 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 1371218 closing signal SIGTERM |
|
W0702 16:30:51.534000 140670700869440 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 1371219 closing signal SIGTERM |
|
W0702 16:30:51.535000 140670700869440 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 1371220 closing signal SIGTERM |
|
W0702 16:30:51.536000 140670700869440 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 1371221 closing signal SIGTERM |
|
[default0]:[rank8]: Traceback (most recent call last): |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module> |
|
[default0]:[rank8]: trainer.train(dataloader) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train |
|
[default0]:[rank8]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step |
|
[default0]:[rank8]: outputs = self.pipeline_engine.train_batch_iter( |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter |
|
[default0]:[rank8]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward |
|
[default0]:[rank8]: output = model(**micro_batch) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default0]:[rank8]: return self._call_impl(*args, **kwargs) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default0]:[rank8]: return forward_call(*args, **kwargs) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward |
|
[default0]:[rank8]: sharded_logits = self.model( |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default0]:[rank8]: return self._call_impl(*args, **kwargs) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default0]:[rank8]: return forward_call(*args, **kwargs) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward |
|
[default0]:[rank8]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0] |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states |
|
[default0]:[rank8]: hidden_encoder_states = encoder_block(**hidden_encoder_states) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default0]:[rank8]: return self._call_impl(*args, **kwargs) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default0]:[rank8]: return forward_call(*args, **kwargs) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward |
|
[default0]:[rank8]: new_kwargs[name] = recv_from_pipeline_state_buffer( |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer |
|
[default0]:[rank8]: pipeline_state.run_communication() |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication |
|
[default0]:[rank8]: recv_activation_tensor = recv_activation() |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__ |
|
[default0]:[rank8]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0] |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors |
|
[default0]:[rank8]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors |
|
[default0]:[rank8]: meta = self._recv_meta(from_rank=from_rank, tag=tag) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta |
|
[default0]:[rank8]: dist.recv( |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper |
|
[default0]:[rank8]: return func(*args, **kwargs) |
|
[default0]:[rank8]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv |
|
[default0]:[rank8]: pg.recv([tensor], group_src_rank, tag).wait() |
|
[default0]:[rank8]: torch.distributed.DistBackendError: [2] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '1:2', but store->get('1:2') got error: Connection reset by peer |
|
[default0]:[rank8]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first): |
|
[default0]:[rank8]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7fa140f6f897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so) |
|
[default0]:[rank8]: frame #1: <unknown function> + 0x5b3a23e (0x7fa17aa8c23e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7fa17aa86c87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7fa17aa86f82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7fa17aa87fd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fa17aa3c371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fa17aa3c371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fa17aa3c371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fa17aa3c371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7fa142249189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default0]:[rank8]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7fa142250610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default0]:[rank8]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7fa14226f978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default0]:[rank8]: frame #12: <unknown function> + 0x5adc309 (0x7fa17aa2e309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #13: <unknown function> + 0x5ae6f10 (0x7fa17aa38f10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #14: <unknown function> + 0x5ae6fa5 (0x7fa17aa38fa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #15: <unknown function> + 0x5124446 (0x7fa17a076446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #16: <unknown function> + 0x1acf4b8 (0x7fa176a214b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #17: <unknown function> + 0x5aee004 (0x7fa17aa40004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #18: <unknown function> + 0x5af36b5 (0x7fa17aa456b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default0]:[rank8]: frame #19: <unknown function> + 0xd2631e (0x7fa18d62f31e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default0]:[rank8]: frame #20: <unknown function> + 0x47def4 (0x7fa18cd86ef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default0]:[rank8]: frame #21: <unknown function> + 0x1445a6 (0x561950c5b5a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #22: _PyObject_MakeTpCall + 0x26b (0x561950c54a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #23: <unknown function> + 0x150866 (0x561950c67866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x561950c50142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #25: _PyFunction_Vectorcall + 0x6c (0x561950c5ba2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #26: PyObject_Call + 0xbc (0x561950c67f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x561950c4e2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #28: _PyFunction_Vectorcall + 0x6c (0x561950c5ba2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x561950c4c8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #30: <unknown function> + 0x150582 (0x561950c67582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x561950c4c8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #32: <unknown function> + 0x150582 (0x561950c67582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x561950c4c8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #34: <unknown function> + 0x150582 (0x561950c67582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x561950c4c8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x561950c53f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #37: _PyObject_Call_Prepend + 0x69 (0x561950c65c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #38: <unknown function> + 0x211239 (0x561950d28239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #39: _PyObject_MakeTpCall + 0x26b (0x561950c54a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x561950c503e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #41: _PyFunction_Vectorcall + 0x6c (0x561950c5ba2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x561950c4bc5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #43: _PyFunction_Vectorcall + 0x6c (0x561950c5ba2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x561950c4c8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #45: <unknown function> + 0x150582 (0x561950c67582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #46: PyObject_Call + 0xbc (0x561950c67f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x561950c4e2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #48: <unknown function> + 0x150582 (0x561950c67582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #49: PyObject_Call + 0xbc (0x561950c67f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x561950c4e2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #51: _PyFunction_Vectorcall + 0x6c (0x561950c5ba2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x561950c54007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #53: _PyObject_Call_Prepend + 0x69 (0x561950c65c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #54: <unknown function> + 0x211239 (0x561950d28239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #55: PyObject_Call + 0x207 (0x561950c68067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x561950c4e2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #57: <unknown function> + 0x150582 (0x561950c67582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x561950c4c8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #59: <unknown function> + 0x150582 (0x561950c67582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #60: PyObject_Call + 0xbc (0x561950c67f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x561950c4e2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #62: <unknown function> + 0x150582 (0x561950c67582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: frame #63: PyObject_Call + 0xbc (0x561950c67f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default0]:[rank8]: . This may indicate a possible application crash on rank 0 or a network set up issue. |
|
[default7]:[rank15]: Traceback (most recent call last): |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module> |
|
[default7]:[rank15]: trainer.train(dataloader) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train |
|
[default7]:[rank15]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step |
|
[default7]:[rank15]: outputs = self.pipeline_engine.train_batch_iter( |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter |
|
[default7]:[rank15]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward |
|
[default7]:[rank15]: output = model(**micro_batch) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default7]:[rank15]: return self._call_impl(*args, **kwargs) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default7]:[rank15]: return forward_call(*args, **kwargs) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward |
|
[default7]:[rank15]: sharded_logits = self.model( |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default7]:[rank15]: return self._call_impl(*args, **kwargs) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default7]:[rank15]: return forward_call(*args, **kwargs) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward |
|
[default7]:[rank15]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0] |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states |
|
[default7]:[rank15]: hidden_encoder_states = encoder_block(**hidden_encoder_states) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default7]:[rank15]: return self._call_impl(*args, **kwargs) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default7]:[rank15]: return forward_call(*args, **kwargs) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward |
|
[default7]:[rank15]: new_kwargs[name] = recv_from_pipeline_state_buffer( |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer |
|
[default7]:[rank15]: pipeline_state.run_communication() |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication |
|
[default7]:[rank15]: recv_activation_tensor = recv_activation() |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__ |
|
[default7]:[rank15]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0] |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors |
|
[default7]:[rank15]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors |
|
[default7]:[rank15]: meta = self._recv_meta(from_rank=from_rank, tag=tag) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta |
|
[default7]:[rank15]: dist.recv( |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper |
|
[default7]:[rank15]: return func(*args, **kwargs) |
|
[default7]:[rank15]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv |
|
[default7]:[rank15]: pg.recv([tensor], group_src_rank, tag).wait() |
|
[default7]:[rank15]: torch.distributed.DistBackendError: [3] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '2:3', but store->get('2:3') got error: Connection reset by peer |
|
[default7]:[rank15]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first): |
|
[default7]:[rank15]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7fb587109897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so) |
|
[default7]:[rank15]: frame #1: <unknown function> + 0x5b3a23e (0x7fb5c0c2623e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7fb5c0c20c87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7fb5c0c20f82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7fb5c0c21fd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fb5c0bd6371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fb5c0bd6371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fb5c0bd6371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fb5c0bd6371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7fb5883e3189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default7]:[rank15]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7fb5883ea610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default7]:[rank15]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7fb588409978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default7]:[rank15]: frame #12: <unknown function> + 0x5adc309 (0x7fb5c0bc8309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #13: <unknown function> + 0x5ae6f10 (0x7fb5c0bd2f10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #14: <unknown function> + 0x5ae6fa5 (0x7fb5c0bd2fa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #15: <unknown function> + 0x5124446 (0x7fb5c0210446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #16: <unknown function> + 0x1acf4b8 (0x7fb5bcbbb4b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #17: <unknown function> + 0x5aee004 (0x7fb5c0bda004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #18: <unknown function> + 0x5af36b5 (0x7fb5c0bdf6b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default7]:[rank15]: frame #19: <unknown function> + 0xd2631e (0x7fb5d37c931e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default7]:[rank15]: frame #20: <unknown function> + 0x47def4 (0x7fb5d2f20ef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default7]:[rank15]: frame #21: <unknown function> + 0x1445a6 (0x560ca76775a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #22: _PyObject_MakeTpCall + 0x26b (0x560ca7670a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #23: <unknown function> + 0x150866 (0x560ca7683866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x560ca766c142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #25: _PyFunction_Vectorcall + 0x6c (0x560ca7677a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #26: PyObject_Call + 0xbc (0x560ca7683f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x560ca766a2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #28: _PyFunction_Vectorcall + 0x6c (0x560ca7677a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x560ca76688fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #30: <unknown function> + 0x150582 (0x560ca7683582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x560ca76688fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #32: <unknown function> + 0x150582 (0x560ca7683582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x560ca76688fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #34: <unknown function> + 0x150582 (0x560ca7683582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x560ca76688fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x560ca766ff50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #37: _PyObject_Call_Prepend + 0x69 (0x560ca7681c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #38: <unknown function> + 0x211239 (0x560ca7744239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #39: _PyObject_MakeTpCall + 0x26b (0x560ca7670a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x560ca766c3e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #41: _PyFunction_Vectorcall + 0x6c (0x560ca7677a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x560ca7667c5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #43: _PyFunction_Vectorcall + 0x6c (0x560ca7677a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x560ca76688fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #45: <unknown function> + 0x150582 (0x560ca7683582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #46: PyObject_Call + 0xbc (0x560ca7683f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x560ca766a2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #48: <unknown function> + 0x150582 (0x560ca7683582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #49: PyObject_Call + 0xbc (0x560ca7683f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x560ca766a2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #51: _PyFunction_Vectorcall + 0x6c (0x560ca7677a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x560ca7670007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #53: _PyObject_Call_Prepend + 0x69 (0x560ca7681c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #54: <unknown function> + 0x211239 (0x560ca7744239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #55: PyObject_Call + 0x207 (0x560ca7684067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x560ca766a2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #57: <unknown function> + 0x150582 (0x560ca7683582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x560ca76688fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #59: <unknown function> + 0x150582 (0x560ca7683582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #60: PyObject_Call + 0xbc (0x560ca7683f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x560ca766a2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #62: <unknown function> + 0x150582 (0x560ca7683582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: frame #63: PyObject_Call + 0xbc (0x560ca7683f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default7]:[rank15]: . This may indicate a possible application crash on rank 0 or a network set up issue. |
|
[default3]:[rank11]: Traceback (most recent call last): |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module> |
|
[default3]:[rank11]: trainer.train(dataloader) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train |
|
[default3]:[rank11]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step |
|
[default3]:[rank11]: outputs = self.pipeline_engine.train_batch_iter( |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter |
|
[default3]:[rank11]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward |
|
[default3]:[rank11]: output = model(**micro_batch) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default3]:[rank11]: return self._call_impl(*args, **kwargs) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default3]:[rank11]: return forward_call(*args, **kwargs) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward |
|
[default3]:[rank11]: sharded_logits = self.model( |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default3]:[rank11]: return self._call_impl(*args, **kwargs) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default3]:[rank11]: return forward_call(*args, **kwargs) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward |
|
[default3]:[rank11]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0] |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states |
|
[default3]:[rank11]: hidden_encoder_states = encoder_block(**hidden_encoder_states) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default3]:[rank11]: return self._call_impl(*args, **kwargs) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default3]:[rank11]: return forward_call(*args, **kwargs) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward |
|
[default3]:[rank11]: new_kwargs[name] = recv_from_pipeline_state_buffer( |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer |
|
[default3]:[rank11]: pipeline_state.run_communication() |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication |
|
[default3]:[rank11]: recv_activation_tensor = recv_activation() |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__ |
|
[default3]:[rank11]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0] |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors |
|
[default3]:[rank11]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors |
|
[default3]:[rank11]: meta = self._recv_meta(from_rank=from_rank, tag=tag) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta |
|
[default3]:[rank11]: dist.recv( |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper |
|
[default3]:[rank11]: return func(*args, **kwargs) |
|
[default3]:[rank11]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv |
|
[default3]:[rank11]: pg.recv([tensor], group_src_rank, tag).wait() |
|
[default3]:[rank11]: torch.distributed.DistBackendError: [2] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '1:2', but store->get('1:2') got error: Connection reset by peer |
|
[default3]:[rank11]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first): |
|
[default3]:[rank11]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7f5fb9b73897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so) |
|
[default3]:[rank11]: frame #1: <unknown function> + 0x5b3a23e (0x7f5ff369023e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7f5ff368ac87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7f5ff368af82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7f5ff368bfd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f5ff3640371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f5ff3640371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f5ff3640371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f5ff3640371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7f5fbae4d189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default3]:[rank11]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7f5fbae54610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default3]:[rank11]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7f5fbae73978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default3]:[rank11]: frame #12: <unknown function> + 0x5adc309 (0x7f5ff3632309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #13: <unknown function> + 0x5ae6f10 (0x7f5ff363cf10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #14: <unknown function> + 0x5ae6fa5 (0x7f5ff363cfa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #15: <unknown function> + 0x5124446 (0x7f5ff2c7a446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #16: <unknown function> + 0x1acf4b8 (0x7f5fef6254b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #17: <unknown function> + 0x5aee004 (0x7f5ff3644004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #18: <unknown function> + 0x5af36b5 (0x7f5ff36496b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default3]:[rank11]: frame #19: <unknown function> + 0xd2631e (0x7f600623331e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default3]:[rank11]: frame #20: <unknown function> + 0x47def4 (0x7f600598aef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default3]:[rank11]: frame #21: <unknown function> + 0x1445a6 (0x5598550ee5a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #22: _PyObject_MakeTpCall + 0x26b (0x5598550e7a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #23: <unknown function> + 0x150866 (0x5598550fa866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x5598550e3142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #25: _PyFunction_Vectorcall + 0x6c (0x5598550eea2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #26: PyObject_Call + 0xbc (0x5598550faf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x5598550e12b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #28: _PyFunction_Vectorcall + 0x6c (0x5598550eea2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x5598550df8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #30: <unknown function> + 0x150582 (0x5598550fa582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x5598550df8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #32: <unknown function> + 0x150582 (0x5598550fa582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x5598550df8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #34: <unknown function> + 0x150582 (0x5598550fa582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x5598550df8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x5598550e6f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #37: _PyObject_Call_Prepend + 0x69 (0x5598550f8c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #38: <unknown function> + 0x211239 (0x5598551bb239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #39: _PyObject_MakeTpCall + 0x26b (0x5598550e7a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x5598550e33e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #41: _PyFunction_Vectorcall + 0x6c (0x5598550eea2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x5598550dec5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #43: _PyFunction_Vectorcall + 0x6c (0x5598550eea2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x5598550df8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #45: <unknown function> + 0x150582 (0x5598550fa582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #46: PyObject_Call + 0xbc (0x5598550faf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x5598550e12b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #48: <unknown function> + 0x150582 (0x5598550fa582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #49: PyObject_Call + 0xbc (0x5598550faf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x5598550e12b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #51: _PyFunction_Vectorcall + 0x6c (0x5598550eea2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x5598550e7007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #53: _PyObject_Call_Prepend + 0x69 (0x5598550f8c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #54: <unknown function> + 0x211239 (0x5598551bb239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #55: PyObject_Call + 0x207 (0x5598550fb067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x5598550e12b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #57: <unknown function> + 0x150582 (0x5598550fa582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x5598550df8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #59: <unknown function> + 0x150582 (0x5598550fa582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #60: PyObject_Call + 0xbc (0x5598550faf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x5598550e12b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #62: <unknown function> + 0x150582 (0x5598550fa582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: frame #63: PyObject_Call + 0xbc (0x5598550faf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default3]:[rank11]: . This may indicate a possible application crash on rank 0 or a network set up issue. |
|
[default1]:[rank9]: Traceback (most recent call last): |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module> |
|
[default1]:[rank9]: trainer.train(dataloader) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train |
|
[default1]:[rank9]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step |
|
[default1]:[rank9]: outputs = self.pipeline_engine.train_batch_iter( |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter |
|
[default1]:[rank9]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward |
|
[default1]:[rank9]: output = model(**micro_batch) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default1]:[rank9]: return self._call_impl(*args, **kwargs) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default1]:[rank9]: return forward_call(*args, **kwargs) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward |
|
[default1]:[rank9]: sharded_logits = self.model( |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default1]:[rank9]: return self._call_impl(*args, **kwargs) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default1]:[rank9]: return forward_call(*args, **kwargs) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward |
|
[default1]:[rank9]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0] |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states |
|
[default1]:[rank9]: hidden_encoder_states = encoder_block(**hidden_encoder_states) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default1]:[rank9]: return self._call_impl(*args, **kwargs) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default1]:[rank9]: return forward_call(*args, **kwargs) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward |
|
[default1]:[rank9]: new_kwargs[name] = recv_from_pipeline_state_buffer( |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer |
|
[default1]:[rank9]: pipeline_state.run_communication() |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication |
|
[default1]:[rank9]: recv_activation_tensor = recv_activation() |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__ |
|
[default1]:[rank9]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0] |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors |
|
[default1]:[rank9]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors |
|
[default1]:[rank9]: meta = self._recv_meta(from_rank=from_rank, tag=tag) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta |
|
[default1]:[rank9]: dist.recv( |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper |
|
[default1]:[rank9]: return func(*args, **kwargs) |
|
[default1]:[rank9]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv |
|
[default1]:[rank9]: pg.recv([tensor], group_src_rank, tag).wait() |
|
[default1]:[rank9]: torch.distributed.DistBackendError: [2] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '1:2', but store->get('1:2') got error: Connection reset by peer |
|
[default1]:[rank9]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first): |
|
[default1]:[rank9]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7f68506b8897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so) |
|
[default1]:[rank9]: frame #1: <unknown function> + 0x5b3a23e (0x7f688a1d523e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7f688a1cfc87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7f688a1cff82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7f688a1d0fd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f688a185371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f688a185371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f688a185371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f688a185371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7f6851992189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default1]:[rank9]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7f6851999610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default1]:[rank9]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7f68519b8978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default1]:[rank9]: frame #12: <unknown function> + 0x5adc309 (0x7f688a177309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #13: <unknown function> + 0x5ae6f10 (0x7f688a181f10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #14: <unknown function> + 0x5ae6fa5 (0x7f688a181fa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #15: <unknown function> + 0x5124446 (0x7f68897bf446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #16: <unknown function> + 0x1acf4b8 (0x7f688616a4b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #17: <unknown function> + 0x5aee004 (0x7f688a189004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #18: <unknown function> + 0x5af36b5 (0x7f688a18e6b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default1]:[rank9]: frame #19: <unknown function> + 0xd2631e (0x7f689cd7831e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default1]:[rank9]: frame #20: <unknown function> + 0x47def4 (0x7f689c4cfef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default1]:[rank9]: frame #21: <unknown function> + 0x1445a6 (0x55a657c7f5a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #22: _PyObject_MakeTpCall + 0x26b (0x55a657c78a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #23: <unknown function> + 0x150866 (0x55a657c8b866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x55a657c74142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #25: _PyFunction_Vectorcall + 0x6c (0x55a657c7fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #26: PyObject_Call + 0xbc (0x55a657c8bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x55a657c722b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #28: _PyFunction_Vectorcall + 0x6c (0x55a657c7fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x55a657c708fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #30: <unknown function> + 0x150582 (0x55a657c8b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x55a657c708fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #32: <unknown function> + 0x150582 (0x55a657c8b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x55a657c708fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #34: <unknown function> + 0x150582 (0x55a657c8b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x55a657c708fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x55a657c77f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #37: _PyObject_Call_Prepend + 0x69 (0x55a657c89c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #38: <unknown function> + 0x211239 (0x55a657d4c239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #39: _PyObject_MakeTpCall + 0x26b (0x55a657c78a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x55a657c743e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #41: _PyFunction_Vectorcall + 0x6c (0x55a657c7fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x55a657c6fc5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #43: _PyFunction_Vectorcall + 0x6c (0x55a657c7fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x55a657c708fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #45: <unknown function> + 0x150582 (0x55a657c8b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #46: PyObject_Call + 0xbc (0x55a657c8bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x55a657c722b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #48: <unknown function> + 0x150582 (0x55a657c8b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #49: PyObject_Call + 0xbc (0x55a657c8bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x55a657c722b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #51: _PyFunction_Vectorcall + 0x6c (0x55a657c7fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x55a657c78007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #53: _PyObject_Call_Prepend + 0x69 (0x55a657c89c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #54: <unknown function> + 0x211239 (0x55a657d4c239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #55: PyObject_Call + 0x207 (0x55a657c8c067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x55a657c722b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #57: <unknown function> + 0x150582 (0x55a657c8b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x55a657c708fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #59: <unknown function> + 0x150582 (0x55a657c8b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #60: PyObject_Call + 0xbc (0x55a657c8bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x55a657c722b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #62: <unknown function> + 0x150582 (0x55a657c8b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: frame #63: PyObject_Call + 0xbc (0x55a657c8bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default1]:[rank9]: . This may indicate a possible application crash on rank 0 or a network set up issue. |
|
[default6]:[rank14]: Traceback (most recent call last): |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module> |
|
[default6]:[rank14]: trainer.train(dataloader) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train |
|
[default6]:[rank14]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step |
|
[default6]:[rank14]: outputs = self.pipeline_engine.train_batch_iter( |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter |
|
[default6]:[rank14]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward |
|
[default6]:[rank14]: output = model(**micro_batch) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default6]:[rank14]: return self._call_impl(*args, **kwargs) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default6]:[rank14]: return forward_call(*args, **kwargs) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward |
|
[default6]:[rank14]: sharded_logits = self.model( |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default6]:[rank14]: return self._call_impl(*args, **kwargs) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default6]:[rank14]: return forward_call(*args, **kwargs) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward |
|
[default6]:[rank14]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0] |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states |
|
[default6]:[rank14]: hidden_encoder_states = encoder_block(**hidden_encoder_states) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default6]:[rank14]: return self._call_impl(*args, **kwargs) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default6]:[rank14]: return forward_call(*args, **kwargs) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward |
|
[default6]:[rank14]: new_kwargs[name] = recv_from_pipeline_state_buffer( |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer |
|
[default6]:[rank14]: pipeline_state.run_communication() |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication |
|
[default6]:[rank14]: recv_activation_tensor = recv_activation() |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__ |
|
[default6]:[rank14]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0] |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors |
|
[default6]:[rank14]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors |
|
[default6]:[rank14]: meta = self._recv_meta(from_rank=from_rank, tag=tag) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta |
|
[default6]:[rank14]: dist.recv( |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper |
|
[default6]:[rank14]: return func(*args, **kwargs) |
|
[default6]:[rank14]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv |
|
[default6]:[rank14]: pg.recv([tensor], group_src_rank, tag).wait() |
|
[default6]:[rank14]: torch.distributed.DistBackendError: [3] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '2:3', but store->get('2:3') got error: Connection reset by peer |
|
[default6]:[rank14]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first): |
|
[default6]:[rank14]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7f46d4075897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so) |
|
[default6]:[rank14]: frame #1: <unknown function> + 0x5b3a23e (0x7f470db9223e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7f470db8cc87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7f470db8cf82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7f470db8dfd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f470db42371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f470db42371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f470db42371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f470db42371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7f46d534f189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default6]:[rank14]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7f46d5356610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default6]:[rank14]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7f46d5375978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default6]:[rank14]: frame #12: <unknown function> + 0x5adc309 (0x7f470db34309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #13: <unknown function> + 0x5ae6f10 (0x7f470db3ef10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #14: <unknown function> + 0x5ae6fa5 (0x7f470db3efa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #15: <unknown function> + 0x5124446 (0x7f470d17c446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #16: <unknown function> + 0x1acf4b8 (0x7f4709b274b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #17: <unknown function> + 0x5aee004 (0x7f470db46004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #18: <unknown function> + 0x5af36b5 (0x7f470db4b6b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default6]:[rank14]: frame #19: <unknown function> + 0xd2631e (0x7f472073531e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default6]:[rank14]: frame #20: <unknown function> + 0x47def4 (0x7f471fe8cef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default6]:[rank14]: frame #21: <unknown function> + 0x1445a6 (0x55db164415a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #22: _PyObject_MakeTpCall + 0x26b (0x55db1643aa6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #23: <unknown function> + 0x150866 (0x55db1644d866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x55db16436142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #25: _PyFunction_Vectorcall + 0x6c (0x55db16441a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #26: PyObject_Call + 0xbc (0x55db1644df1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x55db164342b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #28: _PyFunction_Vectorcall + 0x6c (0x55db16441a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x55db164328fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #30: <unknown function> + 0x150582 (0x55db1644d582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x55db164328fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #32: <unknown function> + 0x150582 (0x55db1644d582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x55db164328fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #34: <unknown function> + 0x150582 (0x55db1644d582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x55db164328fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x55db16439f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #37: _PyObject_Call_Prepend + 0x69 (0x55db1644bc39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #38: <unknown function> + 0x211239 (0x55db1650e239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #39: _PyObject_MakeTpCall + 0x26b (0x55db1643aa6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x55db164363e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #41: _PyFunction_Vectorcall + 0x6c (0x55db16441a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x55db16431c5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #43: _PyFunction_Vectorcall + 0x6c (0x55db16441a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x55db164328fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #45: <unknown function> + 0x150582 (0x55db1644d582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #46: PyObject_Call + 0xbc (0x55db1644df1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x55db164342b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #48: <unknown function> + 0x150582 (0x55db1644d582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #49: PyObject_Call + 0xbc (0x55db1644df1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x55db164342b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #51: _PyFunction_Vectorcall + 0x6c (0x55db16441a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x55db1643a007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #53: _PyObject_Call_Prepend + 0x69 (0x55db1644bc39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #54: <unknown function> + 0x211239 (0x55db1650e239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #55: PyObject_Call + 0x207 (0x55db1644e067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x55db164342b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #57: <unknown function> + 0x150582 (0x55db1644d582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x55db164328fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #59: <unknown function> + 0x150582 (0x55db1644d582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #60: PyObject_Call + 0xbc (0x55db1644df1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x55db164342b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #62: <unknown function> + 0x150582 (0x55db1644d582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: frame #63: PyObject_Call + 0xbc (0x55db1644df1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default6]:[rank14]: . This may indicate a possible application crash on rank 0 or a network set up issue. |
|
[default5]:[rank13]: Traceback (most recent call last): |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module> |
|
[default5]:[rank13]: trainer.train(dataloader) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train |
|
[default5]:[rank13]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step |
|
[default5]:[rank13]: outputs = self.pipeline_engine.train_batch_iter( |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter |
|
[default5]:[rank13]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward |
|
[default5]:[rank13]: output = model(**micro_batch) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default5]:[rank13]: return self._call_impl(*args, **kwargs) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default5]:[rank13]: return forward_call(*args, **kwargs) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward |
|
[default5]:[rank13]: sharded_logits = self.model( |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default5]:[rank13]: return self._call_impl(*args, **kwargs) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default5]:[rank13]: return forward_call(*args, **kwargs) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward |
|
[default5]:[rank13]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0] |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states |
|
[default5]:[rank13]: hidden_encoder_states = encoder_block(**hidden_encoder_states) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default5]:[rank13]: return self._call_impl(*args, **kwargs) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default5]:[rank13]: return forward_call(*args, **kwargs) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward |
|
[default5]:[rank13]: new_kwargs[name] = recv_from_pipeline_state_buffer( |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer |
|
[default5]:[rank13]: pipeline_state.run_communication() |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication |
|
[default5]:[rank13]: recv_activation_tensor = recv_activation() |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__ |
|
[default5]:[rank13]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0] |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors |
|
[default5]:[rank13]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors |
|
[default5]:[rank13]: meta = self._recv_meta(from_rank=from_rank, tag=tag) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta |
|
[default5]:[rank13]: dist.recv( |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper |
|
[default5]:[rank13]: return func(*args, **kwargs) |
|
[default5]:[rank13]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv |
|
[default5]:[rank13]: pg.recv([tensor], group_src_rank, tag).wait() |
|
[default5]:[rank13]: torch.distributed.DistBackendError: [3] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '2:3', but store->get('2:3') got error: Connection reset by peer |
|
[default5]:[rank13]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first): |
|
[default5]:[rank13]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7f16cf2c6897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so) |
|
[default5]:[rank13]: frame #1: <unknown function> + 0x5b3a23e (0x7f1708de323e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7f1708dddc87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7f1708dddf82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7f1708ddefd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f1708d93371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f1708d93371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f1708d93371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7f1708d93371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7f16d05a0189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default5]:[rank13]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7f16d05a7610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default5]:[rank13]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7f16d05c6978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default5]:[rank13]: frame #12: <unknown function> + 0x5adc309 (0x7f1708d85309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #13: <unknown function> + 0x5ae6f10 (0x7f1708d8ff10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #14: <unknown function> + 0x5ae6fa5 (0x7f1708d8ffa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #15: <unknown function> + 0x5124446 (0x7f17083cd446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #16: <unknown function> + 0x1acf4b8 (0x7f1704d784b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #17: <unknown function> + 0x5aee004 (0x7f1708d97004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #18: <unknown function> + 0x5af36b5 (0x7f1708d9c6b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default5]:[rank13]: frame #19: <unknown function> + 0xd2631e (0x7f171b98631e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default5]:[rank13]: frame #20: <unknown function> + 0x47def4 (0x7f171b0ddef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default5]:[rank13]: frame #21: <unknown function> + 0x1445a6 (0x560ecbdc55a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #22: _PyObject_MakeTpCall + 0x26b (0x560ecbdbea6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #23: <unknown function> + 0x150866 (0x560ecbdd1866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x560ecbdba142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #25: _PyFunction_Vectorcall + 0x6c (0x560ecbdc5a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #26: PyObject_Call + 0xbc (0x560ecbdd1f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x560ecbdb82b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #28: _PyFunction_Vectorcall + 0x6c (0x560ecbdc5a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x560ecbdb68fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #30: <unknown function> + 0x150582 (0x560ecbdd1582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x560ecbdb68fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #32: <unknown function> + 0x150582 (0x560ecbdd1582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x560ecbdb68fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #34: <unknown function> + 0x150582 (0x560ecbdd1582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x560ecbdb68fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x560ecbdbdf50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #37: _PyObject_Call_Prepend + 0x69 (0x560ecbdcfc39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #38: <unknown function> + 0x211239 (0x560ecbe92239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #39: _PyObject_MakeTpCall + 0x26b (0x560ecbdbea6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x560ecbdba3e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #41: _PyFunction_Vectorcall + 0x6c (0x560ecbdc5a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x560ecbdb5c5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #43: _PyFunction_Vectorcall + 0x6c (0x560ecbdc5a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x560ecbdb68fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #45: <unknown function> + 0x150582 (0x560ecbdd1582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #46: PyObject_Call + 0xbc (0x560ecbdd1f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x560ecbdb82b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #48: <unknown function> + 0x150582 (0x560ecbdd1582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #49: PyObject_Call + 0xbc (0x560ecbdd1f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x560ecbdb82b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #51: _PyFunction_Vectorcall + 0x6c (0x560ecbdc5a2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x560ecbdbe007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #53: _PyObject_Call_Prepend + 0x69 (0x560ecbdcfc39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #54: <unknown function> + 0x211239 (0x560ecbe92239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #55: PyObject_Call + 0x207 (0x560ecbdd2067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x560ecbdb82b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #57: <unknown function> + 0x150582 (0x560ecbdd1582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x560ecbdb68fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #59: <unknown function> + 0x150582 (0x560ecbdd1582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #60: PyObject_Call + 0xbc (0x560ecbdd1f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x560ecbdb82b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #62: <unknown function> + 0x150582 (0x560ecbdd1582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: frame #63: PyObject_Call + 0xbc (0x560ecbdd1f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default5]:[rank13]: . This may indicate a possible application crash on rank 0 or a network set up issue. |
|
[default2]:[rank10]: Traceback (most recent call last): |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module> |
|
[default2]:[rank10]: trainer.train(dataloader) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train |
|
[default2]:[rank10]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step |
|
[default2]:[rank10]: outputs = self.pipeline_engine.train_batch_iter( |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 252, in train_batch_iter |
|
[default2]:[rank10]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward |
|
[default2]:[rank10]: output = model(**micro_batch) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default2]:[rank10]: return self._call_impl(*args, **kwargs) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default2]:[rank10]: return forward_call(*args, **kwargs) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward |
|
[default2]:[rank10]: sharded_logits = self.model( |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default2]:[rank10]: return self._call_impl(*args, **kwargs) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default2]:[rank10]: return forward_call(*args, **kwargs) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward |
|
[default2]:[rank10]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0] |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states |
|
[default2]:[rank10]: hidden_encoder_states = encoder_block(**hidden_encoder_states) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default2]:[rank10]: return self._call_impl(*args, **kwargs) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default2]:[rank10]: return forward_call(*args, **kwargs) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward |
|
[default2]:[rank10]: new_kwargs[name] = recv_from_pipeline_state_buffer( |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer |
|
[default2]:[rank10]: pipeline_state.run_communication() |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication |
|
[default2]:[rank10]: recv_activation_tensor = recv_activation() |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__ |
|
[default2]:[rank10]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0] |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors |
|
[default2]:[rank10]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors |
|
[default2]:[rank10]: meta = self._recv_meta(from_rank=from_rank, tag=tag) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta |
|
[default2]:[rank10]: dist.recv( |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper |
|
[default2]:[rank10]: return func(*args, **kwargs) |
|
[default2]:[rank10]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv |
|
[default2]:[rank10]: pg.recv([tensor], group_src_rank, tag).wait() |
|
[default2]:[rank10]: torch.distributed.DistBackendError: [2] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '1:2', but store->get('1:2') got error: Connection reset by peer |
|
[default2]:[rank10]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first): |
|
[default2]:[rank10]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7fc92087d897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so) |
|
[default2]:[rank10]: frame #1: <unknown function> + 0x5b3a23e (0x7fc95a39a23e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7fc95a394c87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7fc95a394f82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7fc95a395fd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fc95a34a371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fc95a34a371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fc95a34a371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fc95a34a371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7fc921b57189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default2]:[rank10]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7fc921b5e610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default2]:[rank10]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7fc921b7d978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default2]:[rank10]: frame #12: <unknown function> + 0x5adc309 (0x7fc95a33c309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #13: <unknown function> + 0x5ae6f10 (0x7fc95a346f10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #14: <unknown function> + 0x5ae6fa5 (0x7fc95a346fa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #15: <unknown function> + 0x5124446 (0x7fc959984446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #16: <unknown function> + 0x1acf4b8 (0x7fc95632f4b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #17: <unknown function> + 0x5aee004 (0x7fc95a34e004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #18: <unknown function> + 0x5af36b5 (0x7fc95a3536b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default2]:[rank10]: frame #19: <unknown function> + 0xd2631e (0x7fc96cf3d31e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default2]:[rank10]: frame #20: <unknown function> + 0x47def4 (0x7fc96c694ef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default2]:[rank10]: frame #21: <unknown function> + 0x1445a6 (0x557ca950a5a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #22: _PyObject_MakeTpCall + 0x26b (0x557ca9503a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #23: <unknown function> + 0x150866 (0x557ca9516866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x557ca94ff142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #25: _PyFunction_Vectorcall + 0x6c (0x557ca950aa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #26: PyObject_Call + 0xbc (0x557ca9516f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x557ca94fd2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #28: _PyFunction_Vectorcall + 0x6c (0x557ca950aa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x557ca94fb8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #30: <unknown function> + 0x150582 (0x557ca9516582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x557ca94fb8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #32: <unknown function> + 0x150582 (0x557ca9516582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x557ca94fb8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #34: <unknown function> + 0x150582 (0x557ca9516582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x557ca94fb8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x557ca9502f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #37: _PyObject_Call_Prepend + 0x69 (0x557ca9514c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #38: <unknown function> + 0x211239 (0x557ca95d7239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #39: _PyObject_MakeTpCall + 0x26b (0x557ca9503a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x557ca94ff3e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #41: _PyFunction_Vectorcall + 0x6c (0x557ca950aa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x557ca94fac5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #43: _PyFunction_Vectorcall + 0x6c (0x557ca950aa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x557ca94fb8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #45: <unknown function> + 0x150582 (0x557ca9516582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #46: PyObject_Call + 0xbc (0x557ca9516f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x557ca94fd2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #48: <unknown function> + 0x150582 (0x557ca9516582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #49: PyObject_Call + 0xbc (0x557ca9516f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x557ca94fd2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #51: _PyFunction_Vectorcall + 0x6c (0x557ca950aa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x557ca9503007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #53: _PyObject_Call_Prepend + 0x69 (0x557ca9514c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #54: <unknown function> + 0x211239 (0x557ca95d7239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #55: PyObject_Call + 0x207 (0x557ca9517067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x557ca94fd2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #57: <unknown function> + 0x150582 (0x557ca9516582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x557ca94fb8fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #59: <unknown function> + 0x150582 (0x557ca9516582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #60: PyObject_Call + 0xbc (0x557ca9516f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x557ca94fd2b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #62: <unknown function> + 0x150582 (0x557ca9516582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: frame #63: PyObject_Call + 0xbc (0x557ca9516f1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default2]:[rank10]: . This may indicate a possible application crash on rank 0 or a network set up issue. |
|
[default4]:[rank12]: Traceback (most recent call last): |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py", line 237, in <module> |
|
[default4]:[rank12]: trainer.train(dataloader) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 429, in train |
|
[default4]:[rank12]: outputs, loss_avg = self.training_step(dataloader=self.current_dataloader) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/trainer.py", line 462, in training_step |
|
[default4]:[rank12]: outputs = self.pipeline_engine.train_batch_iter( |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 278, in train_batch_iter |
|
[default4]:[rank12]: output = self.forward(context=context, state=state, micro_batch=micro_batch, model=model) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/engine.py", line 44, in forward |
|
[default4]:[rank12]: output = model(**micro_batch) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default4]:[rank12]: return self._call_impl(*args, **kwargs) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default4]:[rank12]: return forward_call(*args, **kwargs) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 891, in forward |
|
[default4]:[rank12]: sharded_logits = self.model( |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default4]:[rank12]: return self._call_impl(*args, **kwargs) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default4]:[rank12]: return forward_call(*args, **kwargs) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 764, in forward |
|
[default4]:[rank12]: return self.forward_with_hidden_states(input_ids=input_ids, input_mask=input_mask)[0] |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/models/llama.py", line 780, in forward_with_hidden_states |
|
[default4]:[rank12]: hidden_encoder_states = encoder_block(**hidden_encoder_states) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1532, in _wrapped_call_impl |
|
[default4]:[rank12]: return self._call_impl(*args, **kwargs) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/nn/modules/module.py", line 1541, in _call_impl |
|
[default4]:[rank12]: return forward_call(*args, **kwargs) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/block.py", line 126, in forward |
|
[default4]:[rank12]: new_kwargs[name] = recv_from_pipeline_state_buffer( |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/functional.py", line 117, in recv_from_pipeline_state_buffer |
|
[default4]:[rank12]: pipeline_state.run_communication() |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 150, in run_communication |
|
[default4]:[rank12]: recv_activation_tensor = recv_activation() |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/state.py", line 31, in __call__ |
|
[default4]:[rank12]: return self.p2p.recv_tensors(num_tensors=1, from_rank=self.from_rank)[0] |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 353, in recv_tensors |
|
[default4]:[rank12]: buffers, futures = self.irecv_tensors(num_tensors=num_tensors, from_rank=from_rank, tag=tag) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 326, in irecv_tensors |
|
[default4]:[rank12]: meta = self._recv_meta(from_rank=from_rank, tag=tag) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/src/nanotron/parallel/pipeline_parallel/p2p.py", line 246, in _recv_meta |
|
[default4]:[rank12]: dist.recv( |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/c10d_logger.py", line 75, in wrapper |
|
[default4]:[rank12]: return func(*args, **kwargs) |
|
[default4]:[rank12]: File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/distributed_c10d.py", line 1932, in recv |
|
[default4]:[rank12]: pg.recv([tensor], group_src_rank, tag).wait() |
|
[default4]:[rank12]: torch.distributed.DistBackendError: [3] is setting up NCCL communicator and retrieving ncclUniqueId from [0] via c10d key-value store by key '2:3', but store->get('2:3') got error: Connection reset by peer |
|
[default4]:[rank12]: Exception raised from recvBytes at ../torch/csrc/distributed/c10d/Utils.hpp:672 (most recent call first): |
|
[default4]:[rank12]: frame #0: c10::Error::Error(c10::SourceLocation, std::string) + 0x57 (0x7fd91bbd1897 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libc10.so) |
|
[default4]:[rank12]: frame #1: <unknown function> + 0x5b3a23e (0x7fd9556ee23e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #2: c10d::TCPStore::doWait(c10::ArrayRef<std::string>, std::chrono::duration<long, std::ratio<1l, 1000l> >) + 0x2c7 (0x7fd9556e8c87 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #3: c10d::TCPStore::doGet(std::string const&) + 0x32 (0x7fd9556e8f82 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #4: c10d::TCPStore::get(std::string const&) + 0xa1 (0x7fd9556e9fd1 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #5: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fd95569e371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #6: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fd95569e371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #7: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fd95569e371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #8: c10d::PrefixStore::get(std::string const&) + 0x31 (0x7fd95569e371 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #9: c10d::ProcessGroupNCCL::broadcastUniqueNCCLID(ncclUniqueId*, bool, std::string const&, int) + 0xa9 (0x7fd91ceab189 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default4]:[rank12]: frame #10: c10d::ProcessGroupNCCL::getNCCLComm(std::string const&, c10::Device&, c10d::OpType, int, bool) + 0xc50 (0x7fd91ceb2610 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default4]:[rank12]: frame #11: c10d::ProcessGroupNCCL::recv(std::vector<at::Tensor, std::allocator<at::Tensor> >&, int, int) + 0x5f8 (0x7fd91ced1978 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cuda.so) |
|
[default4]:[rank12]: frame #12: <unknown function> + 0x5adc309 (0x7fd955690309 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #13: <unknown function> + 0x5ae6f10 (0x7fd95569af10 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #14: <unknown function> + 0x5ae6fa5 (0x7fd95569afa5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #15: <unknown function> + 0x5124446 (0x7fd954cd8446 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #16: <unknown function> + 0x1acf4b8 (0x7fd9516834b8 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #17: <unknown function> + 0x5aee004 (0x7fd9556a2004 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #18: <unknown function> + 0x5af36b5 (0x7fd9556a76b5 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_cpu.so) |
|
[default4]:[rank12]: frame #19: <unknown function> + 0xd2631e (0x7fd96829131e in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default4]:[rank12]: frame #20: <unknown function> + 0x47def4 (0x7fd9679e8ef4 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/lib/libtorch_python.so) |
|
[default4]:[rank12]: frame #21: <unknown function> + 0x1445a6 (0x564458a8f5a6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #22: _PyObject_MakeTpCall + 0x26b (0x564458a88a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #23: <unknown function> + 0x150866 (0x564458a9b866 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #24: _PyEval_EvalFrameDefault + 0x4c12 (0x564458a84142 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #25: _PyFunction_Vectorcall + 0x6c (0x564458a8fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #26: PyObject_Call + 0xbc (0x564458a9bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #27: _PyEval_EvalFrameDefault + 0x2d83 (0x564458a822b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #28: _PyFunction_Vectorcall + 0x6c (0x564458a8fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #29: _PyEval_EvalFrameDefault + 0x13ca (0x564458a808fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #30: <unknown function> + 0x150582 (0x564458a9b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #31: _PyEval_EvalFrameDefault + 0x13ca (0x564458a808fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #32: <unknown function> + 0x150582 (0x564458a9b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #33: _PyEval_EvalFrameDefault + 0x13ca (0x564458a808fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #34: <unknown function> + 0x150582 (0x564458a9b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #35: _PyEval_EvalFrameDefault + 0x13ca (0x564458a808fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #36: _PyObject_FastCallDictTstate + 0xd0 (0x564458a87f50 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #37: _PyObject_Call_Prepend + 0x69 (0x564458a99c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #38: <unknown function> + 0x211239 (0x564458b5c239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #39: _PyObject_MakeTpCall + 0x26b (0x564458a88a6b in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #40: _PyEval_EvalFrameDefault + 0x4eb6 (0x564458a843e6 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #41: _PyFunction_Vectorcall + 0x6c (0x564458a8fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #42: _PyEval_EvalFrameDefault + 0x72c (0x564458a7fc5c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #43: _PyFunction_Vectorcall + 0x6c (0x564458a8fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #44: _PyEval_EvalFrameDefault + 0x13ca (0x564458a808fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #45: <unknown function> + 0x150582 (0x564458a9b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #46: PyObject_Call + 0xbc (0x564458a9bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #47: _PyEval_EvalFrameDefault + 0x2d83 (0x564458a822b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #48: <unknown function> + 0x150582 (0x564458a9b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #49: PyObject_Call + 0xbc (0x564458a9bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #50: _PyEval_EvalFrameDefault + 0x2d83 (0x564458a822b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #51: _PyFunction_Vectorcall + 0x6c (0x564458a8fa2c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #52: _PyObject_FastCallDictTstate + 0x187 (0x564458a88007 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #53: _PyObject_Call_Prepend + 0x69 (0x564458a99c39 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #54: <unknown function> + 0x211239 (0x564458b5c239 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #55: PyObject_Call + 0x207 (0x564458a9c067 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #56: _PyEval_EvalFrameDefault + 0x2d83 (0x564458a822b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #57: <unknown function> + 0x150582 (0x564458a9b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #58: _PyEval_EvalFrameDefault + 0x13ca (0x564458a808fa in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #59: <unknown function> + 0x150582 (0x564458a9b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #60: PyObject_Call + 0xbc (0x564458a9bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #61: _PyEval_EvalFrameDefault + 0x2d83 (0x564458a822b3 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #62: <unknown function> + 0x150582 (0x564458a9b582 in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: frame #63: PyObject_Call + 0xbc (0x564458a9bf1c in /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10) |
|
[default4]:[rank12]: . This may indicate a possible application crash on rank 0 or a network set up issue. |
|
E0702 16:30:54.049000 140670700869440 torch/distributed/elastic/multiprocessing/api.py:826] failed (exitcode: 1) local_rank: 1 (pid: 1371215) of binary: /fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/python3.10 |
|
Traceback (most recent call last): |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/torchrun", line 8, in <module> |
|
sys.exit(main()) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 347, in wrapper |
|
return f(*args, **kwargs) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/run.py", line 879, in main |
|
run(args) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/run.py", line 870, in run |
|
elastic_launch( |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 132, in __call__ |
|
return launch_agent(self._config, self._entrypoint, list(args)) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 263, in launch_agent |
|
raise ChildFailedError( |
|
torch.distributed.elastic.multiprocessing.errors.ChildFailedError: |
|
============================================================ |
|
/fsx/ferdinandmom/ferdinand-hf/bench_cluster/nanotron/run_train.py FAILED |
|
------------------------------------------------------------ |
|
Failures: |
|
<NO_OTHER_FAILURES> |
|
------------------------------------------------------------ |
|
Root Cause (first observed failure): |
|
[0]: |
|
time : 2024-07-02_16:30:51 |
|
host : ip-26-0-160-225.ec2.internal |
|
rank : 1 (local_rank: 1) |
|
exitcode : 1 (pid: 1371215) |
|
error_file: <N/A> |
|
traceback : To enable traceback see: https://pytorch.org/docs/stable/elastic/errors.html |
|
============================================================ |
|
srun: error: ip-26-0-160-225: task 0: Exited with exit code 1 |
|
W0702 16:30:55.547000 140350647297792 torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1252] The node 'ip-26-0-171-56.ec2.internal_3276546_0' has failed to send a keep-alive heartbeat to the rendezvous 'none' due to an error of type RendezvousConnectionError. |
|
W0702 16:30:56.530000 140356314117952 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 3276696 closing signal SIGTERM |
|
W0702 16:30:56.530000 140356314117952 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 3276697 closing signal SIGTERM |
|
W0702 16:30:56.531000 140356314117952 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 3276698 closing signal SIGTERM |
|
W0702 16:30:56.531000 140356314117952 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 3276699 closing signal SIGTERM |
|
W0702 16:30:56.532000 140356314117952 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 3276700 closing signal SIGTERM |
|
W0702 16:30:56.532000 140356314117952 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 3276701 closing signal SIGTERM |
|
W0702 16:30:56.533000 140356314117952 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 3276702 closing signal SIGTERM |
|
W0702 16:30:56.533000 140356314117952 torch/distributed/elastic/multiprocessing/api.py:851] Sending process 3276703 closing signal SIGTERM |
|
W0702 16:30:58.353000 140356314117952 torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1203] The node 'ip-26-0-171-56.ec2.internal_3276546_0' has failed to shutdown the rendezvous 'none' due to an error of type RendezvousConnectionError. |
|
W0702 16:30:58.362000 140356314117952 torch/distributed/elastic/rendezvous/dynamic_rendezvous.py:1203] The node 'ip-26-0-171-56.ec2.internal_3276546_0' has failed to shutdown the rendezvous 'none' due to an error of type RendezvousConnectionError. |
|
Traceback (most recent call last): |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py", line 113, in _call_store |
|
return getattr(self._store, store_op)(*args, **kwargs) |
|
torch.distributed.DistNetworkError: Broken pipe |
|
|
|
The above exception was the direct cause of the following exception: |
|
|
|
Traceback (most recent call last): |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/bin/torchrun", line 8, in <module> |
|
sys.exit(main()) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/multiprocessing/errors/__init__.py", line 347, in wrapper |
|
return f(*args, **kwargs) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/run.py", line 879, in main |
|
run(args) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/run.py", line 870, in run |
|
elastic_launch( |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 132, in __call__ |
|
return launch_agent(self._config, self._entrypoint, list(args)) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/launcher/api.py", line 254, in launch_agent |
|
result = agent.run() |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/metrics/api.py", line 123, in wrapper |
|
result = f(*args, **kwargs) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/agent/server/api.py", line 733, in run |
|
result = self._invoke_run(role) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/agent/server/api.py", line 908, in _invoke_run |
|
num_nodes_waiting = rdzv_handler.num_nodes_waiting() |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py", line 1174, in num_nodes_waiting |
|
self._state_holder.sync() |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/rendezvous/dynamic_rendezvous.py", line 419, in sync |
|
get_response = self._backend.get_state() |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py", line 73, in get_state |
|
base64_state: bytes = self._call_store("get", self._key) |
|
File "/fsx/ferdinandmom/miniforge3/envs/env-bench-cluster/lib/python3.10/site-packages/torch/distributed/elastic/rendezvous/c10d_rendezvous_backend.py", line 115, in _call_store |
|
raise RendezvousConnectionError( |
|
torch.distributed.elastic.rendezvous.api.RendezvousConnectionError: The connection to the C10d store has failed. See inner exception for details. |
|
srun: error: ip-26-0-171-56: task 1: Exited with exit code 1 |
|
Consider using `hf_transfer` for faster uploads. This solution comes with some limitations. See https://huggingface.co/docs/huggingface_hub/hf_transfer for more details. |
|
|