|
|
| The following have been reloaded with a version change: |
| 1) GCCcore/.14.3.0 => GCCcore/14.3.0 |
|
|
|
|
| Lmod is automatically replacing "GCC/14.3.0" with |
| "nvidia-compilers/25.9-CUDA-13". |
|
|
| Deactivating conda environment: /e/scratch/jureap59/feuer1/miniforge3/envs/otagent |
| Activating RL environment: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl |
| Python executable: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python |
| Python path check: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python |
| [ray] RAY_TMPDIR=/tmp/ray/ray_388832 |
| [triton_cache] Triton cache: /tmp/triton_cache_feuer1_388832 |
| [triton_cache] TorchInductor cache: /tmp/torchinductor_cache_feuer1_388832 |
| [proxy] ✓ Found proxychains binary at /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 |
| [proxy] Setting up SSH tunnel to jpbl-s01-01 |
| [proxy] SSH key: /e/home/jusers/feuer1/jupiter/.ssh/authorized_keys/id_ed25519_jsc |
| [proxy] Tunnel port: 7003 |
| [proxy] Node IP: 10.128.33.33 (workers will connect here) |
| [proxy] ✓ SSH tunnel started successfully |
| [proxy] ✓ Generated proxychains config at /e/home/jusers/feuer1/jupiter/.proxychains/proxychains_388832.conf |
| [proxy] - Internal traffic (10.x.x.x, 172.x.x.x, 169.254.x.x) → DIRECT |
| [proxy] - External traffic (internet) → PROXY via tunnel |
| [proxy] ✓ Daytona timeout settings configured |
| [proxy] Testing proxy connectivity... |
| [proxychains] config file found: /e/home/jusers/feuer1/jupiter/.proxychains/proxychains_388832.conf |
| [proxychains] preloading /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/lib/libproxychains4.so |
| [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e |
| [proxy] ✓ Proxy connectivity test passed (huggingface.co reachable via wrapped binary) |
| [proxy] ⚠ Tunnel not accessible at 10.128.33.33:7003 (workers may fail) |
| [proxy] ✓ Proxy setup complete (using wrapped binary for Ray workers) |
| [container_runtime] Using cloud backend: daytona (no local container setup) |
| === Universal RL Training Runner === |
| Config: /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/configs/a2-rl-stack_jest_v2_rl_config.json |
| Working directory: /e/scratch/jureap59/feuer1/OpenThoughts-Agent |
| Python: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python |
| Python version: Python 3.12.12 |
| Proxy: DISABLED (direct internet or not configured) |
| ======================================== |
| === RLJobRunner: a2-rl-stack_jest_v2 === |
| [wandb_utils] Fixing permissions on: /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/wandb |
| [wandb_utils] WandB directory ready: /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/wandb |
| HF_TOKEN=****pDbg |
| HF_HUB_CACHE=/e/data1/datasets/playground/ot-baf/hf_hub |
| SUPABASE_URL=https: |
| Environment configured: |
| TENSOR_PARALLEL_SIZE=2 |
| NUM_INFERENCE_ENGINES=32 |
| POLICY_NUM_NODES=16 |
| WANDB_DIR=/e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/wandb |
| Starting Ray cluster with 16 nodes, 4 GPUs/node |
| Cleaning up existing Ray instances... |
| === Starting Ray Cluster === |
| Nodes: 16 |
| GPUs per node: 4 |
| CPUs per node: 288 |
| Head node: jpbo-047-01 (10.128.33.33) |
| Ray port: 6379 |
| ============================ |
| Starting Ray head on jpbo-047-01 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_head_jpbo-047-01.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.33 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-01 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --head --node-ip-address=10.128.33.33 --port=6379 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray head on jpbo-047-01 |
| Starting Ray worker on jpbo-047-02 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-02.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.34 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-02 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.34 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 1 on jpbo-047-02 |
| Starting Ray worker on jpbo-047-03 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-03.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.35 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-03 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.35 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 2 on jpbo-047-03 |
| Starting Ray worker on jpbo-047-04 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-04.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.36 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-04 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.36 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 3 on jpbo-047-04 |
| Starting Ray worker on jpbo-047-05 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-05.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.37 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-05 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.37 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 4 on jpbo-047-05 |
| Starting Ray worker on jpbo-047-06 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-06.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.38 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-06 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.38 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 5 on jpbo-047-06 |
| Starting Ray worker on jpbo-047-07 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-07.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.39 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-07 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.39 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 6 on jpbo-047-07 |
| Starting Ray worker on jpbo-047-08 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-08.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.40 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-08 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.40 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 7 on jpbo-047-08 |
| Starting Ray worker on jpbo-047-09 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-09.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.41 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-09 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.41 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 8 on jpbo-047-09 |
| Starting Ray worker on jpbo-047-10 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-10.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.42 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-10 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.42 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 9 on jpbo-047-10 |
| Starting Ray worker on jpbo-047-11 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-11.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.43 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-11 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.43 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 10 on jpbo-047-11 |
| Starting Ray worker on jpbo-047-12 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-12.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.44 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-12 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.44 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 11 on jpbo-047-12 |
| Starting Ray worker on jpbo-047-13 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-13.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.45 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-13 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.45 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 12 on jpbo-047-13 |
| Starting Ray worker on jpbo-047-14 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-14.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.46 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-14 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.46 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 13 on jpbo-047-14 |
| Starting Ray worker on jpbo-047-15 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-15.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.47 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-15 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.47 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 14 on jpbo-047-15 |
| Starting Ray worker on jpbo-047-16 (logging to /e/scratch/jureap59/feuer1/OpenThoughts-Agent/experiments/logs/ray_worker_jpbo-047-16.log)... |
| Command: srun --export=ALL,VLLM_HOST_IP=10.128.33.48 --nodes=1 --ntasks=1 --gres=gpu:4 --gpu-bind=none --overlap --cpu-bind=none -w jpbo-047-16 bash -c /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f "$PROXYCHAINS_CONF_FILE" ray start --address=10.128.33.33:6379 --node-ip-address=10.128.33.48 --num-gpus=4 --num-cpus=288 --block --object-store-memory=42949672960 |
| Started Ray worker 15 on jpbo-047-16 |
| Waiting for cluster (64 GPUs, 16 nodes)... |
| Connecting to Ray at 10.128.33.33:6379 (expecting 16 nodes, 64.0 GPUs) |
| Ray connection established, polling for resources... |
| [Ray wait] nodes=16/16 GPUs=64.0/64.0 resources={'GPU': 64.0, 'node:10.128.33.47': 1.0, 'object_store_memory': 687194767360.0, 'accelerator_type:GH200': 16.0, 'memory': 11834749616128.0, 'CPU': 4608.0, 'node:10.128.33.45': 1.0, 'node:10.128.33.33': 1.0, 'node:__internal_head__': 1.0, 'node:10.128.33.34': 1.0, 'node:10.128.33.37': 1.0, 'node:10.128.33.46': 1.0, 'node:10.128.33.38': 1.0, 'node:10.128.33.35': 1.0, 'node:10.128.33.41': 1.0, 'node:10.128.33.48': 1.0, 'node:10.128.33.36': 1.0, 'node:10.128.33.39': 1.0, 'node:10.128.33.42': 1.0, 'node:10.128.33.43': 1.0, 'node:10.128.33.40': 1.0, 'node:10.128.33.44': 1.0} |
| ✓ Ray cluster ready |
| === Ray Cluster Ready === |
| Address: 10.128.33.33:6379 |
| Total GPUs: 64 |
| ========================= |
| Ray cluster ready at 10.128.33.33:6379 |
| Total GPUs available: 64 |
| [RLJobRunner] Pinggy check: url=False, token=False, needs_tunnel=False (agent=terminus-2, env=daytona) |
| [RLJobRunner] No Pinggy tunnel needed, using local vLLM |
|
|
| Running SkyRL: |
| Python: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python |
| Entrypoint: examples.terminal_bench.entrypoints.main_tbench |
| Args: 119 Hydra arguments |
| Working dir: /e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train |
| Using proxychains binary: /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 |
|
|
| Executing command with srun: /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/bin/proxychains4 -f $PROXYCHAINS_CONF_FILE /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/bin/python -m examples.terminal_bench.entrypoints.main_tbench +terminal_bench_config=terminal_bench trainer.strategy=fsdp2 trainer.algorithm.advantage_estimator=rloo_n trainer.algorithm.use_kl_loss=false trainer.algorithm.kl_loss_coef=0.0 trainer.algorithm.eps_clip_low=0.2 trainer.algorithm.eps_clip_high=0.2 trainer.algorithm.loss_reduction=token_mean trainer.epochs=2 trainer.max_steps=60 trainer.update_epochs_per_batch=1 trainer.train_batch_size=64 trainer.policy_mini_batch_size=64 trainer.eval_batch_size=64 trainer.micro_forward_batch_size_per_gpu=1 trainer.micro_train_batch_size_per_gpu=1 trainer.max_prompt_length=999999 trainer.eval_interval=999999 trainer.eval_before_train=false trainer.ckpt_interval=1 trainer.resume_mode=latest trainer.hf_save_interval=5 ++trainer.hf_hub_repo_id=laion/a2-rl-stack_jest_v2 ++trainer.hf_hub_private=false ++trainer.hf_hub_revision=main ++trainer.enable_db_registration=true trainer.project_name=OpenThoughts-Agent trainer.log_level=INFO trainer.tracker_commit_each_step=true trainer.logger=console trainer.run_name=a2-rl-stack_jest_v2 trainer.ckpt_path=/e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints trainer.export_path=/e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/exports trainer.policy.optimizer_config.lr=9e-6 trainer.policy.optimizer_config.weight_decay=0.0 trainer.policy.optimizer_config.adam_betas=[0.9,0.999] trainer.policy.optimizer_config.max_grad_norm=10.0 trainer.policy.fsdp_config.cpu_offload=true trainer.policy.fsdp_config.reshard_after_forward=true trainer.policy.fsdp_config.fsdp_size=4 trainer.policy.model.path=/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137 trainer.ref.fsdp_config.cpu_offload=true trainer.ref.fsdp_config.reshard_after_forward=true trainer.ref.fsdp_config.fsdp_size=4 trainer.placement.colocate_all=false trainer.placement.policy_num_nodes=4 trainer.placement.ref_num_nodes=4 trainer.placement.policy_num_gpus_per_node=4 trainer.placement.ref_num_gpus_per_node=4 trainer.fully_async.max_staleness_steps=16 trainer.fully_async.num_parallel_generation_workers=512 generator.backend=vllm generator.timeout_multiplier=1.0 generator.model_dtype=bfloat16 generator.inference_engine_tensor_parallel_size=2 generator.num_inference_engines=24 generator.n_samples_per_prompt=8 generator.eval_n_samples_per_prompt=8 generator.gpu_memory_utilization=0.85 generator.max_num_seqs=24 generator.max_num_batched_tokens=16384 generator.enable_prefix_caching=true generator.enable_chunked_prefill=true generator.run_engines_locally=true generator.weight_sync_backend=nccl generator.async_engine=true generator.batched=false generator.enable_http_endpoint=true generator.enable_ray_prometheus_stats=false generator.vllm_stats_interval=1 generator.append_eos_token_after_stop_str_in_multi_turn=true generator.max_turns=999999 generator.sampling_params.max_generate_length=16384 generator.sampling_params.temperature=0.7 generator.sampling_params.top_p=0.95 generator.sampling_params.top_k=20 ++generator.engine_init_kwargs.max_model_len=32768 ++generator.engine_init_kwargs.custom_chat_template_chat_completion_path=chat_templates/qwen3_thinking_acc.jinja2 ++generator.engine_init_kwargs.served_model_name=9216db5781bf21249d130ec9da846c4624c16137 data.train_data=["/e/scratch/jureap59/feuer1/tasks/exp_rpt_stack-jest-v2"] data.val_data=["/e/scratch/jureap59/feuer1/tasks/OpenThoughts-TB-dev"] +terminal_bench_config.trials_dir=/e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/trace_jobs +terminal_bench_config.harbor.name=terminus-2 +terminal_bench_config.harbor.max_episodes=999999 +terminal_bench_config.harbor.enable_summarize=false +terminal_bench_config.harbor.store_all_messages=true +terminal_bench_config.harbor.trajectory_config.raw_content=true +terminal_bench_config.harbor.enable_episode_logging=false +terminal_bench_config.harbor.record_terminal_session=false +terminal_bench_config.harbor.enable_pane_logging=false +terminal_bench_config.harbor.strict_json_parser=true +terminal_bench_config.harbor.interleaved_thinking=true +terminal_bench_config.harbor.extra_body.chat_template_kwargs.enable_thinking=true +terminal_bench_config.harbor.override_timeout_sec=1800 +terminal_bench_config.harbor.override_cpus=1 +terminal_bench_config.harbor.override_memory_mb=2048 +terminal_bench_config.harbor.override_storage_mb=2048 +terminal_bench_config.harbor.auto_snapshot=true +terminal_bench_config.harbor.verifier_override_timeout_sec=120 +terminal_bench_config.harbor.max_retries=3 +terminal_bench_config.harbor.min_wait_sec=60.0 +terminal_bench_config.harbor.max_wait_sec=600.0 +terminal_bench_config.harbor.wait_multiplier=2.0 +terminal_bench_config.harbor.exclude_exceptions=["VerifierTimeoutError","VerifierRuntimeError","RewardFileNotFoundError","RewardFileEmptyError","VerifierOutputParseError"] +terminal_bench_config.harbor.n_concurrent_trials=256 +terminal_bench_config.harbor.log_level=INFO +terminal_bench_config.harbor.enable_reward_shaping=false +terminal_bench_config.harbor.enable_error_classification=true +terminal_bench_config.harbor.mask_exceptions=["DaytonaError","EnvironmentStartTimeoutError","NetworkError","ConnectionError","RewardFileNotFoundError","RewardFileEmptyError","AgentEnvironmentTimeoutError"] +terminal_bench_config.harbor.default_error_treatment=zero +terminal_bench_config.harbor.passthrough_exceptions=["AgentTimeoutError","ContextLengthExceededError"] +terminal_bench_config.model_info.max_input_tokens=32768 +terminal_bench_config.model_info.max_output_tokens=4096 +terminal_bench_config.archiving.enabled=false +terminal_bench_config.trace_upload.enabled=true +terminal_bench_config.trace_upload.repo_org=DCAgent +terminal_bench_config.trace_upload.episodes=last +terminal_bench_config.trace_upload.dataset_type=SFT +terminal_bench_config.trace_upload.cleanup=true |
| [proxychains] config file found: /e/home/jusers/feuer1/jupiter/.proxychains/proxychains_388832.conf |
| [proxychains] preloading /e/scratch/jureap59/feuer1/proxychains-ng-aarch64/lib/libproxychains4.so |
| [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e |
| 2026-04-24 06:37:44.831 | INFO | skyrl_train.utils.utils:prepare_runtime_environment:590 - Exporting wandb api key to ray runtime env |
| 2026-04-24 06:37:44.832 | INFO | skyrl_train.utils.utils:prepare_runtime_environment:609 - Exporting RAY_ADDRESS to ray runtime env |
| 2026-04-24 06:37:44,832 INFO worker.py:1680 -- Using address 10.128.33.33:6379 set in the environment variable RAY_ADDRESS |
| 2026-04-24 06:37:44,861 INFO worker.py:1821 -- Connecting to existing Ray cluster at address: 10.128.33.33:6379... |
| 2026-04-24 06:37:44,870 INFO worker.py:2007 -- Connected to Ray cluster. |
| /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py:2046: FutureWarning: Tip: In future versions of Ray, Ray will no longer override accelerator visible devices env var if num_gpus=0 or num_gpus=None (default). To enable this behavior and turn off this error message, set RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0 |
| warnings.warn( |
| [33m(raylet, ip=10.128.33.46)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e |
| [36m(pid=2093226)[0m [2026-04-24 06:37:45,192 E 2093226 2094275] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 |
| 2026-04-24 06:37:46.913 | INFO | skyrl_train.utils.ppo_utils:sync_registries:546 - Synced registries to ray actor |
| [33m(raylet)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 17x across cluster] (Ray deduplicates logs by default. Set RAY_DEDUP_LOGS=0 to disable log deduplication, or see https: |
| [36m(pid=2093242)[0m [2026-04-24 06:37:50,741 E 2093242 2094821] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 13x across cluster][0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:37:55.806[0m | [1mINFO [0m | [36mskyrl_train.entrypoints.main_base[0m:[36m_configure_log_level[0m:[36m199[0m - [1mSkyRL log level set to: INFO[0m |
| [36m(pid=2093427)[0m [2026-04-24 06:37:55,946 E 2093427 2103528] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 121x across cluster][0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:37:56.047[0m | [1mINFO [0m | [36mexamples.terminal_bench.dataset[0m:[36m_load_data_files[0m:[36m40[0m - [1mLoading data from: /e/scratch/jureap59/feuer1/tasks/exp_rpt_stack-jest-v2[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:37:56.363[0m | [1mINFO [0m | [36mexamples.terminal_bench.dataset[0m:[36m_load_data_files[0m:[36m50[0m - [1mFound 500 valid task directories out of 500 total directories[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:37:56.363[0m | [1mINFO [0m | [36mexamples.terminal_bench.dataset[0m:[36m__init__[0m:[36m27[0m - [1mTerminalBenchTaskDataset initialized with 500 task paths[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:37:56.364[0m | [1mINFO [0m | [36mexamples.terminal_bench.dataset[0m:[36m_load_data_files[0m:[36m40[0m - [1mLoading data from: /e/scratch/jureap59/feuer1/tasks/OpenThoughts-TB-dev[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:37:56.407[0m | [1mINFO [0m | [36mexamples.terminal_bench.dataset[0m:[36m_load_data_files[0m:[36m50[0m - [1mFound 70 valid task directories out of 70 total directories[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:37:56.407[0m | [1mINFO [0m | [36mexamples.terminal_bench.dataset[0m:[36m__init__[0m:[36m27[0m - [1mTerminalBenchTaskDataset initialized with 70 task paths[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:37:56.418[0m | [1mINFO [0m | [36mskyrl_train.entrypoints.main_base[0m:[36m_setup_trainer[0m:[36m352[0m - [1mdata: |
| [36m(skyrl_entrypoint pid=2109818)[0m train_data: |
| [36m(skyrl_entrypoint pid=2109818)[0m - /e/scratch/jureap59/feuer1/tasks/exp_rpt_stack-jest-v2 |
| [36m(skyrl_entrypoint pid=2109818)[0m val_data: |
| [36m(skyrl_entrypoint pid=2109818)[0m - /e/scratch/jureap59/feuer1/tasks/OpenThoughts-TB-dev |
| [36m(skyrl_entrypoint pid=2109818)[0m trainer: |
| [36m(skyrl_entrypoint pid=2109818)[0m placement: |
| [36m(skyrl_entrypoint pid=2109818)[0m colocate_all: false |
| [36m(skyrl_entrypoint pid=2109818)[0m colocate_policy_ref: true |
| [36m(skyrl_entrypoint pid=2109818)[0m policy_num_nodes: 4 |
| [36m(skyrl_entrypoint pid=2109818)[0m policy_num_gpus_per_node: 4 |
| [36m(skyrl_entrypoint pid=2109818)[0m critic_num_nodes: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m critic_num_gpus_per_node: 4 |
| [36m(skyrl_entrypoint pid=2109818)[0m ref_num_nodes: 4 |
| [36m(skyrl_entrypoint pid=2109818)[0m ref_num_gpus_per_node: 4 |
| [36m(skyrl_entrypoint pid=2109818)[0m sequence_parallel_backend: ulysses |
| [36m(skyrl_entrypoint pid=2109818)[0m strategy: fsdp2 |
| [36m(skyrl_entrypoint pid=2109818)[0m policy: |
| [36m(skyrl_entrypoint pid=2109818)[0m model: |
| [36m(skyrl_entrypoint pid=2109818)[0m path: /e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137 |
| [36m(skyrl_entrypoint pid=2109818)[0m lora: |
| [36m(skyrl_entrypoint pid=2109818)[0m rank: 0 |
| [36m(skyrl_entrypoint pid=2109818)[0m alpha: 16 |
| [36m(skyrl_entrypoint pid=2109818)[0m dropout: 0 |
| [36m(skyrl_entrypoint pid=2109818)[0m lora_sync_path: /tmp/skyrl_lora_sync |
| [36m(skyrl_entrypoint pid=2109818)[0m target_modules: all-linear |
| [36m(skyrl_entrypoint pid=2109818)[0m exclude_modules: null |
| [36m(skyrl_entrypoint pid=2109818)[0m deepspeed_config: ${deepspeed_config.train} |
| [36m(skyrl_entrypoint pid=2109818)[0m optimizer_config: |
| [36m(skyrl_entrypoint pid=2109818)[0m optimizer: AdamW |
| [36m(skyrl_entrypoint pid=2109818)[0m lr: 9.0e-06 |
| [36m(skyrl_entrypoint pid=2109818)[0m adam_betas: |
| [36m(skyrl_entrypoint pid=2109818)[0m - 0.9 |
| [36m(skyrl_entrypoint pid=2109818)[0m - 0.999 |
| [36m(skyrl_entrypoint pid=2109818)[0m weight_decay: 0.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m max_grad_norm: 10.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m offload_after_step: true |
| [36m(skyrl_entrypoint pid=2109818)[0m num_warmup_steps: 0 |
| [36m(skyrl_entrypoint pid=2109818)[0m scheduler: constant_with_warmup |
| [36m(skyrl_entrypoint pid=2109818)[0m optimizer_kwargs: {} |
| [36m(skyrl_entrypoint pid=2109818)[0m fsdp_config: |
| [36m(skyrl_entrypoint pid=2109818)[0m cpu_offload: true |
| [36m(skyrl_entrypoint pid=2109818)[0m reshard_after_forward: true |
| [36m(skyrl_entrypoint pid=2109818)[0m fsdp_size: 4 |
| [36m(skyrl_entrypoint pid=2109818)[0m sequence_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m use_torch_compile: false |
| [36m(skyrl_entrypoint pid=2109818)[0m record_memory: false |
| [36m(skyrl_entrypoint pid=2109818)[0m megatron_config: |
| [36m(skyrl_entrypoint pid=2109818)[0m tensor_model_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m pipeline_model_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m context_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m expert_model_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m expert_tensor_parallel_size: null |
| [36m(skyrl_entrypoint pid=2109818)[0m ddp_config: |
| [36m(skyrl_entrypoint pid=2109818)[0m grad_reduce_in_fp32: true |
| [36m(skyrl_entrypoint pid=2109818)[0m overlap_grad_reduce: false |
| [36m(skyrl_entrypoint pid=2109818)[0m overlap_param_gather: false |
| [36m(skyrl_entrypoint pid=2109818)[0m average_in_collective: true |
| [36m(skyrl_entrypoint pid=2109818)[0m model_config_kwargs: {} |
| [36m(skyrl_entrypoint pid=2109818)[0m torch_profiler_config: |
| [36m(skyrl_entrypoint pid=2109818)[0m enable: false |
| [36m(skyrl_entrypoint pid=2109818)[0m ranks: [] |
| [36m(skyrl_entrypoint pid=2109818)[0m save_path: null |
| [36m(skyrl_entrypoint pid=2109818)[0m optimizer_config_kwargs: |
| [36m(skyrl_entrypoint pid=2109818)[0m overlap_cpu_optimizer_d2h_h2d: false |
| [36m(skyrl_entrypoint pid=2109818)[0m use_precision_aware_optimizer: false |
| [36m(skyrl_entrypoint pid=2109818)[0m optimizer_cpu_offload: false |
| [36m(skyrl_entrypoint pid=2109818)[0m optimizer_offload_fraction: 0.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m transformer_config_kwargs: |
| [36m(skyrl_entrypoint pid=2109818)[0m recompute_granularity: full |
| [36m(skyrl_entrypoint pid=2109818)[0m recompute_modules: |
| [36m(skyrl_entrypoint pid=2109818)[0m - core_attn |
| [36m(skyrl_entrypoint pid=2109818)[0m recompute_method: uniform |
| [36m(skyrl_entrypoint pid=2109818)[0m recompute_num_layers: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m empty_cuda_cache: true |
| [36m(skyrl_entrypoint pid=2109818)[0m ref: |
| [36m(skyrl_entrypoint pid=2109818)[0m model: |
| [36m(skyrl_entrypoint pid=2109818)[0m path: ${trainer.policy.model.path} |
| [36m(skyrl_entrypoint pid=2109818)[0m sequence_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m deepspeed_config: ${deepspeed_config.eval} |
| [36m(skyrl_entrypoint pid=2109818)[0m fsdp_config: |
| [36m(skyrl_entrypoint pid=2109818)[0m cpu_offload: true |
| [36m(skyrl_entrypoint pid=2109818)[0m reshard_after_forward: true |
| [36m(skyrl_entrypoint pid=2109818)[0m fsdp_size: 4 |
| [36m(skyrl_entrypoint pid=2109818)[0m megatron_config: |
| [36m(skyrl_entrypoint pid=2109818)[0m tensor_model_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m pipeline_model_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m context_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m expert_model_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m expert_tensor_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m model_config_kwargs: {} |
| [36m(skyrl_entrypoint pid=2109818)[0m transformer_config_kwargs: {} |
| [36m(skyrl_entrypoint pid=2109818)[0m critic: |
| [36m(skyrl_entrypoint pid=2109818)[0m model: |
| [36m(skyrl_entrypoint pid=2109818)[0m path: null |
| [36m(skyrl_entrypoint pid=2109818)[0m lora: |
| [36m(skyrl_entrypoint pid=2109818)[0m rank: 0 |
| [36m(skyrl_entrypoint pid=2109818)[0m alpha: 16 |
| [36m(skyrl_entrypoint pid=2109818)[0m dropout: 0 |
| [36m(skyrl_entrypoint pid=2109818)[0m target_modules: all-linear |
| [36m(skyrl_entrypoint pid=2109818)[0m exclude_modules: null |
| [36m(skyrl_entrypoint pid=2109818)[0m deepspeed_config: ${deepspeed_config.train} |
| [36m(skyrl_entrypoint pid=2109818)[0m optimizer_config: |
| [36m(skyrl_entrypoint pid=2109818)[0m optimizer: AdamW |
| [36m(skyrl_entrypoint pid=2109818)[0m lr: 5.0e-06 |
| [36m(skyrl_entrypoint pid=2109818)[0m adam_betas: |
| [36m(skyrl_entrypoint pid=2109818)[0m - 0.9 |
| [36m(skyrl_entrypoint pid=2109818)[0m - 0.999 |
| [36m(skyrl_entrypoint pid=2109818)[0m weight_decay: 0.01 |
| [36m(skyrl_entrypoint pid=2109818)[0m max_grad_norm: 1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m offload_after_step: true |
| [36m(skyrl_entrypoint pid=2109818)[0m num_warmup_steps: 0 |
| [36m(skyrl_entrypoint pid=2109818)[0m scheduler: constant_with_warmup |
| [36m(skyrl_entrypoint pid=2109818)[0m optimizer_kwargs: {} |
| [36m(skyrl_entrypoint pid=2109818)[0m fsdp_config: |
| [36m(skyrl_entrypoint pid=2109818)[0m cpu_offload: false |
| [36m(skyrl_entrypoint pid=2109818)[0m reshard_after_forward: true |
| [36m(skyrl_entrypoint pid=2109818)[0m fsdp_size: -1 |
| [36m(skyrl_entrypoint pid=2109818)[0m sequence_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m algorithm: |
| [36m(skyrl_entrypoint pid=2109818)[0m advantage_estimator: rloo_n |
| [36m(skyrl_entrypoint pid=2109818)[0m kl_ctrl: |
| [36m(skyrl_entrypoint pid=2109818)[0m type: fixed |
| [36m(skyrl_entrypoint pid=2109818)[0m kl_target: 0.1 |
| [36m(skyrl_entrypoint pid=2109818)[0m horizon: 10000 |
| [36m(skyrl_entrypoint pid=2109818)[0m kl_estimator_type: k3 |
| [36m(skyrl_entrypoint pid=2109818)[0m use_kl_estimator_k3: false |
| [36m(skyrl_entrypoint pid=2109818)[0m use_abs_kl: false |
| [36m(skyrl_entrypoint pid=2109818)[0m use_kl_in_reward: false |
| [36m(skyrl_entrypoint pid=2109818)[0m use_kl_loss: false |
| [36m(skyrl_entrypoint pid=2109818)[0m kl_loss_coef: 0.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m use_entropy_loss: false |
| [36m(skyrl_entrypoint pid=2109818)[0m entropy_loss_coef: 0.01 |
| [36m(skyrl_entrypoint pid=2109818)[0m advantage_batch_normalize: false |
| [36m(skyrl_entrypoint pid=2109818)[0m value_head_prefix: value_head |
| [36m(skyrl_entrypoint pid=2109818)[0m policy_loss_type: regular |
| [36m(skyrl_entrypoint pid=2109818)[0m loss_reduction: token_mean |
| [36m(skyrl_entrypoint pid=2109818)[0m grpo_norm_by_std: true |
| [36m(skyrl_entrypoint pid=2109818)[0m rloo_n_min_group_size: 4 |
| [36m(skyrl_entrypoint pid=2109818)[0m rloo_n_filter_zero_reward_groups: true |
| [36m(skyrl_entrypoint pid=2109818)[0m lambd: 1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m gamma: 1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m eps_clip_low: 0.2 |
| [36m(skyrl_entrypoint pid=2109818)[0m eps_clip_high: 0.2 |
| [36m(skyrl_entrypoint pid=2109818)[0m clip_ratio_c: 3.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m tis_imp_ratio_cap: -1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m use_tis: false |
| [36m(skyrl_entrypoint pid=2109818)[0m sapo: |
| [36m(skyrl_entrypoint pid=2109818)[0m tau_pos: 1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m tau_neg: 1.05 |
| [36m(skyrl_entrypoint pid=2109818)[0m value_clip: 0.2 |
| [36m(skyrl_entrypoint pid=2109818)[0m dynamic_sampling: |
| [36m(skyrl_entrypoint pid=2109818)[0m type: null |
| [36m(skyrl_entrypoint pid=2109818)[0m max_sample_batches: 30 |
| [36m(skyrl_entrypoint pid=2109818)[0m min_replace_ratio: 0.3 |
| [36m(skyrl_entrypoint pid=2109818)[0m clip_cov: |
| [36m(skyrl_entrypoint pid=2109818)[0m clip_ratio: 0.0002 |
| [36m(skyrl_entrypoint pid=2109818)[0m clip_cov_lb: 1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m clip_cov_ub: 5.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m kl_cov: |
| [36m(skyrl_entrypoint pid=2109818)[0m kl_cov_frac: 0.2 |
| [36m(skyrl_entrypoint pid=2109818)[0m ppo_kl_coef: 1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m cispo: |
| [36m(skyrl_entrypoint pid=2109818)[0m cispo_eps_clip_low: 0 |
| [36m(skyrl_entrypoint pid=2109818)[0m cispo_eps_clip_high: 5 |
| [36m(skyrl_entrypoint pid=2109818)[0m max_seq_len: 1016383 |
| [36m(skyrl_entrypoint pid=2109818)[0m fully_async: |
| [36m(skyrl_entrypoint pid=2109818)[0m max_staleness_steps: 16 |
| [36m(skyrl_entrypoint pid=2109818)[0m num_parallel_generation_workers: 512 |
| [36m(skyrl_entrypoint pid=2109818)[0m gradient_checkpointing: true |
| [36m(skyrl_entrypoint pid=2109818)[0m gradient_checkpointing_use_reentrant: false |
| [36m(skyrl_entrypoint pid=2109818)[0m seed: 42 |
| [36m(skyrl_entrypoint pid=2109818)[0m resume_mode: latest |
| [36m(skyrl_entrypoint pid=2109818)[0m resume_path: null |
| [36m(skyrl_entrypoint pid=2109818)[0m ckpt_path: /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints |
| [36m(skyrl_entrypoint pid=2109818)[0m max_ckpts_to_keep: -1 |
| [36m(skyrl_entrypoint pid=2109818)[0m ckpt_interval: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m hf_save_interval: 5 |
| [36m(skyrl_entrypoint pid=2109818)[0m export_path: /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/exports |
| [36m(skyrl_entrypoint pid=2109818)[0m bf16: true |
| [36m(skyrl_entrypoint pid=2109818)[0m epochs: 2 |
| [36m(skyrl_entrypoint pid=2109818)[0m max_steps: 60 |
| [36m(skyrl_entrypoint pid=2109818)[0m update_epochs_per_batch: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m train_batch_size: 64 |
| [36m(skyrl_entrypoint pid=2109818)[0m policy_mini_batch_size: 64 |
| [36m(skyrl_entrypoint pid=2109818)[0m critic_mini_batch_size: 256 |
| [36m(skyrl_entrypoint pid=2109818)[0m micro_train_batch_size_per_gpu: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m micro_forward_batch_size_per_gpu: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m update_ref_every_epoch: false |
| [36m(skyrl_entrypoint pid=2109818)[0m use_sample_packing: true |
| [36m(skyrl_entrypoint pid=2109818)[0m eval_batch_size: 64 |
| [36m(skyrl_entrypoint pid=2109818)[0m eval_before_train: false |
| [36m(skyrl_entrypoint pid=2109818)[0m eval_interval: 999999 |
| [36m(skyrl_entrypoint pid=2109818)[0m max_prompt_length: 999999 |
| [36m(skyrl_entrypoint pid=2109818)[0m flash_attn: true |
| [36m(skyrl_entrypoint pid=2109818)[0m disable_fast_tokenizer: false |
| [36m(skyrl_entrypoint pid=2109818)[0m target_modules: null |
| [36m(skyrl_entrypoint pid=2109818)[0m exclude_modules: null |
| [36m(skyrl_entrypoint pid=2109818)[0m project_name: OpenThoughts-Agent |
| [36m(skyrl_entrypoint pid=2109818)[0m run_name: a2-rl-stack_jest_v2 |
| [36m(skyrl_entrypoint pid=2109818)[0m logger: console |
| [36m(skyrl_entrypoint pid=2109818)[0m tracker_commit_each_step: true |
| [36m(skyrl_entrypoint pid=2109818)[0m dump_data_batch: false |
| [36m(skyrl_entrypoint pid=2109818)[0m dump_eval_results: true |
| [36m(skyrl_entrypoint pid=2109818)[0m log_level: INFO |
| [36m(skyrl_entrypoint pid=2109818)[0m rope_scaling: null |
| [36m(skyrl_entrypoint pid=2109818)[0m rope_theta: null |
| [36m(skyrl_entrypoint pid=2109818)[0m step_wise_training: false |
| [36m(skyrl_entrypoint pid=2109818)[0m hf_hub_repo_id: laion/a2-rl-stack_jest_v2 |
| [36m(skyrl_entrypoint pid=2109818)[0m hf_hub_private: false |
| [36m(skyrl_entrypoint pid=2109818)[0m hf_hub_revision: main |
| [36m(skyrl_entrypoint pid=2109818)[0m enable_db_registration: true |
| [36m(skyrl_entrypoint pid=2109818)[0m generator: |
| [36m(skyrl_entrypoint pid=2109818)[0m model_name: ${trainer.policy.model.path} |
| [36m(skyrl_entrypoint pid=2109818)[0m model_dtype: bfloat16 |
| [36m(skyrl_entrypoint pid=2109818)[0m timeout_multiplier: 1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m run_engines_locally: true |
| [36m(skyrl_entrypoint pid=2109818)[0m num_inference_engines: 24 |
| [36m(skyrl_entrypoint pid=2109818)[0m backend: vllm |
| [36m(skyrl_entrypoint pid=2109818)[0m weight_sync_backend: nccl |
| [36m(skyrl_entrypoint pid=2109818)[0m fuse_weights: false |
| [36m(skyrl_entrypoint pid=2109818)[0m weight_transfer_threshold_cuda_ipc_GB: 1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m inference_engine_tensor_parallel_size: 2 |
| [36m(skyrl_entrypoint pid=2109818)[0m inference_engine_pipeline_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m inference_engine_expert_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m inference_engine_data_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m n_samples_per_prompt: 8 |
| [36m(skyrl_entrypoint pid=2109818)[0m async_engine: true |
| [36m(skyrl_entrypoint pid=2109818)[0m batched: false |
| [36m(skyrl_entrypoint pid=2109818)[0m max_input_length: ${trainer.max_prompt_length} |
| [36m(skyrl_entrypoint pid=2109818)[0m vllm_v1_disable_multiproc: true |
| [36m(skyrl_entrypoint pid=2109818)[0m enable_prefix_caching: true |
| [36m(skyrl_entrypoint pid=2109818)[0m enable_chunked_prefill: true |
| [36m(skyrl_entrypoint pid=2109818)[0m max_num_batched_tokens: 16384 |
| [36m(skyrl_entrypoint pid=2109818)[0m enforce_eager: true |
| [36m(skyrl_entrypoint pid=2109818)[0m fully_sharded_loras: false |
| [36m(skyrl_entrypoint pid=2109818)[0m enable_ray_prometheus_stats: false |
| [36m(skyrl_entrypoint pid=2109818)[0m vllm_stats_interval: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m gpu_memory_utilization: 0.85 |
| [36m(skyrl_entrypoint pid=2109818)[0m max_num_seqs: 24 |
| [36m(skyrl_entrypoint pid=2109818)[0m remote_inference_engine_urls: |
| [36m(skyrl_entrypoint pid=2109818)[0m - 127.0.0.1:8001 |
| [36m(skyrl_entrypoint pid=2109818)[0m enable_http_endpoint: true |
| [36m(skyrl_entrypoint pid=2109818)[0m http_endpoint_host: 127.0.0.1 |
| [36m(skyrl_entrypoint pid=2109818)[0m http_endpoint_port: 8000 |
| [36m(skyrl_entrypoint pid=2109818)[0m max_turns: 999999 |
| [36m(skyrl_entrypoint pid=2109818)[0m chat_template: |
| [36m(skyrl_entrypoint pid=2109818)[0m source: name |
| [36m(skyrl_entrypoint pid=2109818)[0m name_or_path: null |
| [36m(skyrl_entrypoint pid=2109818)[0m chat_template_kwargs: {} |
| [36m(skyrl_entrypoint pid=2109818)[0m engine_init_kwargs: |
| [36m(skyrl_entrypoint pid=2109818)[0m max_model_len: 32768 |
| [36m(skyrl_entrypoint pid=2109818)[0m custom_chat_template_chat_completion_path: chat_templates/qwen3_thinking_acc.jinja2 |
| [36m(skyrl_entrypoint pid=2109818)[0m served_model_name: 9216db5781bf21249d130ec9da846c4624c16137 |
| [36m(skyrl_entrypoint pid=2109818)[0m override_existing_update_group: disable |
| [36m(skyrl_entrypoint pid=2109818)[0m sampling_params: |
| [36m(skyrl_entrypoint pid=2109818)[0m max_generate_length: 16384 |
| [36m(skyrl_entrypoint pid=2109818)[0m repetition_penalty: 1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m temperature: 0.7 |
| [36m(skyrl_entrypoint pid=2109818)[0m top_p: 0.95 |
| [36m(skyrl_entrypoint pid=2109818)[0m min_p: 0.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m top_k: 20 |
| [36m(skyrl_entrypoint pid=2109818)[0m logprobs: null |
| [36m(skyrl_entrypoint pid=2109818)[0m stop: null |
| [36m(skyrl_entrypoint pid=2109818)[0m use_conversation_multi_turn: true |
| [36m(skyrl_entrypoint pid=2109818)[0m append_eos_token_after_stop_str_in_multi_turn: true |
| [36m(skyrl_entrypoint pid=2109818)[0m eval_sampling_params: |
| [36m(skyrl_entrypoint pid=2109818)[0m max_generate_length: ${generator.sampling_params.max_generate_length} |
| [36m(skyrl_entrypoint pid=2109818)[0m repetition_penalty: 1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m temperature: 0.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m top_p: 1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m min_p: 0.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m top_k: -1 |
| [36m(skyrl_entrypoint pid=2109818)[0m logprobs: null |
| [36m(skyrl_entrypoint pid=2109818)[0m stop: null |
| [36m(skyrl_entrypoint pid=2109818)[0m eval_n_samples_per_prompt: 8 |
| [36m(skyrl_entrypoint pid=2109818)[0m zero_reward_on_non_stop: false |
| [36m(skyrl_entrypoint pid=2109818)[0m apply_overlong_filtering: false |
| [36m(skyrl_entrypoint pid=2109818)[0m rope_scaling: ${trainer.rope_scaling} |
| [36m(skyrl_entrypoint pid=2109818)[0m rope_theta: ${trainer.rope_theta} |
| [36m(skyrl_entrypoint pid=2109818)[0m teacher: |
| [36m(skyrl_entrypoint pid=2109818)[0m model_path: null |
| [36m(skyrl_entrypoint pid=2109818)[0m top_k_logprobs: 256 |
| [36m(skyrl_entrypoint pid=2109818)[0m num_inference_engines: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m inference_engine_tensor_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m inference_engine_pipeline_parallel_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m gpu_memory_utilization: 0.9 |
| [36m(skyrl_entrypoint pid=2109818)[0m enforce_eager: false |
| [36m(skyrl_entrypoint pid=2109818)[0m backend: vllm |
| [36m(skyrl_entrypoint pid=2109818)[0m engine_init_kwargs: {} |
| [36m(skyrl_entrypoint pid=2109818)[0m environment: |
| [36m(skyrl_entrypoint pid=2109818)[0m env_class: gsm8k |
| [36m(skyrl_entrypoint pid=2109818)[0m skyrl_gym: |
| [36m(skyrl_entrypoint pid=2109818)[0m max_env_workers: 32 |
| [36m(skyrl_entrypoint pid=2109818)[0m text2sql: |
| [36m(skyrl_entrypoint pid=2109818)[0m db_path: /home/ray/default/sql_data |
| [36m(skyrl_entrypoint pid=2109818)[0m llm_as_a_judge: |
| [36m(skyrl_entrypoint pid=2109818)[0m model: gpt-4o-mini |
| [36m(skyrl_entrypoint pid=2109818)[0m base_url: null |
| [36m(skyrl_entrypoint pid=2109818)[0m search: |
| [36m(skyrl_entrypoint pid=2109818)[0m log_requests: false |
| [36m(skyrl_entrypoint pid=2109818)[0m search_url: http: |
| [36m(skyrl_entrypoint pid=2109818)[0m topk: 3 |
| [36m(skyrl_entrypoint pid=2109818)[0m timeout: 30 |
| [36m(skyrl_entrypoint pid=2109818)[0m deepspeed_config: |
| [36m(skyrl_entrypoint pid=2109818)[0m train: |
| [36m(skyrl_entrypoint pid=2109818)[0m zero_optimization: |
| [36m(skyrl_entrypoint pid=2109818)[0m stage: 3 |
| [36m(skyrl_entrypoint pid=2109818)[0m offload_param: |
| [36m(skyrl_entrypoint pid=2109818)[0m device: none |
| [36m(skyrl_entrypoint pid=2109818)[0m offload_optimizer: |
| [36m(skyrl_entrypoint pid=2109818)[0m device: none |
| [36m(skyrl_entrypoint pid=2109818)[0m pin_memory: true |
| [36m(skyrl_entrypoint pid=2109818)[0m sub_group_size: auto |
| [36m(skyrl_entrypoint pid=2109818)[0m reduce_bucket_size: auto |
| [36m(skyrl_entrypoint pid=2109818)[0m stage3_param_persistence_threshold: auto |
| [36m(skyrl_entrypoint pid=2109818)[0m stage3_prefetch_bucket_size: auto |
| [36m(skyrl_entrypoint pid=2109818)[0m stage3_max_live_parameters: auto |
| [36m(skyrl_entrypoint pid=2109818)[0m stage3_max_reuse_distance: auto |
| [36m(skyrl_entrypoint pid=2109818)[0m round_robin_gradients: true |
| [36m(skyrl_entrypoint pid=2109818)[0m zero_hpz_partition_size: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m zero_quantized_weights: false |
| [36m(skyrl_entrypoint pid=2109818)[0m zero_quantized_gradients: false |
| [36m(skyrl_entrypoint pid=2109818)[0m torch_autocast: |
| [36m(skyrl_entrypoint pid=2109818)[0m enabled: true |
| [36m(skyrl_entrypoint pid=2109818)[0m dtype: bfloat16 |
| [36m(skyrl_entrypoint pid=2109818)[0m disable_trace_cache: false |
| [36m(skyrl_entrypoint pid=2109818)[0m data_types: |
| [36m(skyrl_entrypoint pid=2109818)[0m grad_accum_dtype: fp32 |
| [36m(skyrl_entrypoint pid=2109818)[0m gradient_clipping: 1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m wall_clock_breakdown: false |
| [36m(skyrl_entrypoint pid=2109818)[0m prescale_gradient: false |
| [36m(skyrl_entrypoint pid=2109818)[0m eval: |
| [36m(skyrl_entrypoint pid=2109818)[0m zero_optimization: |
| [36m(skyrl_entrypoint pid=2109818)[0m stage: 3 |
| [36m(skyrl_entrypoint pid=2109818)[0m stage3_param_persistence_threshold: auto |
| [36m(skyrl_entrypoint pid=2109818)[0m offload_param: |
| [36m(skyrl_entrypoint pid=2109818)[0m device: cpu |
| [36m(skyrl_entrypoint pid=2109818)[0m pin_memory: true |
| [36m(skyrl_entrypoint pid=2109818)[0m torch_autocast: |
| [36m(skyrl_entrypoint pid=2109818)[0m enabled: true |
| [36m(skyrl_entrypoint pid=2109818)[0m dtype: bfloat16 |
| [36m(skyrl_entrypoint pid=2109818)[0m gradient_clipping: 1.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m prescale_gradient: false |
| [36m(skyrl_entrypoint pid=2109818)[0m wall_clock_breakdown: false |
| [36m(skyrl_entrypoint pid=2109818)[0m terminal_bench_config: |
| [36m(skyrl_entrypoint pid=2109818)[0m trials_dir: /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/trace_jobs |
| [36m(skyrl_entrypoint pid=2109818)[0m harbor: |
| [36m(skyrl_entrypoint pid=2109818)[0m name: terminus-2 |
| [36m(skyrl_entrypoint pid=2109818)[0m max_episodes: 999999 |
| [36m(skyrl_entrypoint pid=2109818)[0m enable_summarize: false |
| [36m(skyrl_entrypoint pid=2109818)[0m store_all_messages: true |
| [36m(skyrl_entrypoint pid=2109818)[0m trajectory_config: |
| [36m(skyrl_entrypoint pid=2109818)[0m raw_content: true |
| [36m(skyrl_entrypoint pid=2109818)[0m enable_episode_logging: false |
| [36m(skyrl_entrypoint pid=2109818)[0m record_terminal_session: false |
| [36m(skyrl_entrypoint pid=2109818)[0m enable_pane_logging: false |
| [36m(skyrl_entrypoint pid=2109818)[0m strict_json_parser: true |
| [36m(skyrl_entrypoint pid=2109818)[0m interleaved_thinking: true |
| [36m(skyrl_entrypoint pid=2109818)[0m extra_body: |
| [36m(skyrl_entrypoint pid=2109818)[0m chat_template_kwargs: |
| [36m(skyrl_entrypoint pid=2109818)[0m enable_thinking: true |
| [36m(skyrl_entrypoint pid=2109818)[0m override_timeout_sec: 1800 |
| [36m(skyrl_entrypoint pid=2109818)[0m override_cpus: 1 |
| [36m(skyrl_entrypoint pid=2109818)[0m override_memory_mb: 2048 |
| [36m(skyrl_entrypoint pid=2109818)[0m override_storage_mb: 2048 |
| [36m(skyrl_entrypoint pid=2109818)[0m auto_snapshot: true |
| [36m(skyrl_entrypoint pid=2109818)[0m verifier_override_timeout_sec: 120 |
| [36m(skyrl_entrypoint pid=2109818)[0m max_retries: 3 |
| [36m(skyrl_entrypoint pid=2109818)[0m min_wait_sec: 60.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m max_wait_sec: 600.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m wait_multiplier: 2.0 |
| [36m(skyrl_entrypoint pid=2109818)[0m exclude_exceptions: |
| [36m(skyrl_entrypoint pid=2109818)[0m - VerifierTimeoutError |
| [36m(skyrl_entrypoint pid=2109818)[0m - VerifierRuntimeError |
| [36m(skyrl_entrypoint pid=2109818)[0m - RewardFileNotFoundError |
| [36m(skyrl_entrypoint pid=2109818)[0m - RewardFileEmptyError |
| [36m(skyrl_entrypoint pid=2109818)[0m - VerifierOutputParseError |
| [36m(skyrl_entrypoint pid=2109818)[0m n_concurrent_trials: 256 |
| [36m(skyrl_entrypoint pid=2109818)[0m log_level: INFO |
| [36m(skyrl_entrypoint pid=2109818)[0m enable_reward_shaping: false |
| [36m(skyrl_entrypoint pid=2109818)[0m enable_error_classification: true |
| [36m(skyrl_entrypoint pid=2109818)[0m mask_exceptions: |
| [36m(skyrl_entrypoint pid=2109818)[0m - DaytonaError |
| [36m(skyrl_entrypoint pid=2109818)[0m - EnvironmentStartTimeoutError |
| [36m(skyrl_entrypoint pid=2109818)[0m - NetworkError |
| [36m(skyrl_entrypoint pid=2109818)[0m - ConnectionError |
| [36m(skyrl_entrypoint pid=2109818)[0m - RewardFileNotFoundError |
| [36m(skyrl_entrypoint pid=2109818)[0m - RewardFileEmptyError |
| [36m(skyrl_entrypoint pid=2109818)[0m - AgentEnvironmentTimeoutError |
| [36m(skyrl_entrypoint pid=2109818)[0m default_error_treatment: zero |
| [36m(skyrl_entrypoint pid=2109818)[0m passthrough_exceptions: |
| [36m(skyrl_entrypoint pid=2109818)[0m - AgentTimeoutError |
| [36m(skyrl_entrypoint pid=2109818)[0m - ContextLengthExceededError |
| [36m(skyrl_entrypoint pid=2109818)[0m model_info: |
| [36m(skyrl_entrypoint pid=2109818)[0m max_input_tokens: 32768 |
| [36m(skyrl_entrypoint pid=2109818)[0m max_output_tokens: 4096 |
| [36m(skyrl_entrypoint pid=2109818)[0m archiving: |
| [36m(skyrl_entrypoint pid=2109818)[0m enabled: false |
| [36m(skyrl_entrypoint pid=2109818)[0m trace_upload: |
| [36m(skyrl_entrypoint pid=2109818)[0m enabled: true |
| [36m(skyrl_entrypoint pid=2109818)[0m repo_org: DCAgent |
| [36m(skyrl_entrypoint pid=2109818)[0m episodes: last |
| [36m(skyrl_entrypoint pid=2109818)[0m dataset_type: SFT |
| [36m(skyrl_entrypoint pid=2109818)[0m cleanup: true |
| [36m(skyrl_entrypoint pid=2109818)[0m [0m |
| [36m(skyrl_entrypoint pid=2109818)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: |
| [36m(skyrl_entrypoint pid=2109818)[0m No module named 'vllm._version' |
| [36m(skyrl_entrypoint pid=2109818)[0m from .version import __version__, __version_tuple__ # isort:skip |
| [36m(skyrl_entrypoint pid=2109818)[0m W0424 06:38:01.093000 2109818 envs/rl/lib/python3.12/site-packages/torch/utils/cpp_extension.py:117] No CUDA runtime is found, using CUDA_HOME='/e/software/default/stages/2026/software/CUDA/13' |
| [36m(pid=2093357)[0m [2026-04-24 06:37:56,337 E 2093357 2108164] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 147x across cluster][0m |
| [33m(raylet)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e |
| [36m(pid=1698116, ip=10.128.33.39)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: |
| [36m(pid=1698116, ip=10.128.33.39)[0m No module named 'vllm._version' |
| [36m(pid=1698116, ip=10.128.33.39)[0m from .version import __version__, __version_tuple__ # isort:skip |
| [33m(raylet, ip=10.128.33.42)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 41x across cluster][0m |
| [36m(pid=831169, ip=10.128.33.43)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 3x across cluster][0m |
| [36m(pid=831169, ip=10.128.33.43)[0m No module named 'vllm._version'[32m [repeated 3x across cluster][0m |
| [36m(pid=831169, ip=10.128.33.43)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 3x across cluster][0m |
| [33m(raylet, ip=10.128.33.44)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 62x across cluster][0m |
| [2026-04-24 06:38:15,092 E 2109408 2109790] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 |
| [36m(RegistryActor pid=1955327, ip=10.128.33.46)[0m [2026-04-24 06:38:15,813 E 1955327 1955367] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 |
| [36m(pid=1698116, ip=10.128.33.39)[0m W0424 06:38:16.000000 1698116 envs/rl/lib/python3.12/site-packages/torch/utils/cpp_extension.py:117] No CUDA runtime is found, using CUDA_HOME='/e/software/default/stages/2026/software/CUDA/13' |
| [36m(pid=2361797, ip=10.128.33.37)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 9x across cluster][0m |
| [36m(pid=2361797, ip=10.128.33.37)[0m No module named 'vllm._version'[32m [repeated 9x across cluster][0m |
| [36m(pid=2361797, ip=10.128.33.37)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 9x across cluster][0m |
| [33m(raylet, ip=10.128.33.35)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 46x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m 2026-04-24 06:38:17.724 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:110 - creating LLM with bundle_indices=[0, 1] |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m 2026-04-24 06:38:17.724 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:123 - setup_envvars_for_vllm: distributed_executor_backend=ray, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=<unset>, VLLM_ENABLE_V1_MULTIPROCESSING=<unset> |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m 2026-04-24 06:38:17.724 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:481 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=<unset> |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m 2026-04-24 06:38:17.724 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:489 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0 |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m 2026-04-24 06:38:17.724 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1047 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20 |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m 2026-04-24 06:38:17.743 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1078 - Engine startup stagger: sleeping 2.00s to avoid port collisions |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored. |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored. |
| [36m(skyrl_entrypoint pid=2109818)[0m [2026-04-24 06:38:17,652 E 2109818 2109859] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 2x across cluster][0m |
| [36m(pid=831169, ip=10.128.33.43)[0m W0424 06:38:19.971000 831169 envs/rl/lib/python3.12/site-packages/torch/utils/cpp_extension.py:117] No CUDA runtime is found, using CUDA_HOME='/e/software/default/stages/2026/software/CUDA/13'[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m No module named 'vllm._version'[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 72x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m 2026-04-24 06:38:22.824 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:110 - creating LLM with bundle_indices=[8, 9][32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m 2026-04-24 06:38:22.824 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:123 - setup_envvars_for_vllm: distributed_executor_backend=ray, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=<unset>, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m 2026-04-24 06:38:22.824 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:481 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m 2026-04-24 06:38:22.824 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:489 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m 2026-04-24 06:38:22.824 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1047 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m 2026-04-24 06:38:22.844 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1078 - Engine startup stagger: sleeping 1.72s to avoid port collisions[32m [repeated 6x across cluster][0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [2026-04-24 06:38:23] INFO inference_engine_client_http_endpoint.py:350: Starting server on 127.0.0.1:8000 |
| [36m(skyrl_entrypoint pid=2109818)[0m [2026-04-24 06:38:23] INFO inference_engine_client_http_endpoint.py:242: Starting inference HTTP endpoint... |
| [36m(skyrl_entrypoint pid=2109818)[0m [2026-04-24 06:38:24] INFO inference_engine_client_http_endpoint.py:229: Server ready after 2 attempts (2 seconds) |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:38:24.237[0m | [1mINFO [0m | [36mskyrl_train.inference_engines.inference_engine_client[0m:[36m_spin_up_http_endpoint[0m:[36m960[0m - [1mInferenceEngineClient HTTP endpoint started on 127.0.0.1:8000[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:38:24.237[0m | [1mINFO [0m | [36mskyrl_train.inference_engines.inference_engine_client[0m:[36m__init__[0m:[36m61[0m - [1mInferenceEngineClient initialized with 24 engines.[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:38:24.239[0m | [1mINFO [0m | [36mexamples.terminal_bench.terminal_bench_generator[0m:[36m_configure_harbor_logging[0m:[36m172[0m - [1mHarbor logging level set to INFO[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:38:24.240[0m | [1mINFO [0m | [36mexamples.terminal_bench.terminal_bench_generator[0m:[36m__init__[0m:[36m106[0m - [1mTerminalBenchGenerator initialized with HarborConfigBuilder. Exposed fields: ['name', 'max_episodes', 'enable_summarize', 'store_all_messages', 'trajectory_config', 'enable_episode_logging', 'record_terminal_session', 'enable_pane_logging', 'strict_json_parser', 'interleaved_thinking', 'extra_body', 'override_timeout_sec', 'override_cpus', 'override_memory_mb', 'override_storage_mb', 'auto_snapshot', 'verifier_override_timeout_sec', 'max_retries', 'min_wait_sec', 'max_wait_sec', 'wait_multiplier', 'exclude_exceptions', 'n_concurrent_trials', 'log_level', 'enable_reward_shaping', 'enable_error_classification', 'mask_exceptions', 'default_error_treatment', 'passthrough_exceptions']. Retry config: max_retries=3, backoff=60.0-600.0s. Concurrent trials: 256. Reward shaping: enabled=False, shaper=pass_ratio. Error classification: enabled=True[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:38:24.241[0m | [1mINFO [0m | [36mexamples.terminal_bench.terminal_bench_generator[0m:[36m__init__[0m:[36m122[0m - [1mTerminalBenchGenerator initialized with custom chat template read from: chat_templates/qwen3_thinking_acc.jinja2[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:38:24.242[0m | [1mINFO [0m | [36mskyrl_train.utils.trainer_utils[0m:[36mbuild_dataloader[0m:[36m656[0m - [1mTotal steps: 14[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:38:24.242[0m | [1mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_build_train_dataloader_and_compute_training_steps[0m:[36m351[0m - [1mLength of train_dataloader: 500[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:38:24.242[0m | [1mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_build_train_dataloader_and_compute_training_steps[0m:[36m352[0m - [1mNumber of steps per epoch: 7[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:38:24.242[0m | [1mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_build_train_dataloader_and_compute_training_steps[0m:[36m353[0m - [1mTotal training steps: 14[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:38:24.242[0m | [1mINFO [0m | [36mskyrl_train.utils.trainer_utils[0m:[36mbuild_dataloader[0m:[36m658[0m - [1mValidation set size: 2[0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.[32m [repeated 12x across cluster][0m |
| [36m(pid=1955400, ip=10.128.33.46)[0m W0424 06:38:26.204000 1955400 envs/rl/lib/python3.12/site-packages/torch/utils/cpp_extension.py:117] No CUDA runtime is found, using CUDA_HOME='/e/software/default/stages/2026/software/CUDA/13'[32m [repeated 8x across cluster][0m |
| [36m(pid=4151080, ip=10.128.33.41)[0m Using blocking ray.get inside async actor. This blocks the event loop. Please use `await` on object ref with asyncio.gather if you want to yield execution to the event loop instead. |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 11x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m No module named 'vllm._version'[32m [repeated 11x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 11x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 27x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m 2026-04-24 06:38:27.895 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:110 - creating LLM with bundle_indices=[38, 39][32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m 2026-04-24 06:38:27.895 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:123 - setup_envvars_for_vllm: distributed_executor_backend=ray, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=<unset>, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m 2026-04-24 06:38:27.895 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:481 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m 2026-04-24 06:38:27.895 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:489 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m 2026-04-24 06:38:27.896 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1047 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m 2026-04-24 06:38:27.910 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1078 - Engine startup stagger: sleeping 2.10s to avoid port collisions[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) 2026-04-24 06:38:28,069 INFO worker.py:1680 -- Using address 10.128.33.33:6379 set in the environment variable RAY_ADDRESS |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) 2026-04-24 06:38:28,097 INFO worker.py:1821 -- Connecting to existing Ray cluster at address: 10.128.33.33:6379... |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) 2026-04-24 06:38:28,107 INFO worker.py:2007 -- Connected to Ray cluster. |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.[32m [repeated 18x across cluster][0m |
| [36m(pid=1994236, ip=10.128.33.45)[0m W0424 06:38:30.776000 1994236 envs/rl/lib/python3.12/site-packages/torch/utils/cpp_extension.py:117] No CUDA runtime is found, using CUDA_HOME='/e/software/default/stages/2026/software/CUDA/13'[32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py:2046: FutureWarning: Tip: In future versions of Ray, Ray will no longer override accelerator visible devices env var if num_gpus=0 or num_gpus=None (default). To enable this behavior and turn off this error message, set RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0 |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) warnings.warn( |
| [36m(bundle_reservation_check_func pid=2109901)[0m [2026-04-24 06:38:32,809 E 2109901 2109941] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m No module named 'vllm._version'[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 6x across cluster][0m |
| [33m(raylet, ip=10.128.33.43)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 1074x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m 2026-04-24 06:38:32.121 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:110 - creating LLM with bundle_indices=[42, 43][32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m 2026-04-24 06:38:32.121 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:123 - setup_envvars_for_vllm: distributed_executor_backend=ray, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=<unset>, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m 2026-04-24 06:38:32.121 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:481 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m 2026-04-24 06:38:32.121 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:489 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m 2026-04-24 06:38:32.121 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1047 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m 2026-04-24 06:38:32.143 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1078 - Engine startup stagger: sleeping 1.92s to avoid port collisions[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) 2026-04-24 06:38:32,111 INFO worker.py:1680 -- Using address 10.128.33.33:6379 set in the environment variable RAY_ADDRESS[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) 2026-04-24 06:38:32,143 INFO worker.py:1821 -- Connecting to existing Ray cluster at address: 10.128.33.33:6379...[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) 2026-04-24 06:38:32,152 INFO worker.py:2007 -- Connected to Ray cluster.[32m [repeated 4x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m 2026-04-24 06:38:33 INFO [ipv4-debug] hostname=jpbo-047-09.jupiter.internal |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m 2026-04-24 06:38:33 INFO [ipv4-debug] _global_node.node_ip_address=10.128.33.41 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m 2026-04-24 06:38:33 INFO [ipv4-debug] get_node_ip_address()=10.128.33.41 |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:38:33.368[0m | [1m[32mINFO [0m | [36mskyrl_train.workers.worker[0m:[36m_initiate_actors[0m:[36m496[0m - [1m[32mInitializing process group for RayActorGroup[0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [W424 06:38:33.489203296 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-047-09-interconnect-1.jupiter.internal]:45717 (errno: 97 - Address family not supported by protocol). |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.[32m [repeated 12x across cluster][0m |
| [36m(pid=4151157, ip=10.128.33.41)[0m Using blocking ray.get inside async actor. This blocks the event loop. Please use `await` on object ref with asyncio.gather if you want to yield execution to the event loop instead. |
| [36m(pid=1671783, ip=10.128.33.47)[0m W0424 06:38:33.066000 1671783 envs/rl/lib/python3.12/site-packages/torch/utils/cpp_extension.py:117] No CUDA runtime is found, using CUDA_HOME='/e/software/default/stages/2026/software/CUDA/13'[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [33m(raylet, ip=10.128.33.44)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 980x across cluster] (Ray deduplicates logs by default. Set RAY_DEDUP_LOGS=0 to disable log deduplication, or see https://docs.ray.io/en/master/ray-observability/user-guides/configure-logging.html#log-deduplication for more options.)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1989927, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990134) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py:2046: FutureWarning: Tip: In future versions of Ray, Ray will no longer override accelerator visible devices env var if num_gpus=0 or num_gpus=None (default). To enable this behavior and turn off this error message, set RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0 |
| [36m(AsyncVLLMInferenceEngine pid=1989927, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990134) warnings.warn( |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m [2026-04-24 06:38:36,694 E 2215993 2216096] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m No module named 'vllm._version'[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 7x across cluster][0m |
| [33m(raylet, ip=10.128.33.44)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 965x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m 2026-04-24 06:38:34.328 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:110 - creating LLM with bundle_indices=[46, 47][32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m 2026-04-24 06:38:34.329 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:setup_envvars_for_vllm:123 - setup_envvars_for_vllm: distributed_executor_backend=ray, SKYRL_ENABLE_NUMA_AFFINITY=1, CUDA_VISIBLE_DEVICES=<unset>, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m 2026-04-24 06:38:34.329 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:481 - BaseVLLMInferenceEngine: vllm_v1_disable_multiproc=True, vllm.__version__=dev, VLLM_ENABLE_V1_MULTIPROCESSING=<unset>[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m 2026-04-24 06:38:34.329 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:__init__:489 - BaseVLLMInferenceEngine: set VLLM_ENABLE_V1_MULTIPROCESSING=0[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m 2026-04-24 06:38:34.329 | WARNING | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1047 - OpenAI API sampling params overridden: temperature=0.7, top_p=0.95, top_k=20[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m 2026-04-24 06:38:34.350 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1078 - Engine startup stagger: sleeping 2.79s to avoid port collisions[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) 2026-04-24 06:38:37,716 INFO worker.py:1680 -- Using address 10.128.33.33:6379 set in the environment variable RAY_ADDRESS[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) 2026-04-24 06:38:37,744 INFO worker.py:1821 -- Connecting to existing Ray cluster at address: 10.128.33.33:6379...[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) 2026-04-24 06:38:37,804 INFO worker.py:2007 -- Connected to Ray cluster.[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1958270, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958353) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py:2046: FutureWarning: Tip: In future versions of Ray, Ray will no longer override accelerator visible devices env var if num_gpus=0 or num_gpus=None (default). To enable this behavior and turn off this error message, set RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0 |
| [36m(AsyncVLLMInferenceEngine pid=1958270, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958353) warnings.warn( |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m The argument `trust_remote_code` is to be used with Auto classes. It has no effect here and is ignored.[32m [repeated 4x across cluster][0m |
| [36m(pid=2370204, ip=10.128.33.48)[0m Using blocking ray.get inside async actor. This blocks the event loop. Please use `await` on object ref with asyncio.gather if you want to yield execution to the event loop instead.[32m [repeated 14x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151157, ip=10.128.33.41)[0m [W424 06:38:40.901343659 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-047-09.jupiter.internal]:45717 (errno: 97 - Address family not supported by protocol). |
| [36m(AsyncVLLMInferenceEngine pid=1989927, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990134) [33m(raylet, ip=10.128.33.34)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 981x across cluster] (Ray deduplicates logs by default. Set RAY_DEDUP_LOGS=0 to disable log deduplication, or see https://docs.ray.io/en/master/ray-observability/user-guides/configure-logging.html#log-deduplication for more options.)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1989927, ip=10.128.33.40)[0m [2026-04-24 06:38:41,887 E 1989927 1990030] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [33m(raylet, ip=10.128.33.45)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 1086x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m No module named 'vllm._version' |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m from .version import __version__, __version_tuple__ # isort:skip |
| [33m(raylet, ip=10.128.33.45)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 1016x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) 2026-04-24 06:38:43,043 INFO worker.py:1680 -- Using address 10.128.33.33:6379 set in the environment variable RAY_ADDRESS[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) 2026-04-24 06:38:43,075 INFO worker.py:1821 -- Connecting to existing Ray cluster at address: 10.128.33.33:6379...[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) 2026-04-24 06:38:43,084 INFO worker.py:2007 -- Connected to Ray cluster.[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1958270, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958353) [33m(raylet, ip=10.128.33.47)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 1013x across cluster] (Ray deduplicates logs by default. Set RAY_DEDUP_LOGS=0 to disable log deduplication, or see https://docs.ray.io/en/master/ray-observability/user-guides/configure-logging.html#log-deduplication for more options.)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py:2046: FutureWarning: Tip: In future versions of Ray, Ray will no longer override accelerator visible devices env var if num_gpus=0 or num_gpus=None (default). To enable this behavior and turn off this error message, set RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) warnings.warn( |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py:2046: FutureWarning: Tip: In future versions of Ray, Ray will no longer override accelerator visible devices env var if num_gpus=0 or num_gpus=None (default). To enable this behavior and turn off this error message, set RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0 |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) warnings.warn( |
| [36m(FSDPPolicyWorkerBase pid=2370205, ip=10.128.33.48)[0m [W424 06:38:45.073942351 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-047-09.jupiter.internal]:45717 (errno: 97 - Address family not supported by protocol).[32m [repeated 14x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m `torch_dtype` is deprecated! Use `dtype` instead! |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:38:46.477[0m | [1m[32mINFO [0m | [36mskyrl_train.workers.worker[0m:[36m_initiate_actors[0m:[36m498[0m - [1m[32mInitialized process group for RayActorGroup[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:38:46.494[0m | [1m[32mINFO [0m | [36mskyrl_train.workers.worker[0m:[36m_initiate_actors[0m:[36m500[0m - [1m[32mMesh Ranks: [MeshRank(dp=0, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=1, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=2, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=3, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=4, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=5, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=6, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=7, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=8, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=9, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=10, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=11, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=12, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=13, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=14, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1), MeshRank(dp=15, sp=0, tp=0, pp=0, world_size=16, dp_size=16, pp_size=1)][0m |
| [36m(FSDPPolicyWorkerBase pid=2370205, ip=10.128.33.48)[0m
Loading checkpoint shards: 0%| | 0/17 [00:00<?, ?it/s] |
| [36m(FSDPPolicyWorkerBase pid=2370205, ip=10.128.33.48)[0m
Loading checkpoint shards: 35%|███▌ | 6/17 [00:00<00:00, 52.20it/s] |
| [36m(FSDPPolicyWorkerBase pid=2370207, ip=10.128.33.48)[0m
Loading checkpoint shards: 100%|██████████| 17/17 [00:00<00:00, 48.42it/s] |
| [36m(FSDPPolicyWorkerBase pid=2370205, ip=10.128.33.48)[0m
Loading checkpoint shards: 100%|██████████| 17/17 [00:00<00:00, 42.95it/s]
Loading checkpoint shards: 100%|██████████| 17/17 [00:00<00:00, 44.35it/s] |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m
Loading checkpoint shards: 88%|████████▊ | 15/17 [00:00<00:00, 39.43it/s]
Loading checkpoint shards: 100%|██████████| 17/17 [00:00<00:00, 40.96it/s] |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m [2026-04-24 06:38:47,738 E 1706685 1706789] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 8x across cluster][0m |
| [33m(raylet, ip=10.128.33.37)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 531x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1989927, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990134) [33m(raylet, ip=10.128.33.37)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 1089x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [33m(raylet, ip=10.128.33.47)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 1111x across cluster] (Ray deduplicates logs by default. Set RAY_DEDUP_LOGS=0 to disable log deduplication, or see https://docs.ray.io/en/master/ray-observability/user-guides/configure-logging.html#log-deduplication for more options.)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) [33m(raylet)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 149x across cluster] (Ray deduplicates logs by default. Set RAY_DEDUP_LOGS=0 to disable log deduplication, or see https://docs.ray.io/en/master/ray-observability/user-guides/configure-logging.html#log-deduplication for more options.)[0m |
| [36m(FSDPPolicyWorkerBase pid=2417108, ip=10.128.33.38)[0m `torch_dtype` is deprecated! Use `dtype` instead![32m [repeated 15x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2109998)[0m
Loading checkpoint shards: 0%| | 0/17 [00:00<?, ?it/s][32m [repeated 15x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151159, ip=10.128.33.41)[0m
Loading checkpoint shards: 71%|███████ | 12/17 [00:00<00:00, 41.75it/s][32m [repeated 31x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151158, ip=10.128.33.41)[0m
Loading checkpoint shards: 100%|██████████| 17/17 [00:00<00:00, 39.97it/s][32m [repeated 3x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151159, ip=10.128.33.41)[0m
Loading checkpoint shards: 100%|██████████| 17/17 [00:00<00:00, 40.70it/s]
Loading checkpoint shards: 100%|██████████| 17/17 [00:00<00:00, 41.73it/s][32m [repeated 7x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417108, ip=10.128.33.38)[0m
Loading checkpoint shards: 82%|████████▏ | 14/17 [00:00<00:00, 38.52it/s]
Loading checkpoint shards: 100%|██████████| 17/17 [00:00<00:00, 40.29it/s][32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py:2046: FutureWarning: Tip: In future versions of Ray, Ray will no longer override accelerator visible devices env var if num_gpus=0 or num_gpus=None (default). To enable this behavior and turn off this error message, set RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0 |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) warnings.warn( |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) 2026-04-24 06:38:53,042 INFO worker.py:1680 -- Using address 10.128.33.33:6379 set in the environment variable RAY_ADDRESS |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) 2026-04-24 06:38:53,076 INFO worker.py:1821 -- Connecting to existing Ray cluster at address: 10.128.33.33:6379... |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) 2026-04-24 06:38:53,086 INFO worker.py:2007 -- Connected to Ray cluster. |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m [2026-04-24 06:38:52,475 E 1671655 1671736] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 11x across cluster][0m |
| [33m(raylet, ip=10.128.33.42)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 16x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1958270, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958353) [33m(raylet, ip=10.128.33.45)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 436x across cluster][0m[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(pid=2232400)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(pid=2232400)[0m No module named 'vllm._version' |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(pid=2232400)[0m from .version import __version__, __version_tuple__ # isort:skip |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py:2046: FutureWarning: Tip: In future versions of Ray, Ray will no longer override accelerator visible devices env var if num_gpus=0 or num_gpus=None (default). To enable this behavior and turn off this error message, set RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) warnings.warn([32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) 2026-04-24 06:38:54,281 INFO worker.py:1680 -- Using address 10.128.33.33:6379 set in the environment variable RAY_ADDRESS[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) 2026-04-24 06:38:54,334 INFO worker.py:1821 -- Connecting to existing Ray cluster at address: 10.128.33.33:6379...[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) 2026-04-24 06:38:54,354 INFO worker.py:2007 -- Connected to Ray cluster.[32m [repeated 3x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [2026-04-24 06:38:55,189 E 4151080 4151120] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 2x across cluster][0m |
| [33m(raylet, ip=10.128.33.40)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 101x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) [33m(raylet, ip=10.128.33.43)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 52x across cluster][0m[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(pid=1706852)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(pid=1706852)[0m No module named 'vllm._version'[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(pid=1706852)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [33m(raylet, ip=10.128.33.40)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 106x across cluster] (Ray deduplicates logs by default. Set RAY_DEDUP_LOGS=0 to disable log deduplication, or see https://docs.ray.io/en/master/ray-observability/user-guides/configure-logging.html#log-deduplication for more options.)[0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.) |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232400)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706902) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py:2046: FutureWarning: Tip: In future versions of Ray, Ray will no longer override accelerator visible devices env var if num_gpus=0 or num_gpus=None (default). To enable this behavior and turn off this error message, set RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0[32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706902) warnings.warn([32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706902) 2026-04-24 06:39:04,550 INFO worker.py:1680 -- Using address 10.128.33.33:6379 set in the environment variable RAY_ADDRESS[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706902) 2026-04-24 06:39:04,589 INFO worker.py:1821 -- Connecting to existing Ray cluster at address: 10.128.33.33:6379...[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706902) 2026-04-24 06:39:04,600 INFO worker.py:2007 -- Connected to Ray cluster.[32m [repeated 2x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2370206, ip=10.128.33.48)[0m [2026-04-24 06:39:04,629 E 2370206 2370406] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 15x across cluster][0m |
| [33m(raylet, ip=10.128.33.35)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 124x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) [33m(raylet, ip=10.128.33.37)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 74x across cluster][0m[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(pid=3868681)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 17x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(pid=3868681)[0m No module named 'vllm._version'[32m [repeated 17x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(pid=3868681)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 17x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [33m(raylet, ip=10.128.33.37)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 61x across cluster] (Ray deduplicates logs by default. Set RAY_DEDUP_LOGS=0 to disable log deduplication, or see https://docs.ray.io/en/master/ray-observability/user-guides/configure-logging.html#log-deduplication for more options.)[0m[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232400)[0m [W424 06:39:05.625726769 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [localhost]:46057 (errno: 97 - Address family not supported by protocol). |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=1702991, ip=10.128.33.39)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=2232721, ip=10.128.33.42)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) [36m(RayWorkerWrapper pid=1713192)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831854)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) [36m(RayWorkerWrapper pid=1713188)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=1714615)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706851)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m
Loading safetensors checkpoint shards: 0% Completed | 0/17 [00:00<?, ?it/s] |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706851)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831041, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831322) [36m(RayWorkerWrapper pid=847593)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=831041, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831322) [36m(RayWorkerWrapper pid=847595)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706852)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690574) [36m(RayWorkerWrapper pid=1695347)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690574) [36m(RayWorkerWrapper pid=1695344)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/ray/_private/worker.py:2046: FutureWarning: Tip: In future versions of Ray, Ray will no longer override accelerator visible devices env var if num_gpus=0 or num_gpus=None (default). To enable this behavior and turn off this error message, set RAY_ACCEL_ENV_VAR_OVERRIDE_ON_ZERO=0[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) warnings.warn([32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) 2026-04-24 06:39:08,057 INFO worker.py:1680 -- Using address 10.128.33.33:6379 set in the environment variable RAY_ADDRESS |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) 2026-04-24 06:39:08,091 INFO worker.py:1821 -- Connecting to existing Ray cluster at address: 10.128.33.33:6379... |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) 2026-04-24 06:39:08,150 INFO worker.py:2007 -- Connected to Ray cluster. |
| [36m(pid=2216258, ip=10.128.33.42)[0m [2026-04-24 06:39:09,372 E 2216258 2216998] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 32x across cluster][0m |
| [33m(raylet, ip=10.128.33.47)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 59x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [33m(raylet, ip=10.128.33.45)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 115x across cluster][0m[32m [repeated 13x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(pid=3868654)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 16x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(pid=3868654)[0m No module named 'vllm._version'[32m [repeated 16x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(pid=3868654)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 16x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [33m(raylet, ip=10.128.33.47)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 79x across cluster] (Ray deduplicates logs by default. Set RAY_DEDUP_LOGS=0 to disable log deduplication, or see https://docs.ray.io/en/master/ray-observability/user-guides/configure-logging.html#log-deduplication for more options.)[0m[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m
Loading safetensors checkpoint shards: 6% Completed | 1/17 [00:00<00:14, 1.10it/s] |
| [36m(AsyncVLLMInferenceEngine pid=831041, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831322) [36m(RayWorkerWrapper pid=847595)[0m [W424 06:39:10.290638141 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [localhost]:43101 (errno: 97 - Address family not supported by protocol).[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1989927, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990134) [36m(RayWorkerWrapper pid=1990697)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1989927, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990134) [36m(RayWorkerWrapper pid=1990726)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006410)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(pid=1690621, ip=10.128.33.36)[0m [2026-04-24 06:39:12,953 E 1690621 1691159] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1958270, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958353) [36m(RayWorkerWrapper pid=1958810)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1958270, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958353) [36m(RayWorkerWrapper pid=1958900)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=831041, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831322) [36m(RayWorkerWrapper pid=847593)[0m
Loading safetensors checkpoint shards: 0% Completed | 0/17 [00:00<?, ?it/s][32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706852)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)[32m [repeated 13x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m _C._set_float32_matmul_precision(precision) |
| [36m(pid=1690618, ip=10.128.33.36)[0m [2026-04-24 06:39:14,623 E 1690618 1691395] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 24x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955400, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955609) [33m(raylet, ip=10.128.33.47)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 20x across cluster][0m[32m [repeated 18x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(pid=1688119)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(pid=1688119)[0m No module named 'vllm._version'[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(pid=1688119)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [33m(raylet)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 23x across cluster] (Ray deduplicates logs by default. Set RAY_DEDUP_LOGS=0 to disable log deduplication, or see https://docs.ray.io/en/master/ray-observability/user-guides/configure-logging.html#log-deduplication for more options.)[0m[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1958143, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958349) [36m(RayWorkerWrapper pid=1974631)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1958143, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958349) [36m(RayWorkerWrapper pid=1974630)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690574) [36m(RayWorkerWrapper pid=1695344)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) [36m(RayWorkerWrapper pid=1713188)[0m
Loading safetensors checkpoint shards: 18% Completed | 3/17 [00:02<00:11, 1.25it/s][32m [repeated 19x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955400, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955609) [36m(RayWorkerWrapper pid=1971922)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1955400, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955609) [36m(RayWorkerWrapper pid=1971921)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1955400, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955609) [36m(RayWorkerWrapper pid=1971922)[0m [W424 06:39:15.223983834 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [localhost]:35471 (errno: 97 - Address family not supported by protocol).[32m [repeated 11x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378339)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=3867661, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867869) [36m(RayWorkerWrapper pid=3884173)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868681)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868654)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706902) [36m(RayWorkerWrapper pid=1723202)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=3867661, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867869) [36m(RayWorkerWrapper pid=3884172)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706902) [36m(RayWorkerWrapper pid=1723203)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723127)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723126)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(pid=1690622, ip=10.128.33.36)[0m [2026-04-24 06:39:14,014 E 1690622 1691285] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 24x across cluster][0m[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) [36m(RayWorkerWrapper pid=1994958)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) [36m(RayWorkerWrapper pid=1994853)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1958143, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958349) [36m(RayWorkerWrapper pid=1974630)[0m
Loading safetensors checkpoint shards: 0% Completed | 0/17 [00:00<?, ?it/s][32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)[32m [repeated 15x across cluster][0m |
| [36m(pid=2216278, ip=10.128.33.42)[0m [2026-04-24 06:39:19,748 E 2216278 2218720] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 36x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 2x across cluster][0m[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) [36m(RayWorkerWrapper pid=1994853)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 11x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994315) [36m(RayWorkerWrapper pid=2010626)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994315) [36m(RayWorkerWrapper pid=2010627)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m
Loading safetensors checkpoint shards: 6% Completed | 1/17 [00:00<00:13, 1.20it/s][32m [repeated 41x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m [W424 06:39:15.767252031 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-047-14-interconnect-1.jupiter.internal]:57433 (errno: 97 - Address family not supported by protocol).[32m [repeated 14x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688123)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(RayWorkerWrapper pid=1688122)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688118)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(RayWorkerWrapper pid=1688119)[0m _C._set_float32_matmul_precision(precision) |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(pid=2216267, ip=10.128.33.42)[0m [2026-04-24 06:39:19,024 E 2216267 2217592] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 15x across cluster][0m[32m [repeated 24x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994315) [36m(RayWorkerWrapper pid=2010627)[0m
Loading safetensors checkpoint shards: 0% Completed | 0/17 [00:00<?, ?it/s][32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) [36m(RayWorkerWrapper pid=1994853)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)[32m [repeated 11x across cluster][0m |
| [36m(pid=831488, ip=10.128.33.43)[0m [2026-04-24 06:39:24,758 E 831488 836125] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 683x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(RayWorkerWrapper pid=1688122)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868681)[0m
Loading safetensors checkpoint shards: 18% Completed | 3/17 [00:03<00:17, 1.26s/it][32m [repeated 71x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) [36m(RayWorkerWrapper pid=1994853)[0m [W424 06:39:19.101500674 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [localhost]:43851 (errno: 97 - Address family not supported by protocol).[32m [repeated 12x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994315) [36m(pid=1690648, ip=10.128.33.36)[0m [2026-04-24 06:39:23,976 E 1690648 1693692] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 475x across cluster][0m[32m [repeated 24x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(RayWorkerWrapper pid=1688122)[0m
Loading safetensors checkpoint shards: 0% Completed | 0/17 [00:00<?, ?it/s][32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(RayWorkerWrapper pid=1688119)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/__init__.py:1617: UserWarning: Please use the new API settings to control TF32 behavior, such as torch.backends.cudnn.conv.fp32_precision = 'tf32' or torch.backends.cuda.matmul.fp32_precision = 'ieee'. Old settings, e.g, torch.backends.cuda.matmul.allow_tf32 = True, torch.backends.cudnn.allow_tf32 = True, allowTF32CuDNN() and allowTF32CuBLAS() will be deprecated after Pytorch 2.9. Please see https://pytorch.org/docs/main/notes/cuda.html#tensorfloat-32-tf32-on-ampere-and-later-devices (Triggered internally at /pytorch/aten/src/ATen/Context.cpp:80.)[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m |
| [36m(pid=3867955, ip=10.128.33.34)[0m [2026-04-24 06:39:29,788 E 3867955 3869849] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 863x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1958270, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958353) [36m(RayWorkerWrapper pid=1958900)[0m
Loading safetensors checkpoint shards: 53% Completed | 9/17 [00:12<00:11, 1.43s/it][32m [repeated 86x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(RayWorkerWrapper pid=1688119)[0m [W424 06:39:22.606662621 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [localhost]:44931 (errno: 97 - Address family not supported by protocol).[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232400)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=1714615)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) [36m(RayWorkerWrapper pid=1713188)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(pid=1990366, ip=10.128.33.40)[0m [2026-04-24 06:39:28,878 E 1990366 1996374] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 822x across cluster][0m[32m [repeated 24x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831041, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831322) [36m(RayWorkerWrapper pid=847593)[0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m |
| [36m(pid=1671939, ip=10.128.33.47)[0m [2026-04-24 06:39:34,725 E 1671939 1673405] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 1315x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m
Loading safetensors checkpoint shards: 76% Completed | 13/17 [00:16<00:05, 1.28s/it][32m [repeated 80x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 15x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m 2026-04-24 06:39:38.219 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1115 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2 |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/inference_engines/vllm/vllm_engine.py:505: RuntimeWarning: coroutine 'AsyncLLM.collective_rpc' was never awaited |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m self.llm.collective_rpc("set_numa_affinity") |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m RuntimeWarning: Enable tracemalloc to get the object allocation traceback |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(pid=3868186)[0m [2026-04-24 06:39:33,895 E 3868186 3882738] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 1451x across cluster][0m[32m [repeated 24x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706852)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690574) [36m(RayWorkerWrapper pid=1695344)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m [2026-04-24 06:39:38,171 E 1671871 1688062] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 491x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706902) [36m(RayWorkerWrapper pid=1723202)[0m
Loading safetensors checkpoint shards: 65% Completed | 11/17 [00:17<00:09, 1.57s/it][32m [repeated 65x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1989927, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990134) [36m(RayWorkerWrapper pid=1990726)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006410)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1955400, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955609) [36m(RayWorkerWrapper pid=1971922)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378339)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1958143, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958349) [36m(RayWorkerWrapper pid=1974630)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1958270, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958353) [36m(RayWorkerWrapper pid=1958900)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690574) [36m(RayWorkerWrapper pid=1695347)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867661, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867869) [36m(RayWorkerWrapper pid=3884173)[0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868681)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m 2026-04-24 06:39:42.997 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1115 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/inference_engines/vllm/vllm_engine.py:505: RuntimeWarning: coroutine 'AsyncLLM.collective_rpc' was never awaited[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m self.llm.collective_rpc("set_numa_affinity")[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m RuntimeWarning: Enable tracemalloc to get the object allocation traceback[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690574) [36m(pid=1672140, ip=10.128.33.47)[0m [2026-04-24 06:39:38,136 E 1672140 1687932] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 609x across cluster][0m[32m [repeated 18x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(RayWorkerWrapper pid=1688122)[0m
Loading safetensors checkpoint shards: 76% Completed | 13/17 [00:18<00:05, 1.42s/it][32m [repeated 47x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) [36m(RayWorkerWrapper pid=1994853)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994315) [36m(RayWorkerWrapper pid=2010627)[0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868681)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 12x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m 2026-04-24 06:39:47.475 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1115 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/inference_engines/vllm/vllm_engine.py:505: RuntimeWarning: coroutine 'AsyncLLM.collective_rpc' was never awaited[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m self.llm.collective_rpc("set_numa_affinity")[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m RuntimeWarning: Enable tracemalloc to get the object allocation traceback[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706902) [36m(RayWorkerWrapper pid=1723202)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723127)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(RayWorkerWrapper pid=1688122)[0m
Loading safetensors checkpoint shards: 94% Completed | 16/17 [00:23<00:01, 1.56s/it][32m [repeated 18x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688118)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(RayWorkerWrapper pid=1688122)[0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688123)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m 2026-04-24 06:39:51.406 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1115 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/inference_engines/vllm/vllm_engine.py:505: RuntimeWarning: coroutine 'AsyncLLM.collective_rpc' was never awaited[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m self.llm.collective_rpc("set_numa_affinity")[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m RuntimeWarning: Enable tracemalloc to get the object allocation traceback[32m [repeated 10x across cluster][0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:39:56.943[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36mbuild_models[0m:[36m799[0m - [1m[32minit policy/ref/critic models done[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:39:56.944[0m | [1m[32mINFO [0m | [36mharbor.orchestrators.queue[0m:[36mstart[0m:[36m262[0m - [1m[32m[terminal_bench_generator:195] Started 256 workers (status every 120.0s, 0.75s launch grace period)[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:39:56.944[0m | [1m[32mINFO [0m | [36mexamples.terminal_bench.terminal_bench_generator[0m:[36m_create_orchestrator[0m:[36m216[0m - [1m[32mQueueOrchestrator created and started with n_concurrent_trials=256, rollback_hook registered for ContextLengthExceededError/AgentTimeoutError[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:39:56.944[0m | [1m[32mINFO [0m | [36mexamples.terminal_bench.terminal_bench_generator[0m:[36mstartup[0m:[36m185[0m - [1m[32mTerminalBenchGenerator startup complete. Shared orchestrator ready with n_concurrent_trials=256[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:39:56.944[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36mtrain[0m:[36m428[0m - [1m[32mGenerator startup complete[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:39:56.944[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m453[0m - [1m[32mStarted: 'load_checkpoints'[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:39:56.955[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36mload_checkpoints[0m:[36m1587[0m - [1m[32mLoading checkpoint from: /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints/global_step_18[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:39:56.956[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36mload_checkpoints[0m:[36m1593[0m - [1m[32mResuming from global_step: 18[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:39:56.964[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36mload_checkpoints[0m:[36m1609[0m - [1m[32mSuccessfully loaded trainer state[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:39:56.965[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36mload_checkpoints[0m:[36m1619[0m - [1m[32mSuccessfully loaded dataloader state[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:39:56.965[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36mload_checkpoints[0m:[36m1628[0m - [1m[32mLoading policy checkpoint from /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints/global_step_18/policy[0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(RayWorkerWrapper pid=1688122)[0m
Loading safetensors checkpoint shards: 100% Completed | 17/17 [00:24<00:00, 1.45s/it][32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(RayWorkerWrapper pid=1688122)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m 2026-04-24 06:40:00.204 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1115 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/inference_engines/vllm/vllm_engine.py:505: RuntimeWarning: coroutine 'AsyncLLM.collective_rpc' was never awaited[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m self.llm.collective_rpc("set_numa_affinity")[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m RuntimeWarning: Enable tracemalloc to get the object allocation traceback[32m [repeated 5x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/distributed_c10d.py:4876: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning. |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m warnings.warn( # warn only once |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m 2026-04-24 06:40:00.286 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:_create_engine:1115 - Initializing OpenAIServingChat with custom_chat_template read from: chat_templates/qwen3_thinking_acc.jinja2 |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/SkyRL/skyrl-train/skyrl_train/inference_engines/vllm/vllm_engine.py:505: RuntimeWarning: coroutine 'AsyncLLM.collective_rpc' was never awaited |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m self.llm.collective_rpc("set_numa_affinity") |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m RuntimeWarning: Enable tracemalloc to get the object allocation traceback |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:20.165[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36mload_checkpoints[0m:[36m1638[0m - [1m[32mSuccessfully loaded policy checkpoint[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:20.165[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36mload_checkpoints[0m:[36m1654[0m - [1m[32mSuccessfully loaded complete checkpoint state from global_step_18[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:20.165[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m455[0m - [1m[32mResumed training from global_step 18[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:20.166[0m | [33m[1mWARNING [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m460[0m - [33m[1mNo data consumption state found in checkpoint — resume may re-train on already-consumed data[0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash: |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m No module named 'vllm._version' |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m from .version import __version__, __version_tuple__ # isort:skip |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:20.166[0m | [33m[1mWARNING [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m480[0m - [33m[1mData consumption count mismatch on resume: expected 256, got 0. This can happen after epoch boundary transitions or error recovery.[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:20.166[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m453[0m - [1m[32mFinished: 'load_checkpoints', time cost: 23.22s[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:20.166[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m487[0m - [1m[32mStarted: 'init_weight_sync_state'[0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [32m2026-04-24 06:40:25.921[0m | [1m[32mINFO [0m | [36mlogging[0m:[36minfo[0m:[36m2216[0m - [1m[32m[weight-sync] Using master_addr=10.128.33.41, master_port=47639[0m |
| [36m(FSDPPolicyWorkerBase pid=2109999)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/vllm/__init__.py:7: RuntimeWarning: Failed to read commit hash:[32m [repeated 15x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2109999)[0m No module named 'vllm._version'[32m [repeated 15x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2109999)[0m from .version import __version__, __version_tuple__ # isort:skip[32m [repeated 15x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [32m2026-04-24 06:40:25.920[0m | [1m[32mINFO [0m | [36mlogging[0m:[36minfo[0m:[36m2216[0m - [1m[32m[weight-sync] get_node_ip_address()=10.128.33.41[0m[32m [repeated 3x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank0]:[W424 06:40:26.111312814 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-047-09-interconnect-1.jupiter.internal]:47639 (errno: 97 - Address family not supported by protocol). |
| [36m(AsyncVLLMInferenceEngine pid=1958143, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958349) [36m(RayWorkerWrapper pid=1974631)[0m 2026-04-24 06:40:26.009 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:init_weight_update_communicator:189 - torch.distributed.get_rank(): 1, rank_offset: 21, rank: 22, world_size: 49, group_name: skyrl |
| [36m(AsyncVLLMInferenceEngine pid=1958143, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958349) [36m(RayWorkerWrapper pid=1974631)[0m 2026-04-24 06:40:26.014 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:init_weight_update_communicator:200 - init_weight_update_communicator: master_address=10.128.33.41, master_port=47639, |
| [36m(AsyncVLLMInferenceEngine pid=1958143, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958349) [36m(RayWorkerWrapper pid=1974630)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 15x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(pid=1672140, ip=10.128.33.47)[0m [2026-04-24 06:39:38,136 E 1672140 1687932] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 746x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/distributed_c10d.py:4876: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning. |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m warnings.warn( # warn only once |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:26.264[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36minit_weight_sync_state[0m:[36m810[0m - [1m[32mInitialized weight sync state for policy model and inference engines.[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:26.265[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m487[0m - [1m[32mFinished: 'init_weight_sync_state', time cost: 6.10s[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:33.375[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m491[0m - [1m[32mFinished: 'sync_weights_to_inference_engines', time cost: 7.11s[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:33.377[0m | [1m[32mINFO [0m | [36mskyrl_train.callbacks.builtin[0m:[36mon_train_begin[0m:[36m263[0m - [1m[32mHFHubUploadCallback initialized: repo=laion/a2-rl-stack_jest_v2, upload_steps=5, export_path=/e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/exports[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:33.385[0m | [1m[32mINFO [0m | [36mskyrl_train.callbacks.builtin[0m:[36mon_train_begin[0m:[36m418[0m - [1m[32mDatabaseRegistrationCallback: Supabase credentials loaded[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m
Training Step Progress: 18it [00:00, ?it/s]
Training Step Progress: 18it [00:00, ?it/s] |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:26.265[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m491[0m - [1m[32mStarted: 'sync_weights_to_inference_engines'[0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) [36m(RayWorkerWrapper pid=1713192)[0m [rank1]:[W424 06:40:26.172605309 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-047-09.jupiter.internal]:47639 (errno: 97 - Address family not supported by protocol).[32m [repeated 24x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) [36m(RayWorkerWrapper pid=1713192)[0m 2026-04-24 06:40:26.010 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:init_weight_update_communicator:189 - torch.distributed.get_rank(): 1, rank_offset: 15, rank: 16, world_size: 49, group_name: skyrl[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688118)[0m 2026-04-24 06:40:26.137 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:init_weight_update_communicator:200 - init_weight_update_communicator: master_address=10.128.33.41, master_port=47639, [32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) [36m(RayWorkerWrapper pid=1713188)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 15x across cluster][0m[32m [repeated 22x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) [36m(pid=1672140, ip=10.128.33.47)[0m [2026-04-24 06:39:38,136 E 1672140 1687932] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 648x across cluster][0m[32m [repeated 5x across cluster][0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:33.591[0m | [33m[1mWARNING [0m | [36mskyrl_train.callbacks.builtin[0m:[36m_process_pending_uploads[0m:[36m317[0m - [33m[1mHFHubUploadCallback: Model path not found: /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/exports/global_step_19/policy[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:35.065[0m | [1m[32mINFO [0m | [36mskyrl_train.callbacks.builtin[0m:[36mon_train_end[0m:[36m536[0m - [1m[32mDatabaseRegistrationCallback: Registering model to database (agent=terminus-2, base_model=/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137, datasets=[['/e/scratch/jureap59/feuer1/tasks/exp_rpt_stack-jest-v2']])[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:37.221[0m | [1m[32mINFO [0m | [36mskyrl_train.callbacks.builtin[0m:[36mon_train_end[0m:[36m548[0m - [1m[32mDatabaseRegistrationCallback: Model 'laion/a2-rl-stack_jest_v2' already exists in database[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:40:37.221[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m757[0m - [1m[32mStarted: 'save_checkpoints'[0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m /e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/torch/distributed/distributed_c10d.py:4876: UserWarning: barrier(): using the device under current context. You can specify `device_id` in `init_process_group` to mute this warning. |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m warnings.warn( # warn only once |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:41:29.827[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36msave_checkpoints[0m:[36m1500[0m - [1m[32mSaved dataloader state to /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints/global_step_19/data.pt[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:41:29.836[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36msave_checkpoints[0m:[36m1512[0m - [1m[32mSaved trainer state to /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints/global_step_19/trainer_state.pt[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:41:29.837[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36msave_checkpoints[0m:[36m1519[0m - [1m[32mSuccessfully saved checkpoint for global_step_19 to: /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints/global_step_19[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:41:29.837[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36msave_checkpoints[0m:[36m1522[0m - [1m[32mStarted: 'cleanup_old_checkpoints'[0m |
| [33m(raylet, ip=10.128.33.38)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m 2026-04-24 06:40:26.007 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:init_weight_update_communicator:189 - torch.distributed.get_rank(): 0, rank_offset: 7, rank: 7, world_size: 49, group_name: skyrl |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m [rank0]:[W424 06:40:26.943492137 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-047-09.jupiter.internal]:47639 (errno: 97 - Address family not supported by protocol). |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m 2026-04-24 06:40:26.011 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:init_weight_update_communicator:200 - init_weight_update_communicator: master_address=10.128.33.41, master_port=47639, |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [33m(raylet, ip=10.128.33.38)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 2x across cluster][0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:41:38.104[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36msave_checkpoints[0m:[36m1522[0m - [1m[32mFinished: 'cleanup_old_checkpoints', time cost: 8.27s[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:41:38.105[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m759[0m - [1m[32mSaved final checkpoint.[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:41:38.106[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m757[0m - [1m[32mFinished: 'save_checkpoints', time cost: 60.88s[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:41:38.106[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m761[0m - [1m[32mStarted: 'save_hf_model'[0m |
| [33m(raylet, ip=10.128.33.41)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 36x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) [36m(RayWorkerWrapper pid=1713188)[0m 2026-04-24 06:40:26.009 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:init_weight_update_communicator:189 - torch.distributed.get_rank(): 0, rank_offset: 15, rank: 15, world_size: 49, group_name: skyrl[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) [36m(RayWorkerWrapper pid=1713188)[0m [rank0]:[W424 06:40:26.173113171 socket.cpp:767] [c10d] The client socket cannot be initialized to connect to [jpbo-047-09.jupiter.internal]:47639 (errno: 97 - Address family not supported by protocol).[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) [36m(RayWorkerWrapper pid=1713188)[0m 2026-04-24 06:40:26.014 | INFO | skyrl_train.inference_engines.vllm.vllm_engine:init_weight_update_communicator:200 - init_weight_update_communicator: master_address=10.128.33.41, master_port=47639, [32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) [33m(raylet, ip=10.128.33.38)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 2x across cluster][0m[32m [repeated 22x across cluster][0m |
| [36m(cleanup_old_checkpoints pid=2370673, ip=10.128.33.48)[0m [2026-04-24 06:42:00,752 E 2370673 2370747] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14 |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) [33m(raylet, ip=10.128.33.41)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 34x across cluster][0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:42:10.861[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36msave_models[0m:[36m1672[0m - [1m[32mSuccessfully saved model weights.[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:42:10.861[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m763[0m - [1m[32mSaved final model.[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:42:10.862[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m761[0m - [1m[32mFinished: 'save_hf_model', time cost: 32.76s[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:42:10.862[0m | [1m[32mINFO [0m | [36mskyrl_train.fully_async_trainer[0m:[36m_train_loop[0m:[36m764[0m - [1m[32mTraining done![0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(pid=2370672, ip=10.128.33.48)[0m [2026-04-24 06:42:00,801 E 2370672 2370775] core_worker_process.cc:842: Failed to establish connection to the metrics exporter agent. Metrics will not be exported. Exporter agent status: RpcError: Running out of retries to initialize the metrics agent. rpc_code: 14[32m [repeated 29x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [33m(raylet, ip=10.128.33.41)[0m [proxychains] DLL init: proxychains-ng 4.17-git-9-g78ead0e[32m [repeated 34x across cluster][0m[32m [repeated 23x across cluster][0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:42:11.868[0m | [1m[32mINFO [0m | [36mskyrl_train.inference_engines.inference_engine_client_http_endpoint[0m:[36mshutdown_server[0m:[36m203[0m - [1m[32mServer shut down after 2 seconds[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:42:11.868[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36m_guarded_sync[0m:[36m225[0m - [1m[32mHTTP endpoint shutdown complete[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:42:11.868[0m | [1m[32mINFO [0m | [36mexamples.terminal_bench.terminal_bench_generator[0m:[36mshutdown[0m:[36m287[0m - [1m[32mShutting down shared QueueOrchestrator...[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:42:11.871[0m | [1m[32mINFO [0m | [36mharbor.orchestrators.queue[0m:[36mshutdown[0m:[36m377[0m - [1m[32m[terminal_bench_generator:195] Shutdown complete. Total completed: 0[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:42:11.871[0m | [1m[32mINFO [0m | [36mexamples.terminal_bench.terminal_bench_generator[0m:[36mshutdown[0m:[36m289[0m - [1m[32mQueueOrchestrator shutdown complete[0m |
| [36m(skyrl_entrypoint pid=2109818)[0m [32m2026-04-24 06:42:11.871[0m | [1m[32mINFO [0m | [36mskyrl_train.trainer[0m:[36m_guarded_async[0m:[36m214[0m - [1m[32mGenerator shutdown complete[0m |
| 2026-04-24 06:42:12.288 | INFO | __main__:main:128 - Shutting down Ray on head node... |
| [36m(skyrl_entrypoint pid=2109818)[0m ⚙️ Running in WANDB offline mode |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m WARNING 04-24 06:38:19 [arg_utils.py:1256] The global random seed is set to 42. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM. |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m INFO 04-24 06:38:19 [model.py:529] Resolved architecture: Qwen3ForCausalLM |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m INFO 04-24 06:38:19 [model.py:1549] Using max model len 32768 |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m INFO 04-24 06:38:20 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}} |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m INFO 04-24 06:38:20 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=16384. |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m WARNING 04-24 06:38:20 [vllm.py:672] Async scheduling will be disabled because it is not supported with the `ray` distributed executor backend (only `mp`, `uni`, and `external_launcher` are supported). |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m INFO 04-24 06:38:20 [vllm.py:690] Asynchronous scheduling is disabled. |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m WARNING 04-24 06:38:20 [vllm.py:728] Enforce eager set, overriding optimization level to -O0 |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m INFO 04-24 06:38:20 [vllm.py:846] Cudagraph is disabled under eager mode |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m WARNING 04-24 06:38:20 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m WARNING 04-24 06:38:20 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m WARNING 04-24 06:38:24 [arg_utils.py:1256] The global random seed is set to 46. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m INFO 04-24 06:38:24 [model.py:529] Resolved architecture: Qwen3ForCausalLM[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m INFO 04-24 06:38:24 [model.py:1549] Using max model len 32768[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m INFO 04-24 06:38:25 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m INFO 04-24 06:38:25 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=16384.[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m WARNING 04-24 06:38:25 [vllm.py:672] Async scheduling will be disabled because it is not supported with the `ray` distributed executor backend (only `mp`, `uni`, and `external_launcher` are supported).[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m INFO 04-24 06:38:25 [vllm.py:690] Asynchronous scheduling is disabled.[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m WARNING 04-24 06:38:25 [vllm.py:728] Enforce eager set, overriding optimization level to -O0[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m INFO 04-24 06:38:25 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m WARNING 04-24 06:38:25 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m WARNING 04-24 06:38:25 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) INFO 04-24 06:38:28 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=43, served_model_name=9216db5781bf21249d130ec9da846c4624c16137, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [16384], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []} |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m WARNING 04-24 06:38:30 [arg_utils.py:1256] The global random seed is set to 61. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m INFO 04-24 06:38:30 [model.py:529] Resolved architecture: Qwen3ForCausalLM[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m INFO 04-24 06:38:30 [model.py:1549] Using max model len 32768[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m INFO 04-24 06:38:30 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m INFO 04-24 06:38:30 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=16384.[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m WARNING 04-24 06:38:30 [vllm.py:672] Async scheduling will be disabled because it is not supported with the `ray` distributed executor backend (only `mp`, `uni`, and `external_launcher` are supported).[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m INFO 04-24 06:38:30 [vllm.py:690] Asynchronous scheduling is disabled.[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m WARNING 04-24 06:38:30 [vllm.py:728] Enforce eager set, overriding optimization level to -O0[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m INFO 04-24 06:38:30 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m WARNING 04-24 06:38:30 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m WARNING 04-24 06:38:30 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned[32m [repeated 9x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) INFO 04-24 06:38:32 [ray_utils.py:469] Using the existing placement group |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) INFO 04-24 06:38:32 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=46, served_model_name=9216db5781bf21249d130ec9da846c4624c16137, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [16384], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}[32m [repeated 4x across cluster][0m |
| [36m(pid=4151080, ip=10.128.33.41)[0m ⚙️ Running in WANDB offline mode |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m WARNING 04-24 06:38:34 [arg_utils.py:1256] The global random seed is set to 62. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m INFO 04-24 06:38:34 [model.py:529] Resolved architecture: Qwen3ForCausalLM[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m INFO 04-24 06:38:34 [model.py:1549] Using max model len 32768[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m INFO 04-24 06:38:34 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m INFO 04-24 06:38:34 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=16384.[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m WARNING 04-24 06:38:34 [vllm.py:672] Async scheduling will be disabled because it is not supported with the `ray` distributed executor backend (only `mp`, `uni`, and `external_launcher` are supported).[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m INFO 04-24 06:38:34 [vllm.py:690] Asynchronous scheduling is disabled.[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m WARNING 04-24 06:38:34 [vllm.py:728] Enforce eager set, overriding optimization level to -O0[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m INFO 04-24 06:38:34 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m WARNING 04-24 06:38:35 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m WARNING 04-24 06:38:35 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1989927, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990134) INFO 04-24 06:38:35 [ray_utils.py:469] Using the existing placement group |
| [36m(AsyncVLLMInferenceEngine pid=1958270, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958353) INFO 04-24 06:38:38 [ray_utils.py:469] Using the existing placement group |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) INFO 04-24 06:38:37 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=60, served_model_name=9216db5781bf21249d130ec9da846c4624c16137, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [16384], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}[32m [repeated 7x across cluster][0m |
| [36m(pid=4151157, ip=10.128.33.41)[0m ⚙️ Running in WANDB offline mode |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m WARNING 04-24 06:38:37 [arg_utils.py:1256] The global random seed is set to 65. Since VLLM_ENABLE_V1_MULTIPROCESSING is set to False, this may affect the random state of the Python process that launched vLLM.[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m INFO 04-24 06:38:37 [model.py:529] Resolved architecture: Qwen3ForCausalLM[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m INFO 04-24 06:38:37 [model.py:1549] Using max model len 32768[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m INFO 04-24 06:38:37 [arg_utils.py:1469] Using ray runtime env (env vars redacted): {'env_vars': {'NCCL_CUMEM_ENABLE': '***', 'RAY_ADDRESS': '***', 'VLLM_ALLOW_INSECURE_SERIALIZATION': '***', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING': '***', 'VLLM_DISABLE_COMPILE_CACHE': '***', 'WANDB_API_KEY': '***'}}[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m INFO 04-24 06:38:37 [scheduler.py:224] Chunked prefill is enabled with max_num_batched_tokens=16384.[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m WARNING 04-24 06:38:37 [vllm.py:672] Async scheduling will be disabled because it is not supported with the `ray` distributed executor backend (only `mp`, `uni`, and `external_launcher` are supported).[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m INFO 04-24 06:38:37 [vllm.py:690] Asynchronous scheduling is disabled.[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m WARNING 04-24 06:38:37 [vllm.py:728] Enforce eager set, overriding optimization level to -O0[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m INFO 04-24 06:38:37 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m WARNING 04-24 06:38:37 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m WARNING 04-24 06:38:37 [system_utils.py:140] We must use the `spawn` multiprocessing start method. Overriding VLLM_WORKER_MULTIPROC_METHOD to 'spawn'. See https://docs.vllm.ai/en/latest/usage/troubleshooting.html#python-multiprocessing for more information. Reasons: In a Ray actor and can only be spawned[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) INFO 04-24 06:38:43 [ray_utils.py:469] Using the existing placement group[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) INFO 04-24 06:38:43 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=64, served_model_name=9216db5781bf21249d130ec9da846c4624c16137, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [16384], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}[32m [repeated 5x across cluster][0m |
| [36m(pid=2370205, ip=10.128.33.48)[0m ⚙️ Running in WANDB offline mode[32m [repeated 14x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) INFO 04-24 06:38:52 [ray_utils.py:469] Using the existing placement group |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) INFO 04-24 06:38:53 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=48, served_model_name=9216db5781bf21249d130ec9da846c4624c16137, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [16384], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []} |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) INFO 04-24 06:38:53 [ray_utils.py:469] Using the existing placement group |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) INFO 04-24 06:38:54 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=49, served_model_name=9216db5781bf21249d130ec9da846c4624c16137, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [16384], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) INFO 04-24 06:38:59 [ray_utils.py:469] Using the existing placement group[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) INFO 04-24 06:39:00 [ray_env.py:66] RAY_NON_CARRY_OVER_ENV_VARS from config: set() |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) INFO 04-24 06:39:00 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_ENABLE_V1_MULTIPROCESSING', 'HF_TOKEN', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'CUDA_HOME', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_CACHE_ROOT', 'VLLM_RAY_BUNDLE_INDICES', 'LD_LIBRARY_PATH', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_RAY_PER_WORKER_GPUS'] |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) INFO 04-24 06:39:00 [ray_env.py:74] If certain env vars should NOT be copied, add them to /e/home/jusers/feuer1/jupiter/.config/vllm/ray_non_carry_over_env_vars.json file |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) INFO 04-24 06:39:02 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_RAY_BUNDLE_INDICES', 'HF_TOKEN', 'VLLM_CACHE_ROOT', 'CUDA_HOME', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_DISABLE_COMPILE_CACHE', 'LD_LIBRARY_PATH', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'VLLM_ENABLE_V1_MULTIPROCESSING'] |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) INFO 04-24 06:39:03 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_DISABLE_COMPILE_CACHE', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'HF_TOKEN', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'VLLM_RAY_BUNDLE_INDICES', 'LD_LIBRARY_PATH', 'VLLM_CACHE_ROOT', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'CUDA_HOME'] |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) INFO 04-24 06:39:03 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_RAY_BUNDLE_INDICES', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_RAY_PER_WORKER_GPUS', 'HF_TOKEN', 'VLLM_CACHE_ROOT', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_WORKER_MULTIPROC_METHOD', 'LD_LIBRARY_PATH', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'CUDA_HOME', 'VLLM_ALLOW_INSECURE_SERIALIZATION'] |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) INFO 04-24 06:39:03 [ray_env.py:69] Copying the following environment variables to workers: ['LD_LIBRARY_PATH', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_CACHE_ROOT', 'CUDA_HOME', 'VLLM_USE_FLASHINFER_SAMPLER', 'HF_TOKEN', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'VLLM_ALLREDUCE_USE_SYMM_MEM'] |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) INFO 04-24 06:39:03 [ray_env.py:69] Copying the following environment variables to workers: ['HF_TOKEN', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'LD_LIBRARY_PATH', 'VLLM_CACHE_ROOT', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_RAY_PER_WORKER_GPUS', 'CUDA_HOME', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING'] |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690574) INFO 04-24 06:39:03 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_CACHE_ROOT', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_RAY_PER_WORKER_GPUS', 'HF_TOKEN', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'LD_LIBRARY_PATH', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'CUDA_HOME'] |
| [36m(AsyncVLLMInferenceEngine pid=831041, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831322) INFO 04-24 06:39:04 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_RAY_BUNDLE_INDICES', 'CUDA_HOME', 'HF_TOKEN', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'LD_LIBRARY_PATH', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_CACHE_ROOT'] |
| [36m(AsyncVLLMInferenceEngine pid=1989927, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990134) INFO 04-24 06:39:07 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'CUDA_HOME', 'VLLM_USE_FLASHINFER_SAMPLER', 'HF_TOKEN', 'VLLM_RAY_BUNDLE_INDICES', 'LD_LIBRARY_PATH', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'VLLM_CACHE_ROOT', 'VLLM_ALLOW_INSECURE_SERIALIZATION'] |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706902) INFO 04-24 06:39:04 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=56, served_model_name=9216db5781bf21249d130ec9da846c4624c16137, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [16384], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []}[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706902) INFO 04-24 06:39:04 [ray_utils.py:469] Using the existing placement group[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1989927, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990134) INFO 04-24 06:39:07 [ray_env.py:66] RAY_NON_CARRY_OVER_ENV_VARS from config: set()[32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1989927, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990134) INFO 04-24 06:39:07 [ray_env.py:74] If certain env vars should NOT be copied, add them to /e/home/jusers/feuer1/jupiter/.config/vllm/ray_non_carry_over_env_vars.json file[32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) INFO 04-24 06:39:07 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_RAY_PER_WORKER_GPUS', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'VLLM_CACHE_ROOT', 'LD_LIBRARY_PATH', 'HF_TOKEN', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'VLLM_WORKER_MULTIPROC_METHOD', 'CUDA_HOME', 'VLLM_RAY_BUNDLE_INDICES'] |
| [36m(AsyncVLLMInferenceEngine pid=1958270, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958353) INFO 04-24 06:39:09 [ray_env.py:69] Copying the following environment variables to workers: ['LD_LIBRARY_PATH', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_CACHE_ROOT', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'VLLM_RAY_PER_WORKER_GPUS', 'HF_TOKEN', 'VLLM_USE_FLASHINFER_SAMPLER', 'CUDA_HOME', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING'] |
| [36m(AsyncVLLMInferenceEngine pid=1958143, ip=10.128.33.44)[0m (EngineCore_DP0 pid=1958349) INFO 04-24 06:39:09 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_ALLOW_INSECURE_SERIALIZATION', 'HF_TOKEN', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_RAY_BUNDLE_INDICES', 'LD_LIBRARY_PATH', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'CUDA_HOME', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_CACHE_ROOT'] |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) INFO 04-24 06:39:09 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_ALLREDUCE_USE_SYMM_MEM', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'CUDA_HOME', 'LD_LIBRARY_PATH', 'VLLM_USE_FLASHINFER_SAMPLER', 'HF_TOKEN', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_CACHE_ROOT'] |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) INFO 04-24 06:39:10 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'CUDA_HOME', 'LD_LIBRARY_PATH', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_CACHE_ROOT', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'HF_TOKEN', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_DISABLE_COMPILE_CACHE'] |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) INFO 04-24 06:39:10 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_RAY_BUNDLE_INDICES', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_CACHE_ROOT', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'HF_TOKEN', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'LD_LIBRARY_PATH', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'CUDA_HOME', 'VLLM_WORKER_MULTIPROC_METHOD'] |
| [36m(AsyncVLLMInferenceEngine pid=1955400, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955609) INFO 04-24 06:39:10 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'LD_LIBRARY_PATH', 'VLLM_USE_FLASHINFER_SAMPLER', 'CUDA_HOME', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_RAY_PER_WORKER_GPUS', 'HF_TOKEN', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_CACHE_ROOT', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'VLLM_ENABLE_V1_MULTIPROCESSING'] |
| [36m(AsyncVLLMInferenceEngine pid=3867661, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867869) INFO 04-24 06:39:11 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_RAY_BUNDLE_INDICES', 'LD_LIBRARY_PATH', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_CACHE_ROOT', 'CUDA_HOME', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'HF_TOKEN', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_ALLOW_INSECURE_SERIALIZATION'] |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) INFO 04-24 06:39:11 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_ALLREDUCE_USE_SYMM_MEM', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_CACHE_ROOT', 'HF_TOKEN', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'LD_LIBRARY_PATH', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'CUDA_HOME'] |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) INFO 04-24 06:39:12 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_CACHE_ROOT', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'HF_TOKEN', 'VLLM_WORKER_MULTIPROC_METHOD', 'LD_LIBRARY_PATH', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'CUDA_HOME', 'VLLM_RAY_PER_WORKER_GPUS'] |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) INFO 04-24 06:39:08 [core.py:97] Initializing a V1 LLM engine (vdev) with config: model='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', speculative_config=None, tokenizer='/e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137', skip_tokenizer_init=False, tokenizer_mode=auto, revision=None, tokenizer_revision=None, trust_remote_code=True, dtype=torch.bfloat16, max_seq_len=32768, download_dir=None, load_format=auto, tensor_parallel_size=2, pipeline_parallel_size=1, data_parallel_size=1, disable_custom_all_reduce=False, quantization=None, enforce_eager=True, enable_return_routed_experts=False, kv_cache_dtype=auto, device_config=cuda, structured_outputs_config=StructuredOutputsConfig(backend='auto', disable_fallback=False, disable_any_whitespace=False, disable_additional_properties=False, reasoning_parser='', reasoning_parser_plugin='', enable_in_reasoning=False), observability_config=ObservabilityConfig(show_hidden_metrics_for_version=None, otlp_traces_endpoint=None, collect_detailed_traces=None, kv_cache_metrics=False, kv_cache_metrics_sample=0.01, cudagraph_metrics=False, enable_layerwise_nvtx_tracing=False, enable_mfu_metrics=False, enable_mm_processor_stats=False, enable_logging_iteration_details=False), seed=65, served_model_name=9216db5781bf21249d130ec9da846c4624c16137, enable_prefix_caching=True, enable_chunked_prefill=True, pooler_config=None, compilation_config={'level': None, 'mode': <CompilationMode.NONE: 0>, 'debug_dump_path': None, 'cache_dir': '', 'compile_cache_save_format': 'binary', 'backend': 'inductor', 'custom_ops': ['all'], 'splitting_ops': [], 'compile_mm_encoder': False, 'compile_sizes': [], 'compile_ranges_split_points': [16384], 'inductor_compile_config': {'enable_auto_functionalized_v2': False, 'combo_kernels': True, 'benchmark_combo_kernel': True}, 'inductor_passes': {}, 'cudagraph_mode': <CUDAGraphMode.NONE: 0>, 'cudagraph_num_of_warmups': 0, 'cudagraph_capture_sizes': [], 'cudagraph_copy_inputs': False, 'cudagraph_specialize_lora': True, 'use_inductor_graph_partition': False, 'pass_config': {'fuse_norm_quant': False, 'fuse_act_quant': False, 'fuse_attn_quant': False, 'eliminate_noops': False, 'enable_sp': False, 'fuse_gemm_comms': False, 'fuse_allreduce_rms': False, 'fuse_act_padding': False}, 'max_cudagraph_capture_size': 0, 'dynamic_shapes_config': {'type': <DynamicShapesType.BACKED: 'backed'>, 'evaluate_guards': False, 'assume_32_bit_indexing': False}, 'local_cache_dir': None, 'fast_moe_cold_start': True, 'static_all_moe_layers': []} |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) INFO 04-24 06:39:08 [ray_utils.py:469] Using the existing placement group[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) INFO 04-24 06:39:12 [ray_env.py:66] RAY_NON_CARRY_OVER_ENV_VARS from config: set()[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) INFO 04-24 06:39:12 [ray_env.py:74] If certain env vars should NOT be copied, add them to /e/home/jusers/feuer1/jupiter/.config/vllm/ray_non_carry_over_env_vars.json file[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706902) INFO 04-24 06:39:12 [ray_env.py:69] Copying the following environment variables to workers: ['CUDA_HOME', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'LD_LIBRARY_PATH', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'HF_TOKEN', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_CACHE_ROOT', 'VLLM_WORKER_MULTIPROC_METHOD'] |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) INFO 04-24 06:39:14 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_RAY_PER_WORKER_GPUS', 'VLLM_WORKER_MULTIPROC_METHOD', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'LD_LIBRARY_PATH', 'VLLM_USE_FLASHINFER_SAMPLER', 'CUDA_HOME', 'VLLM_DISABLE_COMPILE_CACHE', 'HF_TOKEN', 'VLLM_CACHE_ROOT', 'VLLM_ENABLE_V1_MULTIPROCESSING'] |
| [36m(AsyncVLLMInferenceEngine pid=1994236, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994315) INFO 04-24 06:39:15 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_DISABLE_COMPILE_CACHE', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_CACHE_ROOT', 'LD_LIBRARY_PATH', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'VLLM_WORKER_MULTIPROC_METHOD', 'HF_TOKEN', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'CUDA_HOME', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'VLLM_RAY_PER_WORKER_GPUS'] |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) INFO 04-24 06:39:16 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_DISABLE_COMPILE_CACHE', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_CACHE_ROOT', 'VLLM_ALLOW_INSECURE_SERIALIZATION', 'CUDA_HOME', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'VLLM_WORKER_MULTIPROC_METHOD', 'LD_LIBRARY_PATH', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'HF_TOKEN'] |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) INFO 04-24 06:39:16 [ray_env.py:69] Copying the following environment variables to workers: ['VLLM_ALLOW_INSECURE_SERIALIZATION', 'HF_TOKEN', 'VLLM_ALLREDUCE_USE_SYMM_MEM', 'LD_LIBRARY_PATH', 'VLLM_RAY_PER_WORKER_GPUS', 'VLLM_DISABLE_COMPILE_CACHE', 'VLLM_CACHE_ROOT', 'VLLM_USE_FLASHINFER_SAMPLER', 'VLLM_ALLOW_RUNTIME_LORA_UPDATING', 'VLLM_RAY_BUNDLE_INDICES', 'VLLM_ENABLE_V1_MULTIPROCESSING', 'CUDA_HOME', 'VLLM_WORKER_MULTIPROC_METHOD'] |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151152 [0] NCCL INFO NCCL_SOCKET_FAMILY set by environment to AF_INET |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151152 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to ib0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151152 [0] NCCL INFO Bootstrap: Using ib0:10.128.33.41<0> |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151152 [0] NCCL INFO cudaDriverVersion 13010 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151152 [0] NCCL INFO NCCL version 2.27.7+cuda13.0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151152 [0] NCCL INFO Comm config Blocking set to 1 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so. |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO NCCL_SOCKET_FAMILY set by environment to AF_INET |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to ib0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO NET/IB : Using [0]mlx5_0:1/IB [1]mlx5_1:1/IB [2]mlx5_2:1/IB [3]mlx5_3:1/IB [RO]; OOB ib0:10.128.33.41<0> |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Initialized NET plugin IB |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Assigned NET plugin IB to comm |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Using network IB |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO DMA-BUF is available on GPU device 0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO ncclCommInitRankConfig comm 0x400eb25f55e0 rank 0 nranks 16 cudaDev 0 nvmlDev 0 busId 901000 commId 0x993b2bbee9ef7c7e - Init START |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO RAS client listening socket at 127.0.0.1<28028> |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Bootstrap timings total 6.524683 (create 0.000027, send 0.000092, recv 0.000167, ring 6.522021, delay 0.000000) |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO NCCL_CUMEM_ENABLE set by environment to 0. |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO NCCL_NET_GDR_LEVEL set by environment to LOC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Setting affinity for GPU 0 to 0-71 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO NVLS multicast support is not available on dev 0 (NVLS_NCHANNELS 0) |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO comm 0x400eb25f55e0 rank 0 nRanks 16 nNodes 4 localRanks 4 localRank 0 MNNVL 0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 00/08 : 0 1 2 3 7 6 5 4 8 9 10 11 15 14 13 12 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 01/08 : 0 4 5 6 7 11 10 9 8 12 13 14 15 3 2 1 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 02/08 : 0 3 1 5 4 7 6 10 8 11 9 13 12 15 14 2 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 03/08 : 0 3 2 6 4 7 5 9 8 11 10 14 12 15 13 1 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 04/08 : 0 1 2 3 7 6 5 4 8 9 10 11 15 14 13 12 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 05/08 : 0 4 5 6 7 11 10 9 8 12 13 14 15 3 2 1 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 06/08 : 0 3 1 5 4 7 6 10 8 11 9 13 12 15 14 2 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 07/08 : 0 3 2 6 4 7 5 9 8 11 10 14 12 15 13 1 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Trees [0] 1/8/-1->0->-1 [1] -1/-1/-1->0->3 [2] 1/-1/-1->0->2 [3] 2/-1/-1->0->3 [4] 1/-1/-1->0->4 [5] -1/-1/-1->0->3 [6] 1/-1/-1->0->2 [7] 2/-1/-1->0->3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO P2P Chunksize set to 131072 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO PROFILER/Plugin: Could not find: libnccl-profiler.so. |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Check P2P Type isAllDirectP2p 1 directMode 0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151480 [0] NCCL INFO [Proxy Service] Device 0 CPU core 1 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151482 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 2 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 00/0 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 04/0 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 02/0 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 03/0 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 06/0 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 07/0 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151486 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 00/0 : 12[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 04/0 : 12[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 01/0 : 0[0] -> 4[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 05/0 : 0[0] -> 4[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151480 [0] NCCL INFO NCCL_IB_TIMEOUT set by environment to 60. |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Connected all rings, use ring PXN 0 GDR 0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 02/0 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 06/0 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 02/0 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 03/0 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 06/0 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 07/0 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 01/0 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 05/0 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 04/0 : 0[0] -> 4[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 00/0 : 8[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 00/0 : 0[0] -> 8[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Channel 04/0 : 4[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Connected all trees |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO threadThresholds 8/8/64 | 128/8/64 | 512 | 512 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO 8 coll channels, 8 collnet channels, 0 nvls channels, 8 p2p channels, 2 p2p channels per peer |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO CC Off, workFifoBytes 1048576 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO TUNER/Plugin: Could not find: libnccl-tuner.so. Using internal tuner plugin. |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO ncclCommInitRankConfig comm 0x400eb25f55e0 rank 0 nranks 16 cudaDev 0 nvmlDev 0 busId 901000 commId 0x993b2bbee9ef7c7e - Init COMPLETE |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151446 [0] NCCL INFO Init timings - ncclCommInitRankConfig: rank 0 nranks 16 total 8.70 (kernels 0.23, alloc 1.41, bootstrap 6.52, allgathers 0.02, topo 0.06, graphs 0.01, connections 0.45, rest 0.01) |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151152 [0] NCCL INFO Comm config Blocking set to 1 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Assigned NET plugin IB to comm |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Using network IB |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO DMA-BUF is available on GPU device 0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO ncclCommInitRankConfig comm 0x400eb12764b0 rank 0 nranks 4 cudaDev 0 nvmlDev 0 busId 901000 commId 0xfc2340fcc323d1bc - Init START |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Bootstrap timings total 0.014487 (create 0.000035, send 0.000075, recv 0.013798, ring 0.000376, delay 0.000000) |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Setting affinity for GPU 0 to 0-71 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO comm 0x400eb12764b0 rank 0 nRanks 4 nNodes 4 localRanks 1 localRank 0 MNNVL 0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 00/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 01/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 02/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 03/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 04/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 05/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 06/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 07/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 08/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 09/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 10/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 11/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 12/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 13/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 14/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 15/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 16/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 17/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 18/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 19/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 20/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 21/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 22/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 23/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 24/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 25/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 26/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 27/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 28/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 29/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 30/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 31/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 32/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 33/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 34/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 35/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 36/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 37/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 38/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 39/40 : 0 1 2 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Trees [0] 2/-1/-1->0->-1 [1] 2/-1/-1->0->-1 [2] 2/-1/-1->0->-1 [3] 2/-1/-1->0->-1 [4] 2/-1/-1->0->-1 [5] 2/-1/-1->0->-1 [6] 2/-1/-1->0->-1 [7] 2/-1/-1->0->-1 [8] 2/-1/-1->0->-1 [9] 2/-1/-1->0->-1 [10] 2/-1/-1->0->-1 [11] 2/-1/-1->0->-1 [12] 2/-1/-1->0->-1 [13] 2/-1/-1->0->-1 [14] 2/-1/-1->0->-1 [15] 2/-1/-1->0->-1 [16] 2/-1/-1->0->-1 [17] 2/-1/-1->0->-1 [18] 2/-1/-1->0->-1 [19] 2/-1/-1->0->-1 [20] -1/-1/-1->0->1 [21] -1/-1/-1->0->1 [22] -1/-1/-1->0->1 [23] -1/-1/-1->0->1 [24] -1/-1/-1->0->1 [25] -1/-1/-1->0->1 [26] -1/-1/-1->0->1 [27] -1/-1/-1->0->1 [28] -1/-1/-1->0->1 [29] -1/-1/-1->0->1 [30] -1/-1/-1->0->1 [31] -1/-1/-1->0->1 [32] -1/-1/-1->0->1 [33] -1/-1/-1->0->1 [34] -1/-1/-1->0->1 [35] -1/-1/-1->0->1 [36] -1/-1/-1->0->1 [37] -1/-1/-1->0->1 [38] -1/-1/-1->0->1 [39] -1/-1/-1->0->1 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO P2P Chunksize set to 131072 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Check P2P Type isAllDirectP2p 0 directMode 0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151502 [0] NCCL INFO [Proxy Service] Device 0 CPU core 2 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151503 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 4 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151507 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 5 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 00/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 01/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 02/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 03/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 04/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 05/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 06/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 07/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 08/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 09/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 10/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 11/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 12/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 13/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 14/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 15/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 16/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 17/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 18/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 19/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 20/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 21/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 22/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 23/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 24/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 25/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 26/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 27/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 28/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 29/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 30/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 31/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 32/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 33/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 34/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 35/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 36/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 37/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 38/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 39/0 : 3[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 00/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 01/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 02/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 03/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 04/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 05/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 06/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 07/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 08/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 09/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 10/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 11/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 12/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 13/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 14/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151491 [0] NCCL INFO Channel 15/0 : 0[0] -> 1[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:415149 |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) INFO 04-24 06:39:16 [ray_env.py:66] RAY_NON_CARRY_OVER_ENV_VARS from config: set()[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) INFO 04-24 06:39:16 [ray_env.py:74] If certain env vars should NOT be copied, add them to /e/home/jusers/feuer1/jupiter/.config/vllm/ray_non_carry_over_env_vars.json file[32m [repeated 5x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151157, ip=10.128.33.41)[0m j |
| [36m(FSDPPolicyWorkerBase pid=4151158, ip=10.128.33.41)[0m jpbo-047-09:4151158:4 |
| [36m(FSDPPolicyWorkerBase pid=2109996)[0m jpbo-047-01:2109996:2110423 [0] |
| [36m(FSDPPolicyWorkerBase pid=4151159, ip=10.128.33.41)[0m jpbo-047-09:4151159:4151493 [0] NCCL INFO Chan |
| [36m(FSDPPolicyWorkerBase pid=2109997)[0m jpbo-047-01:2109997:2110424 [0] NCCL INFO Channel 06/0 : 2[3] -> 3[3] [receive] via NE |
| [36m(FSDPPolicyWorkerBase pid=2109998)[0m jpbo-047-01:2109998:2110422 [0] NCCL INFO Channel 08/0 : 2[1] -> 3[1] [recei |
| [36m(FSDPPolicyWorkerBase pid=2370207, ip=10.128.33.48)[0m jpbo-047-16 |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232400)[0m WARNING 04-24 06:39:00 [system_utils.py:37] Overwriting environment variable CUDA_VISIBLE_DEVICES from '0,1,2,3' to '1,2' |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232400)[0m WARNING 04-24 06:39:00 [system_utils.py:37] Overwriting environment variable LD_LIBRARY_PATH from '/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib' to '/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib' |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m INFO 04-24 06:39:04 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'set_numa_affinity', 'test_rpc'] |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m WARNING 04-24 06:39:04 [worker_base.py:307] Missing `shared_worker_lock` argument from executor. This argument is needed for mm_processor_cache_type='shm'. |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232400)[0m INFO 04-24 06:39:05 [parallel_state.py:1234] world_size=2 rank=1 local_rank=1 distributed_init_method=tcp://127.0.0.1:46057 backend=nccl |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m INFO 04-24 06:39:06 [pynccl.py:111] vLLM is using nccl==2.27.7 |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m WARNING 04-24 06:39:00 [system_utils.py:37] Overwriting environment variable CUDA_VISIBLE_DEVICES from '0,1,2,3' to '1,2' |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m WARNING 04-24 06:39:00 [system_utils.py:37] Overwriting environment variable LD_LIBRARY_PATH from '/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib' to '/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib' |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232400)[0m INFO 04-24 06:39:07 [parallel_state.py:1445] rank 1 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 1, EP rank N/A |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m INFO 04-24 06:39:08 [gpu_model_runner.py:4125] Starting to load model /e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137... |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m INFO 04-24 06:39:08 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION']. |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m INFO 04-24 06:39:29 [default_loader.py:293] Loading weights took 20.51 seconds |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232400)[0m INFO 04-24 06:39:04 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'set_numa_affinity', 'test_rpc'] |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232400)[0m WARNING 04-24 06:39:04 [worker_base.py:307] Missing `shared_worker_lock` argument from executor. This argument is needed for mm_processor_cache_type='shm'. |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m INFO 04-24 06:39:05 [parallel_state.py:1234] world_size=2 rank=0 local_rank=0 distributed_init_method=tcp://127.0.0.1:46057 backend=nccl |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m INFO 04-24 06:39:07 [parallel_state.py:1445] rank 0 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m INFO 04-24 06:39:30 [gpu_model_runner.py:4222] Model loading took 30.59 GiB memory and 21.093942 seconds |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232399)[0m INFO 04-24 06:39:35 [gpu_worker.py:373] Available KV cache memory: 47.96 GiB |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) INFO 04-24 06:39:35 [kv_cache_utils.py:1307] GPU KV cache size: 392,864 tokens |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) INFO 04-24 06:39:35 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 11.99x |
| [36m(FSDPPolicyWorkerBase pid=2370205, ip=10.128.33.48)[0m jpbo-047-16:2370205:23 |
| [36m(FSDPPolicyWorkerBase pid=2417105, ip=10.128.33.38)[0m jpbo-047-06:2417105:2417504 [0] NCCL INFO Channel 05/0 : 0[1] -> 1[1] [receive] via NET/I |
| [36m(FSDPPolicyWorkerBase pid=2370204, ip=10.128.33.48)[0m jpbo-047-16:2370204:2370605 [0] |
| [36m(FSDPPolicyWorkerBase pid=2109997)[0m T/IB/3 |
| [36m(FSDPPolicyWorkerBase pid=2109997)[0m jpbo-047-01:2109997:2110424 [0] NCCL |
| [36m(FSDPPolicyWorkerBase pid=2109998)[0m ve] via NET/IB/1 |
| [36m(FSDPPolicyWorkerBase pid=2109998)[0m jpbo-047-01:2109998:2110422 |
| [36m(FSDPPolicyWorkerBase pid=2370205, ip=10.128.33.48)[0m jpbo-047-16:2370205:2370606 [0] NCCL INFO Channel 10/0 : 0[1] -> 2 |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m nel 04/0 : 0[0] -> 1[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m jpbo-047-06:2417106:2417506 [0] NCCL INFO Channel 32/0 : 3[0] -> 1[0] [receive] via NET/IB |
| [36m(FSDPPolicyWorkerBase pid=2370207, ip=10.128.33.48)[0m jpbo-047-16:2370207:2370607 [0] NCCL INFO Channel 14/0 |
| [36m(FSDPPolicyWorkerBase pid=2417105, ip=10.128.33.38)[0m B/1 |
| [36m(FSDPPolicyWorkerBase pid=2417105, ip=10.128.33.38)[0m jpbo-047-06:2417105:2417504 [0] NCCL INF |
| [36m(FSDPPolicyWorkerBase pid=2370206, ip=10.128.33.48)[0m 2] -> 2[2] [receive] via NET/IB/2 |
| [36m(FSDPPolicyWorkerBase pid=2370206, ip=10.128.33.48)[0m jpbo-047-1 |
| [36m(FSDPPolicyWorkerBase pid=2417107, ip=10.128.33.38)[0m jpbo-047-06:2417107:2417505 [0] NCCL INFO Ch |
| [36m(FSDPPolicyWorkerBase pid=2109999)[0m T/IB/2 |
| [36m(FSDPPolicyWorkerBase pid=2370204, ip=10.128.33.48)[0m jpbo-047-16:2370204:2370605 [0] NCCL INFO Channel 13/0 : 0[3] -> 2[3] [rece |
| [36m(FSDPPolicyWorkerBase pid=2370205, ip=10.128.33.48)[0m [1] [receive] via NET/IB/1 |
| [36m(FSDPPolicyWorkerBase pid=2370205, ip=10.128.33.48)[0m jpbo-0 |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m /0 |
| [36m(FSDPPolicyWorkerBase pid=2417105, ip=10.128.33.38)[0m O Channel 34/0 : 3[1] -> 1[1] [receive] via NET/IB/1 |
| [36m(FSDPPolicyWorkerBase pid=2417105, ip=10.128.33.38)[0m jpbo-047-06:2417105:24 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080: |
| [36m(FSDPPolicyWorkerBase pid=4151159, ip=10.128.33.41)[0m nel 25/0 : 0[1] -> 1[1] [send] via NET/IB/1 |
| [36m(FSDPPolicyWorkerBase pid=4151159, ip=10.128.33.41)[0m jpbo-047-09:4 |
| [36m(FSDPPolicyWorkerBase pid=2109998)[0m jpbo-047-01:210999 |
| [36m(FSDPPolicyWorkerBase pid=2370207, ip=10.128.33.48)[0m : 0[0] -> 2[0] [receive] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=2370207, ip=10.128.33.48)[0m jpbo-047-16:2370207:2370607 [0] NCCL INFO Channel 23/0 : 3[0] -> 2[0] [receive] via NET/ |
| [36m(FSDPPolicyWorkerBase pid=2370206, ip=10.128.33.48)[0m jpbo-047-16:2370206:2370604 [0] NCCL INFO C |
| [36m(FSDPPolicyWorkerBase pid=2370204, ip=10.128.33.48)[0m ive] via NET/IB/3 |
| [36m(FSDPPolicyWorkerBase pid=2370204, ip=10.128.33.48)[0m jpbo-047-16:237 |
| [36m(FSDPPolicyWorkerBase pid=2417107, ip=10.128.33.38)[0m annel 32/0 : 3[3] -> 1[3] [receive] via NET/IB/3 |
| [36m(FSDPPolicyWorkerBase pid=2417107, ip=10.128.33.38)[0m jpbo-047-06:2417107: |
| [36m(FSDPPolicyWorkerBase pid=2417108, ip=10.128.33.38)[0m jpbo-047-06:2417108:241750 |
| [36m(FSDPPolicyWorkerBase pid=2109997)[0m INFO Channel 35/0 : 1[3] -> 3[3] [receive] via NET/IB/3 |
| [36m(FSDPPolicyWorkerBase pid=2109997)[0m jpbo-047-01:2109997:21 |
| [36m(FSDPPolicyWorkerBase pid=2109999)[0m jpbo-047-01:2109999:21 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo |
| [36m(FSDPPolicyWorkerBase pid=4151159, ip=10.128.33.41)[0m jpbo-047-09:4151159:4151493 [0] NCCL INFO Connected binomial trees |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) [36m(RayWorkerWrapper pid=2232400)[0m INFO 04-24 06:39:35 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled. |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) INFO 04-24 06:39:36 [core.py:278] init engine (profile, create kv cache, warmup model) took 6.29 seconds |
| [36m(FSDPPolicyWorkerBase pid=4151157, ip=10.128.33.41)[0m jpbo-047-09:4151157:4151495 [0] NCCL INFO Init timings - ncclCommI |
| [36m(FSDPPolicyWorkerBase pid=4151158, ip=10.128.33.41)[0m jpbo-047-09:4151158:4151489 [0] NCCL INFO ncclCommInitRankConfig comm 0x400e414f5f0 |
| [36m(FSDPPolicyWorkerBase pid=2109996)[0m : 3[0] -> 2[0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=2109996)[0m jpbo-047-01:2109996:2110438 [0] NCCL IN |
| [36m(FSDPPolicyWorkerBase pid=2370207, ip=10.128.33.48)[0m IB/0 |
| [36m(FSDPPolicyWorkerBase pid=2370207, ip=10.128.33.48)[0m jpbo-047-16:2370207:2370621 [0] NCCL INFO Cha |
| [36m(FSDPPolicyWorkerBase pid=2370206, ip=10.128.33.48)[0m hannel 23/0 : 3[2] -> 2[2] [receive] via NET/IB/2 |
| [36m(FSDPPolicyWorkerBase pid=2370206, ip=10.128.33.48)[0m jpbo-047-16:2370206:2370624 [0] NCCL INFO Channel 20/0 : 2[2] -> 3[3] v |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m [0] [send] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m jpbo-047-06:2417106:241752 |
| [36m(FSDPPolicyWorkerBase pid=2370204, ip=10.128.33.48)[0m jpbo-047-16:2370204:2370623 [0] NCC |
| [36m(FSDPPolicyWorkerBase pid=2417105, ip=10.128.33.38)[0m jpbo-047-06:2417105:24 |
| [36m(FSDPPolicyWorkerBase pid=2417107, ip=10.128.33.38)[0m jpbo-047-06:2417107:2417522 [0] NCCL INFO Channel 15/0 : 3[3] -> 0[0] via P |
| [36m(FSDPPolicyWorkerBase pid=2417108, ip=10.128.33.38)[0m jpbo-047-06:2417108:241 |
| [36m(FSDPPolicyWorkerBase pid=2109997)[0m jpbo-047-01:2109997 |
| [36m(FSDPPolicyWorkerBase pid=2109998)[0m jpbo-047-01:2109998:2110439 [0] NCCL INFO |
| [36m(FSDPPolicyWorkerBase pid=2109999)[0m jpbo-047-01:2109999:2110441 |
| [36m(FSDPPolicyWorkerBase pid=2370207, ip=10.128.33.48)[0m nnel 21/24 : 0 2 1 3 |
| [36m(FSDPPolicyWorkerBase pid=2109996)[0m FO Channel 20/24 : 0 2 3 1 |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) WARNING 04-24 06:39:38 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) INFO 04-24 06:39:38 [vllm.py:690] Asynchronous scheduling is disabled. |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) WARNING 04-24 06:39:38 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored. |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216202) INFO 04-24 06:39:38 [vllm.py:846] Cudagraph is disabled under eager mode |
| [36m(AsyncVLLMInferenceEngine pid=2215993, ip=10.128.33.42)[0m WARNING 04-24 06:39:38 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`. |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=2232721, ip=10.128.33.42)[0m jpbo-047-10:2232721:2232721 [0] NCCL INFO ncclCommInitRank comm 0xaaaaf132b390 rank 1 nranks 2 cudaDev 0 nvmlDev 0 busId 901000 commId 0xd5e8ac5e065ebbfe - Init START |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=2232721, ip=10.128.33.42)[0m jpbo-047-1 |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=1714615)[0m jp |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=1714615)[0m WARNING 04-24 06:39:10 [custom_all_reduce.py:92] Custom allreduce is disabled because this process group spans across nodes. |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=1714615)[0m jpbo-047-07:1714615:1714615 [0] NCCL INFO NCCL_SOCKET_FAMILY set by environment to AF_INET[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=1714615)[0m jpbo-047-07:1714615:1714615 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to ib0[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=1714615)[0m jpbo-047-07:1714615:1714615 [0] NCCL INFO ncclCommInitRank comm 0xaaab0b49e600 rank 0 nranks 2 cudaDev 0 nvmlDev 0 busId 901000 commId 0xd5e8ac5e065ebbfe - Init START |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=1714615)[0m jpbo-047-07:1714615:1714615 [0] NCCL INFO Channel 26/0 : 1[0] -> 0[0] [receive] via NET/IB/0[32m [repeated 27x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=2232721, ip=10.128.33.42)[0m WARNING 04-24 06:39:10 [custom_all_reduce.py:92] Custom allreduce is disabled because this process group spans across nodes. |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=2232721, ip=10.128.33.42)[0m jpbo-047-10:2232721:2233406 [0] NCCL INFO Channel 27/0 : 1[0] -> 0[0] [send] via NET/ |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=2232721, ip=10.128.33.42)[0m jpbo-047-10:2232721:2233406 [0] NCCL INFO Channel 39/0 : 0[0] -> 1[0] [receive] via NET/IB/0[32m [repeated 40x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=1702991, ip=10.128.33.39)[0m jpbo-04 |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NC |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=1702991, ip=10.128.33.39)[0m jpbo-047-07:1702991:1715538 [0] NCCL INFO Channel 27/0 : 1[1] -> 0[3] [send] via N |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=1702991, ip=10.128.33.39)[0m jpbo-047-07:1702991:1702991 [0] NCCL INFO NCCL_SOCKET_FAMILY set by environment to AF_INET[32m [repeated 34x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=1702991, ip=10.128.33.39)[0m jpbo-047-07:1702991:1702991 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to ib0[32m [repeated 34x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO Bootstrap: Using ib0:10.128.33.42<0>[32m [repeated 19x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO cudaDriverVersion 13010[32m [repeated 19x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO NCCL version 2.27.7+cuda13.0[32m [repeated 19x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO Comm config Blocking set to 1[32m [repeated 45x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so. [32m [repeated 19x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO NET/IB : Using [0]mlx5_0:1/IB [1]mlx5_1:1/IB [2]mlx5_2:1/IB [3]mlx5_3:1/IB [RO]; OOB ib0:10.128.33.42<0>[32m [repeated 19x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO Initialized NET plugin IB[32m [repeated 19x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO Assigned NET plugin IB to comm[32m [repeated 49x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO Using network IB[32m [repeated 49x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO DMA-BUF is available on GPU device 0[32m [repeated 49x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO ncclCommInitRankConfig comm 0xaaaaf7609700 rank 0 nranks 2 cudaDev 0 nvmlDev 3 busId 3901000 commId 0x7a3727cb8540b9c2 - Init START[32m [repeated 45x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO RAS client listening socket at 127.0.0.1<28028>[32m [repeated 19x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO Bootstrap timings total 0.002699 (create 0.000028, send 0.000090, recv 0.002113, ring 0.000051, delay 0.000000)[32m [repeated 49x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO NCCL_CUMEM_ENABLE set by environment to 0.[32m [repeated 19x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO NCCL_NET_GDR_LEVEL set by environment to LOC[32m [repeated 19x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO Setting affinity for GPU 3 to 216-287[32m [repeated 49x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2109999)[0m jpbo-047-01:2109999:2110441 [0] NCCL INFO NVLS multicast support is not available on dev 0 (NVLS_NCHANNELS 0)[32m [repeated 28x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO comm 0xaaaaf7609700 rank 0 nRanks 2 nNodes 2 localRanks 1 localRank 0 MNNVL 0[32m [repeated 49x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO Channel 39/40 : 0 1[32m [repeated 376x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO Trees [0] 1/-1/-1->0->-1 [1] 1/-1/-1->0->-1 [2] 1/-1/-1->0->-1 [3] 1/-1/-1->0->-1 [4] 1/-1/-1->0->-1 [5] 1/-1/-1->0->-1 [6] 1/-1/-1->0->-1 [7] 1/-1/-1->0->-1 [8] 1/-1/-1->0->-1 [9] 1/-1/-1->0->-1 [10] 1/-1/-1->0->-1 [11] 1/-1/-1->0->-1 [12] 1/-1/-1->0->-1 [13] 1/-1/-1->0->-1 [14] 1/-1/-1->0->-1 [15] 1/-1/-1->0->-1 [16] 1/-1/-1->0->-1 [17] 1/-1/-1->0->-1 [18] 1/-1/-1->0->-1 [19] 1/-1/-1->0->-1 [20] -1/-1/-1->0->1 [21] -1/-1/-1->0->1 [22] -1/-1/-1->0->1 [23] -1/-1/-1->0->1 [24] -1/-1/-1->0->1 [25] -1/-1/-1->0->1 [26] -1/-1/-1->0->1 [27] -1/-1/-1->0->1 [28] -1/-1/-1->0->1 [29] -1/-1/-1->0->1 [30] -1/-1/-1->0->1 [31] -1/-1/-1->0->1 [32] -1/-1/-1->0->1 [33] -1/-1/-1->0->1 [34] -1/-1/-1->0->1 [35] -1/-1/-1->0->1 [36] -1/-1/-1->0->1 [37] -1/-1/-1->0->1 [38] -1/-1/-1->0->1 [39] -1/-1/-1->0->1[32m [repeated 49x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO P2P Chunksize set to 131072[32m [repeated 49x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO PROFILER/Plugin: Could not find: libnccl-profiler.so. [32m [repeated 19x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO Check P2P Type isAllDirectP2p 0 directMode 0[32m [repeated 11x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233409 [0] NCCL INFO [Proxy Service] Device 0 CPU core 245[32m [repeated 49x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233410 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 247[32m [repeated 49x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m jpbo-047-06:2417106:2417539 [0] NCCL INFO Channel 19/1 : 0[0] -> 2[2] via P2P/IPC[32m [repeated 545x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233411 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 278[32m [repeated 39x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=1702991, ip=10.128.33.39)[0m jpbo-047-07:1702991:1702991 [0] NCCL INFO Channel 39/0 : 0[3] -> 1[1] [receive] via NET/IB/1[32m [repeated 2012x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=1702991, ip=10.128.33.39)[0m jpbo-047-07:1702991:1715538 [0] NCCL INFO Channel 26/0 : 1[1] -> 0[3] [send] via NET/IB/1[32m [repeated 2057x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417108, ip=10.128.33.38)[0m jpbo-047-06:2417108:2417491 [0] NCCL INFO NCCL_IB_TIMEOUT set by environment to 60.[32m [repeated 15x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m jpbo-047-06:2417106:2417520 [0] NCCL INFO Connected all rings, use ring PXN 0 GDR 1[32m [repeated 35x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m jpbo-047-06:2417106:2417520 [0] NCCL INFO Connected all trees[32m [repeated 34x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m jpbo-047-06:2417106:2417520 [0] NCCL INFO threadThresholds 8/8/64 | 32/8/64 | 512 | 512[32m [repeated 34x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m jpbo-047-06:2417106:2417520 [0] NCCL INFO 24 coll channels, 24 collnet channels, 0 nvls channels, 32 p2p channels, 16 p2p channels per peer[32m [repeated 34x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m jpbo-047-06:2417106:2417520 [0] NCCL INFO CC Off, workFifoBytes 1048576[32m [repeated 6x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417108, ip=10.128.33.38)[0m jpbo-047-06:2417108:2417463 [0] NCCL INFO TUNER/Plugin: Could not find: libnccl-tuner.so. Using internal tuner plugin.[32m [repeated 15x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m jpbo-047-06:2417106:2417520 [0] NCCL INFO ncclCommInitRankConfig comm 0x400e427d4340 rank 0 nranks 4 cudaDev 0 nvmlDev 0 busId 901000 commId 0xd132ae93804c20f3 - Init COMPLETE[32m [repeated 32x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m jpbo-047-06:2417106:2417520 [0] NCCL INFO Init timings - ncclCommInitRankConfig: rank 0 nranks 4 total 0.34 (kernels 0.00, alloc 0.00, bootstrap 0.17, allgathers 0.00, topo 0.02, graphs 0.01, connections 0.14, rest 0.00)[32m [repeated 31x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m jpbo-047-06:2417106:2417506 [0] NCCL INFO Chan |
| [36m(FSDPPolicyWorkerBase pid=2109999)[0m jpbo-047-01:2109999:2110421 [0] NCCL INFO Channel 06/0 : 2[2] -> 3[2] [receive] via NE |
| [36m(FSDPPolicyWorkerBase pid=2109996)[0m jpbo-047-01:2109996:2110423 [0] NCCL INFO Channel 33/0 : 1[0] -> 3[0] [recei |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m WARNING 04-24 06:39:03 [system_utils.py:37] Overwriting environment variable CUDA_VISIBLE_DEVICES from '0,1,2,3' to '2,3'[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m WARNING 04-24 06:39:03 [system_utils.py:37] Overwriting environment variable LD_LIBRARY_PATH from '/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib' to '/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib'[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m INFO 04-24 06:39:08 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'set_numa_affinity', 'test_rpc'][32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m WARNING 04-24 06:39:08 [worker_base.py:307] Missing `shared_worker_lock` argument from executor. This argument is needed for mm_processor_cache_type='shm'.[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m INFO 04-24 06:39:08 [parallel_state.py:1234] world_size=2 rank=0 local_rank=0 distributed_init_method=tcp: |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m INFO 04-24 06:39:09 [pynccl.py:111] vLLM is using nccl==2.27.7[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m INFO 04-24 06:39:10 [parallel_state.py:1445] rank 0 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 0, EP rank N/A[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m INFO 04-24 06:39:10 [gpu_model_runner.py:4125] Starting to load model /e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137...[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m INFO 04-24 06:39:11 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m INFO 04-24 06:39:33 [default_loader.py:293] Loading weights took 21.74 seconds[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831914)[0m INFO 04-24 06:39:34 [gpu_model_runner.py:4222] Model loading took 30.59 GiB memory and 22.816749 seconds[32m [repeated 7x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m INFO 04-24 06:39:39 [gpu_worker.py:373] Available KV cache memory: 47.99 GiB[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) INFO 04-24 06:39:40 [kv_cache_utils.py:1307] GPU KV cache size: 392,864 tokens[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) INFO 04-24 06:39:40 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 11.99x[32m [repeated 5x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151509 [0] NCCL [32m [repeated 2x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2109996)[0m ve] via NET/IB/0 |
| [36m(FSDPPolicyWorkerBase pid=2109996)[0m jpbo-047-01:2109996:2110456 [0] NCCL INFO Channel 23/1 : 0[0] -> 2[32m [repeated 2x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151157, ip=10.128.33.41)[0m jpbo-047-09:4151157:4151495 [0] NCCL INFO Channel[32m [repeated 2x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417108, ip=10.128.33.38)[0m jpbo-047-06:2417108:2417503 [0] NCCL INFO Ch |
| [36m(FSDPPolicyWorkerBase pid=4151158, ip=10.128.33.41)[0m jpbo-047-09:4151158:4151489 [0] NCCL INFO Channel 14/0 : 1[2] -> 0[2] [rece |
| [36m(FSDPPolicyWorkerBase pid=4151157, ip=10.128.33.41)[0m 16/0 : 1[3] -> 0[3] [receive] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO Channel 26/0 : 1[1] -> 0[3] [receive] via NET/ |
| [36m(FSDPPolicyWorkerBase pid=4151158, ip=10.128.33.41)[0m ive] via NET/IB/2 |
| [36m(FSDPPolicyWorkerBase pid=2417108, ip=10.128.33.38)[0m annel 34/0 : 3[2] -> 1[2] [receive] via NET/IB/2 |
| [36m(FSDPPolicyWorkerBase pid=2109999)[0m INFO Channel 35/0 : 1[2] -> 3[2] [receive] via NET/IB/2 |
| [36m(FSDPPolicyWorkerBase pid=2370205, ip=10.128.33.48)[0m jpbo |
| [36m(FSDPPolicyWorkerBase pid=2109999)[0m jpbo-047-01:2109999:2110421 [0] NCCL INFO Connected binomial trees[32m [repeated 15x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) [36m(RayWorkerWrapper pid=831854)[0m INFO 04-24 06:39:40 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=831169, ip=10.128.33.43)[0m (EngineCore_DP0 pid=831330) INFO 04-24 06:39:41 [core.py:278] init engine (profile, create kv cache, warmup model) took 6.49 seconds[32m [repeated 5x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=2417106, ip=10.128.33.38)[0m jpbo-047-06:2417106:2417539 [0] NCCL INFO Channel 21/1 : 0[0] -> 2[2] v |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=1714615)[0m jpbo-047-07:1714615:1715537 [0] NCCL INFO |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) WARNING 04-24 06:39:42 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 6x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) INFO 04-24 06:39:42 [vllm.py:690] Asynchronous scheduling is disabled.[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) WARNING 04-24 06:39:42 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698330) INFO 04-24 06:39:42 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698245, ip=10.128.33.39)[0m WARNING 04-24 06:39:42 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.[32m [repeated 5x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO ncclCommInitRank comm 0xaaaadca8ea30 rank 0 nranks 2 cudaDev 0 nvmlDev 3 busId 3901000 commId 0x94afb523beebf273 - Init START[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m WARNING 04-24 06:39:08 [custom_all_reduce.py:92] Custom allreduce is disabled because this process group spans across nodes.[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO NCCL_SOCKET_FAMILY set by environment to AF_INET[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2232505 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to ib0[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO Channel 38/0 : 1[1] -> 0[3] [receive] via NET/IB/3[32m [repeated 39x across cluster][0m[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706851)[0m WARNING 04-24 06:39:03 [system_utils.py:37] Overwriting environment variable CUDA_VISIBLE_DEVICES from '0,1,2,3' to '0,1'[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706851)[0m WARNING 04-24 06:39:03 [system_utils.py:37] Overwriting environment variable LD_LIBRARY_PATH from '/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib' to '/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib'[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706852)[0m INFO 04-24 06:39:09 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'set_numa_affinity', 'test_rpc'][32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706852)[0m WARNING 04-24 06:39:09 [worker_base.py:307] Missing `shared_worker_lock` argument from executor. This argument is needed for mm_processor_cache_type='shm'.[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706851)[0m INFO 04-24 06:39:10 [parallel_state.py:1234] world_size=2 rank=1 local_rank=1 distributed_init_method=tcp://127.0.0.1:39121 backend=nccl[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706852)[0m INFO 04-24 06:39:10 [pynccl.py:111] vLLM is using nccl==2.27.7[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706851)[0m INFO 04-24 06:39:12 [parallel_state.py:1445] rank 1 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 1, EP rank N/A[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706852)[0m INFO 04-24 06:39:13 [gpu_model_runner.py:4125] Starting to load model /e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137...[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706852)[0m INFO 04-24 06:39:14 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706852)[0m INFO 04-24 06:39:38 [default_loader.py:293] Loading weights took 24.36 seconds[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706852)[0m INFO 04-24 06:39:39 [gpu_model_runner.py:4222] Model loading took 30.59 GiB memory and 25.477003 seconds[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) [36m(RayWorkerWrapper pid=1706852)[0m INFO 04-24 06:39:44 [gpu_worker.py:373] Available KV cache memory: 47.96 GiB[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) INFO 04-24 06:39:44 [kv_cache_utils.py:1307] GPU KV cache size: 392,864 tokens[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690356, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690566) INFO 04-24 06:39:44 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 11.99x[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO NCCL_SOCKET_FAMILY set by environment to AF_INET |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to ib0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Bootstrap: Using ib0:10.128.33.37<0> |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO cudaDriverVersion 13010 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO NCCL version 2.27.7+cuda13.0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so. |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO NCCL_SOCKET_FAMILY set by environment to AF_INET |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to ib0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO NET/IB : Using [0]mlx5_0:1/IB [1]mlx5_1:1/IB [2]mlx5_2:1/IB [3]mlx5_3:1/IB [RO]; OOB ib0:10.128.33.37<0> |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Initialized NET plugin IB |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Assigned NET plugin IB to comm |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Using network IB |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO DMA-BUF is available on GPU device 0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO RAS client listening socket at 127.0.0.1<28028> |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Bootstrap timings total 0.002899 (create 0.000050, send 0.000101, recv 0.000218, ring 0.000397, delay 0.000001) |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO NCCL_CUMEM_ENABLE set by environment to 0. |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO NCCL_NET_GDR_LEVEL set by environment to LOC |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Setting affinity for GPU 0 to 0-71 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO comm 0xaaab12a1dd60 rank 0 nRanks 2 nNodes 2 localRanks 1 localRank 0 MNNVL 0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 00/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 01/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 02/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 03/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 04/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 05/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 06/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 07/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 08/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 09/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 10/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 11/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 12/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 13/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 14/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 15/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 16/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 17/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 18/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 19/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 20/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 21/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 22/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 23/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 24/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 25/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 26/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 27/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 28/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 29/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 30/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 31/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 32/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 33/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 34/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 35/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 36/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 37/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 38/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 39/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Trees [0] 1/-1/-1->0->-1 [1] 1/-1/-1->0->-1 [2] 1/-1/-1->0->-1 [3] 1/-1/-1->0->-1 [4] 1/-1/-1->0->-1 [5] 1/-1/-1->0->-1 [6] 1/-1/-1->0->-1 [7] 1/-1/-1->0->-1 [8] 1/-1/-1->0->-1 [9] 1/-1/-1->0->-1 [10] 1/-1/-1->0->-1 [11] 1/-1/-1->0->-1 [12] 1/-1/-1->0->-1 [13] 1/-1/-1->0->-1 [14] 1/-1/-1->0->-1 [15] 1/-1/-1->0->-1 [16] 1/-1/-1->0->-1 [17] 1/-1/-1->0->-1 [18] 1/-1/-1->0->-1 [19] 1/-1/-1->0->-1 [20] -1/-1/-1->0->1 [21] -1/-1/-1->0->1 [22] -1/-1/-1->0->1 [23] -1/-1/-1->0->1 [24] -1/-1/-1->0->1 [25] -1/-1/-1->0->1 [26] -1/-1/-1->0->1 [27] -1/-1/-1->0->1 [28] -1/-1/-1->0->1 [29] -1/-1/-1->0->1 [30] -1/-1/-1->0->1 [31] -1/-1/-1->0->1 [32] -1/-1/-1->0->1 [33] -1/-1/-1->0->1 [34] -1/-1/-1->0->1 [35] -1/-1/-1->0->1 [36] -1/-1/-1->0->1 [37] -1/-1/-1->0->1 [38] -1/-1/-1->0->1 [39] -1/-1/-1->0->1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO P2P Chunksize set to 131072 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO PROFILER/Plugin: Could not find: libnccl-profiler.so. |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Check P2P Type isAllDirectP2p 0 directMode 0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379053 [0] NCCL INFO [Proxy Service] Device 0 CPU core 51 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379054 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 28 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379055 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 69 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 00/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 01/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 02/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 03/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 04/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 05/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 06/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 07/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 08/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 09/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 10/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 11/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 12/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 13/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 14/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 15/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 16/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 17/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 18/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 19/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 20/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 21/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 22/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 23/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 24/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 25/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Channel 26/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jp |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 00/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 01/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 02/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 03/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 04/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 05/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 06/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 07/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 08/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 09/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 10/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 11/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 12/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 13/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Channel 14/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO NCCL_SOCKET_FAMILY set by environment to AF_INET[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to ib0[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Bootstrap: Using ib0:10.128.33.46<0> |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO cudaDriverVersion 13010 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO NCCL version 2.27.7+cuda13.0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so. |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO NET/IB : Using [0]mlx5_0:1/IB [1]mlx5_1:1/IB [2]mlx5_2:1/IB [3]mlx5_3:1/IB [RO]; OOB ib0:10.128.33.46<0> |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Initialized NET plugin IB |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Assigned NET plugin IB to comm |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Using network IB |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO DMA-BUF is available on GPU device 0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO RAS client listening socket at 127.0.0.1<28028> |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Bootstrap timings total 0.236228 (create 0.000037, send 0.000137, recv 0.233538, ring 0.000056, delay 0.000000) |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO NCCL_CUMEM_ENABLE set by environment to 0. |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO NCCL_NET_GDR_LEVEL set by environment to LOC |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Setting affinity for GPU 0 to 0-71 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO comm 0xaaaaec6bd650 rank 1 nRanks 2 nNodes 2 localRanks 1 localRank 0 MNNVL 0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Trees [0] -1/-1/-1->1->0 [1] -1/-1/-1->1->0 [2] -1/-1/-1->1->0 [3] -1/-1/-1->1->0 [4] -1/-1/-1->1->0 [5] -1/-1/-1->1->0 [6] -1/-1/-1->1->0 [7] -1/-1/-1->1->0 [8] -1/-1/-1->1->0 [9] -1/-1/-1->1->0 [10] -1/-1/-1->1->0 [11] -1/-1/-1->1->0 [12] -1/-1/-1->1->0 [13] -1/-1/-1->1->0 [14] -1/-1/-1->1->0 [15] -1/-1/-1->1->0 [16] -1/-1/-1->1->0 [17] -1/-1/-1->1->0 [18] -1/-1/-1->1->0 [19] -1/-1/-1->1->0 [20] 0/-1/-1->1->-1 [21] 0/-1/-1->1->-1 [22] 0/-1/-1->1->-1 [23] 0/-1/-1->1->-1 [24] 0/-1/-1->1->-1 [25] 0/-1/-1->1->-1 [26] 0/-1/-1->1->-1 [27] 0/-1/-1->1->-1 [28] 0/-1/-1->1->-1 [29] 0/-1/-1->1->-1 [30] 0/-1/-1->1->-1 [31] 0/-1/-1->1->-1 [32] 0/-1/-1->1->-1 [33] 0/-1/-1->1->-1 [34] 0/-1/-1->1->-1 [35] 0/-1/-1->1->-1 [36] 0/-1/-1->1->-1 [37] 0/-1/-1->1->-1 [38] 0/-1/-1->1->-1 [39] 0/-1/-1->1->-1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO P2P Chunksize set to 131072 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO PROFILER/Plugin: Could not find: libnccl-profiler.so. |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972670 [0] NCCL INFO [Proxy Service] Device 0 CPU core 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972671 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 2 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972672 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 3 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2378244 [0] NCCL INFO Comm config Blocking set to 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Assigned NET plugin IB to comm |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Using network IB |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO DMA-BUF is available on GPU device 0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO ncclCommInitRankConfig comm 0xaaab2d76fa20 rank 0 nranks 2 cudaDev 0 nvmlDev 0 busId 901000 commId 0xf87a25994ccf61d3 - Init START |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Bootstrap timings total 0.001121 (create 0.000028, send 0.000097, recv 0.000597, ring 0.000056, delay 0.000000) |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Setting affinity for GPU 0 to 0-71 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO comm 0xaaab2d76fa20 rank 0 nRanks 2 nNodes 2 localRanks 1 localRank 0 MNNVL 0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 00/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 01/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 02/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 03/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 04/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 05/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 06/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 07/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 08/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 09/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 10/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 11/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 12/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 13/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 14/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 15/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 16/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 17/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 18/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 19/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 20/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 21/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 22/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 23/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 24/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 25/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 26/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 27/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 28/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 29/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 30/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 31/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 32/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 33/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 34/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 35/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 36/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 37/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 38/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 39/40 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Trees [0] 1/-1/-1->0->-1 [1] 1/-1/-1->0->-1 [2] 1/-1/-1->0->-1 [3] 1/-1/-1->0->-1 [4] 1/-1/-1->0->-1 [5] 1/-1/-1->0->-1 [6] 1/-1/-1->0->-1 [7] 1/-1/-1->0->-1 [8] 1/-1/-1->0->-1 [9] 1/-1/-1->0->-1 [10] 1/-1/-1->0->-1 [11] 1/-1/-1->0->-1 [12] 1/-1/-1->0->-1 [13] 1/-1/-1->0->-1 [14] 1/-1/-1->0->-1 [15] 1/-1/-1->0->-1 [16] 1/-1/-1->0->-1 [17] 1/-1/-1->0->-1 [18] 1/-1/-1->0->-1 [19] 1/-1/-1->0->-1 [20] -1/-1/-1->0->1 [21] -1/-1/-1->0->1 [22] -1/-1/-1->0->1 [23] -1/-1/-1->0->1 [24] -1/-1/-1->0->1 [25] -1/-1/-1->0->1 [26] -1/-1/-1->0->1 [27] -1/-1/-1->0->1 [28] -1/-1/-1->0->1 [29] -1/-1/-1->0->1 [30] -1/-1/-1->0->1 [31] -1/-1/-1->0->1 [32] -1/-1/-1->0->1 [33] -1/-1/-1->0->1 [34] -1/-1/-1->0->1 [35] -1/-1/-1->0->1 [36] -1/-1/-1->0->1 [37] -1/-1/-1->0->1 [38] -1/-1/-1->0->1 [39] -1/-1/-1->0->1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO P2P Chunksize set to 131072 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Check P2P Type isAllDirectP2p 0 directMode 0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379225 [0] NCCL INFO [Proxy Service] Device 0 CPU core 52 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379226 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 27 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379229 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 37 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 00/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 01/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 02/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 03/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 04/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 05/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 06/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 07/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 08/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 09/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 10/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 11/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 12/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 13/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 14/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 15/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 16/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 17/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 18/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 19/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 20/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 21/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 22/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 23/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 24/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 25/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 26/0 : 1[0] -> 0[0] [send] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 27/0 : 1[0] -> 0[0] [send] via NET/IB/ |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690574) [36m(RayWorkerWrapper pid=1695347)[0m INFO 04-24 06:39:44 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690574) INFO 04-24 06:39:45 [core.py:278] init engine (profile, create kv cache, warmup model) took 6.55 seconds[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-0 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO Channel 26/0 : 1[1] -> 0[3] [receive] via NET/ |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NC |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690574) WARNING 04-24 06:39:47 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690574) INFO 04-24 06:39:47 [vllm.py:690] Asynchronous scheduling is disabled.[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690574) WARNING 04-24 06:39:47 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1690483, ip=10.128.33.36)[0m (EngineCore_DP0 pid=1690574) INFO 04-24 06:39:47 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m WARNING 04-24 06:39:49 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.[32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO ncclCommInitRank comm 0xaaab399cebc0 rank 0 nranks 2 cudaDev 0 nvmlDev 3 busId 3901000 commId 0x3bfe8e79642bfe68 - Init START[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m WARNING 04-24 06:39:16 [custom_all_reduce.py:92] Custom allreduce is disabled because this process group spans across nodes.[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 39/0 : 0[3] -> 1[1] [receive] via NET/IB/1[32m [repeated 40x across cluster][0m[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868654)[0m WARNING 04-24 06:39:11 [system_utils.py:37] Overwriting environment variable CUDA_VISIBLE_DEVICES from '0,1,2,3' to '2,3'[32m [repeated 20x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868654)[0m WARNING 04-24 06:39:11 [system_utils.py:37] Overwriting environment variable LD_LIBRARY_PATH from '/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib' to '/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib'[32m [repeated 20x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868654)[0m INFO 04-24 06:39:16 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'set_numa_affinity', 'test_rpc'][32m [repeated 20x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868654)[0m WARNING 04-24 06:39:16 [worker_base.py:307] Missing `shared_worker_lock` argument from executor. This argument is needed for mm_processor_cache_type='shm'.[32m [repeated 20x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868654)[0m INFO 04-24 06:39:17 [parallel_state.py:1234] world_size=2 rank=1 local_rank=1 distributed_init_method=tcp: |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868681)[0m INFO 04-24 06:39:17 [pynccl.py:111] vLLM is using nccl==2.27.7[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868654)[0m INFO 04-24 06:39:19 [parallel_state.py:1445] rank 1 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 1, EP rank N/A[32m [repeated 20x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868681)[0m INFO 04-24 06:39:20 [gpu_model_runner.py:4125] Starting to load model /e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137...[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868681)[0m INFO 04-24 06:39:20 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].[32m [repeated 12x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868681)[0m INFO 04-24 06:39:43 [default_loader.py:293] Loading weights took 21.78 seconds[32m [repeated 12x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) [36m(RayWorkerWrapper pid=3868681)[0m INFO 04-24 06:39:43 [gpu_model_runner.py:4222] Model loading took 30.59 GiB memory and 22.829386 seconds[32m [repeated 12x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867661, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867869) [36m(RayWorkerWrapper pid=3884173)[0m INFO 04-24 06:39:48 [gpu_worker.py:373] Available KV cache memory: 47.96 GiB[32m [repeated 11x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) INFO 04-24 06:39:48 [kv_cache_utils.py:1307] GPU KV cache size: 392,864 tokens[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=3867788, ip=10.128.33.34)[0m (EngineCore_DP0 pid=3867877) INFO 04-24 06:39:48 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 11.99x[32m [repeated 10x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2362727 [0] NCCL INFO NCCL_SOCKET_FAMILY set by environment to AF_INET[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2362727 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to ib0[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO Bootstrap: Using ib0:10.128.33.46<0>[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO cudaDriverVersion 13010[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO NCCL version 2.27.7+cuda13.0[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO NET/Plugin: Could not find: libnccl-net.so. [32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO NET/IB : Using [0]mlx5_0:1/IB [1]mlx5_1:1/IB [2]mlx5_2:1/IB [3]mlx5_3:1/IB [RO]; OOB ib0:10.128.33.46<0>[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO Initialized NET plugin IB[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Assigned NET plugin IB to comm[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Using network IB[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO DMA-BUF is available on GPU device 0[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO RAS client listening socket at 127.0.0.1<28028>[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Bootstrap timings total 0.000624 (create 0.000027, send 0.000099, recv 0.000258, ring 0.000072, delay 0.000000)[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO NCCL_CUMEM_ENABLE set by environment to 0.[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO NCCL_NET_GDR_LEVEL set by environment to LOC[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Setting affinity for GPU 1 to 72-143[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO comm 0xaaab24408ae0 rank 1 nRanks 2 nNodes 2 localRanks 1 localRank 0 MNNVL 0[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 39/40 : 0 1[32m [repeated 80x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Trees [0] -1/-1/-1->1->0 [1] -1/-1/-1->1->0 [2] -1/-1/-1->1->0 [3] -1/-1/-1->1->0 [4] -1/-1/-1->1->0 [5] -1/-1/-1->1->0 [6] -1/-1/-1->1->0 [7] -1/-1/-1->1->0 [8] -1/-1/-1->1->0 [9] -1/-1/-1->1->0 [10] -1/-1/-1->1->0 [11] -1/-1/-1->1->0 [12] -1/-1/-1->1->0 [13] -1/-1/-1->1->0 [14] -1/-1/-1->1->0 [15] -1/-1/-1->1->0 [16] -1/-1/-1->1->0 [17] -1/-1/-1->1->0 [18] -1/-1/-1->1->0 [19] -1/-1/-1->1->0 [20] 0/-1/-1->1->-1 [21] 0/-1/-1->1->-1 [22] 0/-1/-1->1->-1 [23] 0/-1/-1->1->-1 [24] 0/-1/-1->1->-1 [25] 0/-1/-1->1->-1 [26] 0/-1/-1->1->-1 [27] 0/-1/-1->1->-1 [28] 0/-1/-1->1->-1 [29] 0/-1/-1->1->-1 [30] 0/-1/-1->1->-1 [31] 0/-1/-1->1->-1 [32] 0/-1/-1->1->-1 [33] 0/-1/-1->1->-1 [34] 0/-1/-1->1->-1 [35] 0/-1/-1->1->-1 [36] 0/-1/-1->1->-1 [37] 0/-1/-1->1->-1 [38] 0/-1/-1->1->-1 [39] 0/-1/-1->1->-1[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO P2P Chunksize set to 131072[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO PROFILER/Plugin: Could not find: libnccl-profiler.so. [32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Check P2P Type isAllDirectP2p 0 directMode 0[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379227 [0] NCCL INFO [Proxy Service] Device 0 CPU core 74[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379228 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 108[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379230 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 135[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2362727 [0] NCCL INFO Channel 39/0 : 0[3] -> 1[1] [receive] via NET/IB/1[32m [repeated 40x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 27/0 : 1[1] -> 0[3] [send] via[32m [repeated 43x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO NCCL_SOCKET_FAMILY set by environment to AF_INET[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1956420 [0] NCCL INFO NCCL_SOCKET_IFNAME set by environment to ib0[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2362727 [0] NCCL INFO Comm config Blocking set to 1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO ncclCommInitRankConfig comm 0xaaab24408ae0 rank 1 nranks 2 cudaDev 0 nvmlDev 1 busId 1901000 commId 0xc0090a9fe9957abd - Init START |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) [36m(RayWorkerWrapper pid=1994958)[0m INFO 04-24 06:39:52 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.[32m [repeated 11x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) INFO 04-24 06:39:53 [core.py:278] init engine (profile, create kv cache, warmup model) took 6.19 seconds[32m [repeated 11x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) WARNING 04-24 06:39:54 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 11x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) INFO 04-24 06:39:54 [vllm.py:690] Asynchronous scheduling is disabled.[32m [repeated 11x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) WARNING 04-24 06:39:54 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.[32m [repeated 11x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m (EngineCore_DP0 pid=1994319) INFO 04-24 06:39:54 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 11x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1994108, ip=10.128.33.45)[0m WARNING 04-24 06:39:54 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.[32m [repeated 5x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Loading model from /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints/global_step_18/policy/model_world_size_16_rank_0.pt |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Loading extra_state from /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints/global_step_18/policy/extra_state_world_size_16_rank_0.pt |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Loading optim from /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints/global_step_18/policy/optim_world_size_16_rank_0.pt |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723126)[0m WARNING 04-24 06:39:12 [system_utils.py:37] Overwriting environment variable CUDA_VISIBLE_DEVICES from '0,1,2,3' to '2,3'[32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723126)[0m WARNING 04-24 06:39:12 [system_utils.py:37] Overwriting environment variable LD_LIBRARY_PATH from '/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib' to '/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib'[32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723126)[0m INFO 04-24 06:39:17 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'set_numa_affinity', 'test_rpc'][32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723126)[0m WARNING 04-24 06:39:17 [worker_base.py:307] Missing `shared_worker_lock` argument from executor. This argument is needed for mm_processor_cache_type='shm'.[32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723126)[0m INFO 04-24 06:39:18 [parallel_state.py:1234] world_size=2 rank=1 local_rank=1 distributed_init_method=tcp://127.0.0.1:56411 backend=nccl[32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723127)[0m INFO 04-24 06:39:19 [pynccl.py:111] vLLM is using nccl==2.27.7[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723126)[0m INFO 04-24 06:39:20 [parallel_state.py:1445] rank 1 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 1, EP rank N/A[32m [repeated 8x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723127)[0m INFO 04-24 06:39:21 [gpu_model_runner.py:4125] Starting to load model /e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137...[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723127)[0m INFO 04-24 06:39:22 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723127)[0m INFO 04-24 06:39:49 [default_loader.py:293] Loading weights took 26.10 seconds[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) [36m(RayWorkerWrapper pid=1723127)[0m INFO 04-24 06:39:49 [gpu_model_runner.py:4222] Model loading took 30.59 GiB memory and 27.341083 seconds[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706902) [36m(RayWorkerWrapper pid=1723202)[0m INFO 04-24 06:39:54 [gpu_worker.py:373] Available KV cache memory: 47.96 GiB[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) INFO 04-24 06:39:55 [kv_cache_utils.py:1307] GPU KV cache size: 392,864 tokens[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706814, ip=10.128.33.35)[0m (EngineCore_DP0 pid=1706894) INFO 04-24 06:39:55 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 11.99x[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688118)[0m INFO 04-24 06:39:57 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) INFO 04-24 06:39:58 [core.py:278] init engine (profile, create kv cache, warmup model) took 6.75 seconds[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) WARNING 04-24 06:40:00 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) INFO 04-24 06:40:00 [vllm.py:690] Asynchronous scheduling is disabled.[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) WARNING 04-24 06:40:00 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored.[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) INFO 04-24 06:40:00 [vllm.py:846] Cudagraph is disabled under eager mode[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1706685, ip=10.128.33.35)[0m WARNING 04-24 06:39:57 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.[32m [repeated 3x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Successfully loaded model state dict |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Successfully loaded optimizer state |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Successfully loaded scheduler state |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688123)[0m WARNING 04-24 06:39:16 [system_utils.py:37] Overwriting environment variable CUDA_VISIBLE_DEVICES from '0,1,2,3' to '0,1'[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688123)[0m WARNING 04-24 06:39:16 [system_utils.py:37] Overwriting environment variable LD_LIBRARY_PATH from '/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib' to '/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/e/scratch/jureap59/feuer1/OpenThoughts-Agent/envs/rl/lib/python3.12/site-packages/cv2/../../lib64:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real/:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nvshmem/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/comm_libs/13.0/nccl/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/math_libs/13.0/targets/sbsa-linux/lib:/e/software/default/stages/2026/software/nvidia-compilers/25.9-CUDA-13/Linux_aarch64/25.9/compilers/lib:/e/software/default/stages/2026/software/numactl/2.0.19-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/GCCcore/14.3.0/lib64:/e/software/default/stages/2026/software/CUDA/13/nvvm/lib64:/e/software/default/stages/2026/software/CUDA/13/extras/CUPTI/lib64:/e/software/default/stages/2026/software/CUDA/13/targets/sbsa-linux/lib:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/.singularity.d/libs:/usr/local/cuda/compat/lib.real:/usr/local/cuda-13/lib64/stubs:/usr/local/cuda/lib64/stubs:/usr/local/cuda-13/compat/lib.real:/e/software/default/stages/2026/software/binutils/2.44-GCCcore-14.3.0/lib:/e/software/default/stages/2026/software/zlib/1.3.1-GCCcore-14.3.0/lib'[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688118)[0m INFO 04-24 06:39:21 [worker_base.py:289] Injected <class 'skyrl_train.inference_engines.vllm.vllm_engine.WorkerWrap'> into <class 'vllm.v1.worker.gpu_worker.Worker'> for extended collective_rpc calls ['_apply_fp8_weight_loader_patches', '_is_fp8_model', '_quantize_weights_for_fp8', '_restore_param_subclasses', '_undo_param_subclasses', 'begin_weight_update', 'destroy_weights_update_group', 'end_weight_update', 'init_weight_update_communicator', 'load_weights', 'set_numa_affinity', 'test_rpc'][32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688118)[0m WARNING 04-24 06:39:21 [worker_base.py:307] Missing `shared_worker_lock` argument from executor. This argument is needed for mm_processor_cache_type='shm'.[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688123)[0m INFO 04-24 06:39:22 [parallel_state.py:1234] world_size=2 rank=1 local_rank=1 distributed_init_method=tcp: |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688118)[0m INFO 04-24 06:39:22 [pynccl.py:111] vLLM is using nccl==2.27.7[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688123)[0m INFO 04-24 06:39:24 [parallel_state.py:1445] rank 1 in world size 2 is assigned as DP rank 0, PP rank 0, PCP rank 0, TP rank 1, EP rank N/A[32m [repeated 4x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688118)[0m INFO 04-24 06:39:25 [gpu_model_runner.py:4125] Starting to load model /e/data1/datasets/playground/ot-baf/hf_hub/models--Qwen--Qwen3-32B/snapshots/9216db5781bf21249d130ec9da846c4624c16137...[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688118)[0m INFO 04-24 06:39:26 [cuda.py:367] Using FLASH_ATTN attention backend out of potential backends: ['FLASH_ATTN', 'FLASHINFER', 'TRITON_ATTN', 'FLEX_ATTENTION'].[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688118)[0m INFO 04-24 06:39:51 [default_loader.py:293] Loading weights took 24.70 seconds[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) [36m(RayWorkerWrapper pid=1688118)[0m INFO 04-24 06:39:51 [gpu_model_runner.py:4222] Model loading took 30.59 GiB memory and 25.752638 seconds[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(RayWorkerWrapper pid=1688122)[0m INFO 04-24 06:39:57 [gpu_worker.py:373] Available KV cache memory: 47.96 GiB[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) INFO 04-24 06:39:57 [kv_cache_utils.py:1307] GPU KV cache size: 392,864 tokens[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671655, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671863) INFO 04-24 06:39:57 [kv_cache_utils.py:1312] Maximum concurrency for 32,768 tokens per request: 11.99x[32m [repeated 2x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) [36m(RayWorkerWrapper pid=1688122)[0m INFO 04-24 06:39:57 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled. |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) INFO 04-24 06:39:58 [core.py:278] init engine (profile, create kv cache, warmup model) took 6.82 seconds |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) WARNING 04-24 06:40:00 [serial_utils.py:57] Allowing insecure serialization using pickle due to VLLM_ALLOW_INSECURE_SERIALIZATION=1 |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) INFO 04-24 06:40:00 [vllm.py:690] Asynchronous scheduling is disabled. |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) WARNING 04-24 06:40:00 [vllm.py:735] Inductor compilation was disabled by user settings, optimizations settings that are only active during inductor compilation will be ignored. |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m (EngineCore_DP0 pid=1671871) INFO 04-24 06:40:00 [vllm.py:846] Cudagraph is disabled under eager mode |
| [36m(AsyncVLLMInferenceEngine pid=1671783, ip=10.128.33.47)[0m WARNING 04-24 06:40:00 [model.py:1350] Default vLLM sampling parameters have been overridden by the model's `generation_config.json`: `{'temperature': 0.6, 'top_k': 20, 'top_p': 0.95}`. If this is not intended, please relaunch vLLM instance with `--generation-config vllm`.[32m [repeated 2x across cluster][0m |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Checkpoint loaded successfully from /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints/global_step_18/policy |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m INFO Channel 18/0 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151509 [0] NCCL INFO Channel 19/0 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151509 [0] NCCL INFO Channel 08/0 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151509 [0] NCCL INFO Channel 09/0 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151509 [0] NCCL INFO Channel 20/0 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151509 [0] NCCL INFO Channel 21/0 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151509 [0] NCCL INFO Connected all trees |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151521 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151509 [0] NCCL INFO threadThresholds 8/8/64 | 32/8/64 | 512 | 512 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151509 [0] NCCL INFO 24 coll channels, 24 collnet channels, 0 nvls channels, 32 p2p channels, 16 p2p channels per peer |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151509 [0] NCCL INFO CC Off, workFifoBytes 1048576 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151509 [0] NCCL INFO ncclCommInitRankConfig comm 0x400eb3980500 rank 0 nranks 4 cudaDev 0 nvmlDev 0 busId 901000 commId 0x9ec23970e326a807 - Init COMPLETE |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151509 [0] NCCL INFO Init timings - ncclCommInitRankConfig: rank 0 nranks 4 total 0.32 (kernels 0.00, alloc 0.00, bootstrap 0.16, allgathers 0.00, topo 0.02, graphs 0.01, connections 0.14, rest 0.00) |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 01/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 03/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 05/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 07/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 09/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 11/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 13/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 15/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 17/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 19/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 21/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 23/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 25/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 27/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 29/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 31/1 : 0[0] -> 1[1] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 01/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 03/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 05/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 07/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 09/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 11/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 13/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 15/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 17/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 19/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 21/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 23/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 25/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 27/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 29/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 31/1 : 0[0] -> 2[2] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 00/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 02/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 04/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 06/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 08/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 10/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 12/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 14/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 16/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 18/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 20/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 22/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 24/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 26/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 28/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m jpbo-047-09:4151080:4151528 [0] NCCL INFO Channel 30/1 : 0[0] -> 3[3] via P2P/IPC |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m Channel 39/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Connected all rings, use ring PXN 0 GDR 0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Connected binomial trees |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1961339 [0] NCCL INFO Comm config Blocking set to 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Assigned NET plugin IB to comm |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Using network IB |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO DMA-BUF is available on GPU device 0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO ncclCommInitRankConfig comm 0xaaab073f4ce0 rank 1 nranks 2 cudaDev 0 nvmlDev 0 busId 901000 commId 0xf87a25994ccf61d3 - Init START |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Bootstrap timings total 0.000787 (create 0.000061, send 0.000182, recv 0.000179, ring 0.000109, delay 0.000000) |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Setting affinity for GPU 0 to 0-71 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO comm 0xaaab073f4ce0 rank 1 nRanks 2 nNodes 2 localRanks 1 localRank 0 MNNVL 0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Trees [0] -1/-1/-1->1->0 [1] -1/-1/-1->1->0 [2] -1/-1/-1->1->0 [3] -1/-1/-1->1->0 [4] -1/-1/-1->1->0 [5] -1/-1/-1->1->0 [6] -1/-1/-1->1->0 [7] -1/-1/-1->1->0 [8] -1/-1/-1->1->0 [9] -1/-1/-1->1->0 [10] -1/-1/-1->1->0 [11] -1/-1/-1->1->0 [12] -1/-1/-1->1->0 [13] -1/-1/-1->1->0 [14] -1/-1/-1->1->0 [15] -1/-1/-1->1->0 [16] -1/-1/-1->1->0 [17] -1/-1/-1->1->0 [18] -1/-1/-1->1->0 [19] -1/-1/-1->1->0 [20] 0/-1/-1->1->-1 [21] 0/-1/-1->1->-1 [22] 0/-1/-1->1->-1 [23] 0/-1/-1->1->-1 [24] 0/-1/-1->1->-1 [25] 0/-1/-1->1->-1 [26] 0/-1/-1->1->-1 [27] 0/-1/-1->1->-1 [28] 0/-1/-1->1->-1 [29] 0/-1/-1->1->-1 [30] 0/-1/-1->1->-1 [31] 0/-1/-1->1->-1 [32] 0/-1/-1->1->-1 [33] 0/-1/-1->1->-1 [34] 0/-1/-1->1->-1 [35] 0/-1/-1->1->-1 [36] 0/-1/-1->1->-1 [37] 0/-1/-1->1->-1 [38] 0/-1/-1->1->-1 [39] 0/-1/-1->1->-1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO P2P Chunksize set to 131072 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972827 [0] NCCL INFO [Proxy Service] Device 0 CPU core 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972828 [0] NCCL INFO [Proxy Service UDS] Device 0 CPU core 2 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972824 [0] NCCL INFO Channel 39/0 : 0[0] -> 1[0] [receive] via NET/IB/0[32m [repeated 40x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Channel 39/0 : 0[0] -> 1[0] [send] via NET/IB/0[32m [repeated 40x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m INFO 04-24 06:39:47 [gpu_worker.py:373] Available KV cache memory: 47.99 GiB |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m INFO 04-24 06:39:47 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled. |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m 0 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) INFO 04-24 06:40:33 [block_pool.py:452] Successfully reset prefix cache |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Connected all trees |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=1961339, ip=10.128.33.46)[0m jpbo-047-14:1961339:1972831 [0] NCCL INFO [Proxy Progress] Device 0 CPU core 4 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO threadThresholds 8/8/64 | 16/8/64 | 512 | 512 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO 40 coll channels, 40 collnet channels, 0 nvls channels, 64 p2p channels, 2 p2p channels per peer |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO CC Off, workFifoBytes 1048576 |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO ncclCommInitRankConfig comm 0xaaab2d76fa20 rank 0 nranks 2 cudaDev 0 nvmlDev 0 busId 901000 commId 0xf87a25994ccf61d3 - Init COMPLETE |
| [36m(AsyncVLLMInferenceEngine pid=2361797, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362007) [36m(RayWorkerWrapper pid=2378244)[0m jpbo-047-05:2378244:2379223 [0] NCCL INFO Init timings - ncclCommInitRankConfig: rank 0 nranks 2 total 0.33 (kernels 0.00, alloc 0.00, bootstrap 0.00, allgathers 0.00, topo 0.04, graphs 0.00, connections 0.29, rest 0.00) |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379215 [1] NCCL INFO Channel 00/0 : 1[3] -> 0[2] via P2P/IPC |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379215 [1] NCCL INFO Channel 01/0 : 1[3] -> 0[2] via P2P/IPC |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379215 [1] NCCL INFO Channel 02/0 : 1[3] -> 0[2] via P2P/IPC |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379215 [1] NCCL INFO Channel 03/0 : 1[3] -> 0[2] via P2P/IPC |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379215 [1] NCCL INFO Channel 04/0 : 1[3] -> 0[2] via P2P/IPC |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379215 [1] NCCL INFO Channel 05/0 : 1[3] -> 0[2] via P2P/IPC |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379215 [1] NCCL INFO Channel 06/0 : 1[3] -> 0[2] via P2P/IPC |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379215 [1] NCCL INFO Channel 07/0 : 1[3] -> 0[2] via P2P/IPC |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379215 [1] NCCL INFO Connected all trees |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379220 [1] NCCL INFO [Proxy Progress] Device 1 CPU core 220 |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379215 [1] NCCL INFO threadThresholds 8/8/64 | 16/8/64 | 512 | 512 |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379215 [1] NCCL INFO 8 coll channels, 8 collnet channels, 0 nvls channels, 8 p2p channels, 8 p2p channels per peer |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379215 [1] NCCL INFO ncclCommInitRankConfig comm 0xaaaaf496ae50 rank 1 nranks 2 cudaDev 1 nvmlDev 3 busId 3901000 commId 0x56b155e92217a124 - Init COMPLETE |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378344)[0m jpbo-047-05:2378344:2379215 [1] NCCL INFO Init timings - ncclCommInitRankConfig: rank 1 nranks 2 total 0.03 (kernels 0.00, alloc 0.00, bootstrap 0.00, allgathers 0.00, topo 0.01, graphs 0.00, connections 0.02, rest 0.00) |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378339)[0m jpbo-047-05:2378339:2379214 [0] NCCL INFO Channel 00/08 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378339)[0m jpbo-047-05:2378339:2379214 [0] NCCL INFO Channel 01/08 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378339)[0m jpbo-047-05:2378339:2379214 [0] NCCL INFO Channel 02/08 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378339)[0m jpbo-047-05:2378339:2379214 [0] NCCL INFO Channel 03/08 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378339)[0m jpbo-047-05:2378339:2379214 [0] NCCL INFO Channel 04/08 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378339)[0m jpbo-047-05:2378339:2379214 [0] NCCL INFO Channel 05/08 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378339)[0m jpbo-047-05:2378339:2379214 [0] NCCL INFO Channel 06/08 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378339)[0m jpbo-047-05:2378339:2379214 [0] NCCL INFO Channel 07/08 : 0 1 |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378339)[0m jpbo-047-05:2378339:2379214 [0] NCCL INFO Check P2P Type isAllDirectP2p 1 directMode 0 |
| [36m(AsyncVLLMInferenceEngine pid=2361929, ip=10.128.33.37)[0m (EngineCore_DP0 pid=2362015) [36m(RayWorkerWrapper pid=2378339)[0m jpbo-047-05:2378339:2379214 [0] NCCL INFO CC Off, workFifoBytes 1048576 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m NET/IB/1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 28/0 : 1[1] -> 0[3] [send] via NET/IB/1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 29/0 : 1[1] -> 0[3] [send] via NET/IB/1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 30/0 : 1[1] -> 0[3] [send] via NET/IB/1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 31/0 : 1[1] -> 0[3] [send] via NET/IB/1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 32/0 : 1[1] -> 0[3] [send] via NET/IB/1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 33/0 : 1[1] -> 0[3] [send] via NET/IB/1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 34/0 : 1[1] -> 0[3] [send] via NET/IB/1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 35/0 : 1[1] -> 0[3] [send] via NET/IB/1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 36/0 : 1[1] -> 0[3] [send] via NET/IB/1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 37/0 : 1[1] -> 0[3] [send] via NET/IB/1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 38/0 : 1[1] -> 0[3] [send] via NET/IB/1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=2362727, ip=10.128.33.37)[0m jpbo-047-05:2362727:2379224 [0] NCCL INFO Channel 39/0 : 1[1] -> 0[3] [send] via NET/IB/1 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m CL INFO Channel 39/0 : 1[1] -> 0[3] [receive] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 00/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 01/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 02/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 03/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 04/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 05/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 06/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 07/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 08/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 09/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 10/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 11/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 12/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 13/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 14/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 15/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 16/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 17/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 18/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 19/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 20/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 21/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 22/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 23/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 24/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 25/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 26/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 27/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 28/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 29/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 30/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 31/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 32/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 33/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 34/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 35/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 36/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 37/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 38/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1955530, ip=10.128.33.46)[0m (EngineCore_DP0 pid=1955613) [36m(RayWorkerWrapper pid=1956420)[0m jpbo-047-14:1956420:1972826 [0] NCCL INFO Channel 39/0 : 0[3] -> 1[1] [send] via NET/IB/3 |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=2232721, ip=10.128.33.42)[0m IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=1702991, ip=10.128.33.39)[0m ET/IB/1 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Saving model to /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints/global_step_19/policy/model_world_size_16_rank_0.pt |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=1714615)[0m Channel 39/0 : 1[0] -> 0[0] [receive] via NET/IB/0 |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO Connected all rings, use ring PXN 0 GDR 1[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO Connected binomial trees[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2006411 [1] NCCL INFO Comm config Blocking set to 1[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO Assigned NET plugin IB to comm[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO Using network IB[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO DMA-BUF is available on GPU device 1[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO ncclCommInitRankConfig comm 0xaaab4dbdfb80 rank 1 nranks 2 cudaDev 1 nvmlDev 3 busId 3901000 commId 0x1cea5b0b33ddab3 - Init START[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO Bootstrap timings total 0.000316 (create 0.000029, send 0.000076, recv 0.000103, ring 0.000010, delay 0.000000)[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO Setting affinity for GPU 3 to 216-287[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO comm 0xaaab4dbdfb80 rank 1 nRanks 2 nNodes 1 localRanks 2 localRank 1 MNNVL 0[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO Trees [0] -1/-1/-1->1->0 [1] -1/-1/-1->1->0 [2] 0/-1/-1->1->-1 [3] 0/-1/-1->1->-1 [4] -1/-1/-1->1->0 [5] -1/-1/-1->1->0 [6] 0/-1/-1->1->-1 [7] 0/-1/-1->1->-1[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO P2P Chunksize set to 524288[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007344 [1] NCCL INFO [Proxy Service] Device 1 CPU core 255[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007345 [1] NCCL INFO [Proxy Service UDS] Device 1 CPU core 258[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=1702991, ip=10.128.33.39)[0m jpbo-047-07:1702991:1715538 [0] NCCL INFO Channel 39/0 : 0[3] -> 1[1] [receive] via NET/IB/1[32m [repeated 40x across cluster][0m[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m jpbo-047-10:2232505:2233405 [0] NCCL INFO Channel 39/0 : 0[3] -> 1[1] [send] via NET/IB/3[32m [repeated 40x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=1702991, ip=10.128.33.39)[0m INFO 04-24 06:39:40 [gpu_worker.py:373] Available KV cache memory: 47.99 GiB[32m [repeated 3x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006410)[0m INFO 04-24 06:39:46 [kernel_warmup.py:44] Skipping FlashInfer autotune because it is disabled.[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) INFO 04-24 06:40:33 [block_pool.py:452] Successfully reset prefix cache[32m [repeated 23x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO Channel 07/0 : 1[3] -> 0[2] via P2P/IPC[32m [repeated 152x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO Connected all trees[32m [repeated 22x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007351 [1] NCCL INFO [Proxy Progress] Device 1 CPU core 267[32m [repeated 22x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO threadThresholds 8/8/64 | 16/8/64 | 512 | 512[32m [repeated 22x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO 8 coll channels, 8 collnet channels, 0 nvls channels, 8 p2p channels, 8 p2p channels per peer[32m [repeated 22x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO ncclCommInitRankConfig comm 0xaaab4dbdfb80 rank 1 nranks 2 cudaDev 1 nvmlDev 3 busId 3901000 commId 0x1cea5b0b33ddab3 - Init COMPLETE[32m [repeated 22x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006411)[0m jpbo-047-08:2006411:2007339 [1] NCCL INFO Init timings - ncclCommInitRankConfig: rank 1 nranks 2 total 0.04 (kernels 0.00, alloc 0.00, bootstrap 0.00, allgathers 0.00, topo 0.01, graphs 0.00, connections 0.03, rest 0.00)[32m [repeated 22x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006410)[0m jpbo-047-08:2006410:2007338 [0] NCCL INFO Channel 07/08 : 0 1[32m [repeated 152x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006410)[0m jpbo-047-08:2006410:2007338 [0] NCCL INFO Check P2P Type isAllDirectP2p 1 directMode 0[32m [repeated 19x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1990057, ip=10.128.33.40)[0m (EngineCore_DP0 pid=1990138) [36m(RayWorkerWrapper pid=2006410)[0m jpbo-047-08:2006410:2007338 [0] NCCL INFO CC Off, workFifoBytes 1048576[32m [repeated 22x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=1698116, ip=10.128.33.39)[0m (EngineCore_DP0 pid=1698322) [36m(RayWorkerWrapper pid=1714615)[0m jpbo-047-07:1714615:1715537 [0] NCCL INFO Channel 39/0 : 0[0] -> 1[0] [send] via NET/IB/0[32m [repeated 52x across cluster][0m |
| [36m(AsyncVLLMInferenceEngine pid=2216121, ip=10.128.33.42)[0m (EngineCore_DP0 pid=2216210) [36m(RayWorkerWrapper pid=2232505)[0m CL INFO Channel 39/0 : 1[1] -> 0[3] [receive] via NET/IB/3 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Saving optim to /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints/global_step_19/policy/optim_world_size_16_rank_0.pt |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Saving extra_state to /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints/global_step_19/policy/extra_state_world_size_16_rank_0.pt |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Checkpoint saved to /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/checkpoints/global_step_19/policy |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Created output directory: /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/exports/global_step_19/policy |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Detected FSDP version: 2 |
| [36m(FSDPPolicyWorkerBase pid=4151080, ip=10.128.33.41)[0m [rank-0]: Successfully saved model to /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2/exports/global_step_19/policy |
| Stopping Ray cluster... |
| Ray cluster stopped |
| [RLJobRunner] Launching trace upload (training exit code: 0): |
| repo_id: DCAgent/a2-rl-stack_jest_v2 |
| job_dir: /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/a2-rl-stack_jest_v2 |
| episodes: last |
| log: /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/logs/a2-rl-stack_jest_v2_trace_upload.log |
| [RLJobRunner] Waiting for trace upload to complete... |
| [RLJobRunner] Trace upload failed with exit code 1. |
| Preserving Ray logs to /e/data1/datasets/playground/ot-baf/a2-rl-stack_jest_v2/ray_logs/ |
| Collecting Ray logs from worker jpbo-047-02... |
| Collecting Ray logs from worker jpbo-047-03... |
| Collecting Ray logs from worker jpbo-047-04... |
| Collecting Ray logs from worker jpbo-047-05... |
| Collecting Ray logs from worker jpbo-047-06... |
| Collecting Ray logs from worker jpbo-047-07... |
| Collecting Ray logs from worker jpbo-047-08... |
| Collecting Ray logs from worker jpbo-047-09... |
| Collecting Ray logs from worker jpbo-047-10... |
| Collecting Ray logs from worker jpbo-047-11... |
| Collecting Ray logs from worker jpbo-047-12... |
| Collecting Ray logs from worker jpbo-047-13... |
| Collecting Ray logs from worker jpbo-047-14... |
| Collecting Ray logs from worker jpbo-047-15... |
| Collecting Ray logs from worker jpbo-047-16... |
| Ray log preservation complete |
| |